Conceptio › Archive › arXiv CS
arXiv CSopen access

DACP: A Scientific Data Access and Collaboration Protocol

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

DACP: A Scientific Data Access and Collaboration Protocol Zhihong Shen1,2,∗ , Xiaojie Zhu1,2 , Zhenjing Cheng1 , Hao Ren1 , Zhaoji Liang1,2 , Changfa Lu1

arXiv:2605.03411v1 [cs.NI] 5 May 2026

Computer Network Information Center, Chinese Academy of Sciences1 , Beijing, China University of Chinese Academy of Sciences2 , Beijing, China {bluejoe, xjzhu, zjcheng, rh, zjliang, luchangfa}@cnic.cn

Abstract—Scientific computing is rapidly entering a dataintensive era. However, existing general-purpose network protocol stacks face limitations in eliminating data silos and improving data accessibility and interoperability, making it difficult to effectively meet the demands of emerging paradigms such as AI4Science. To address these challenges, we propose the Data Access and Collaboration Protocol (DACP). DACP defines the Streaming Data Frame (SDF) as its core data model. Through Unified Resource Identification, columnar stream framing, and a reverse supply mechanism, DACP enables data discovery, in-situ computation, and the streaming return of results across scientific data centers, thereby facilitating efficient cross-domain collaboration. Furthermore, this paper introduces faird, a reference server implementation of DACP. This work provides a viable path for building scalable and collaborative scientific data infrastructures. Index Terms—Application Layer Protocol, Streaming Data Frame, Lazy Loading, In-situ Computation, Cross-domain Collaboration, Research Data Network.

I. I NTRODUCTION In recent years, with the deepening of scientific inquiry, scientific computing is rapidly evolving toward a data-centric paradigm [1]. Domains such as astronomical observation, high-energy physics, and life sciences are continuously generating data at the Petabyte (PB) or even Exabyte (EB) scales. As data becomes a core infrastructure, the capability to acquire, comprehend, and utilize it directly determines the efficiency and boundaries of scientific discovery. Currently, the FAIR principles (Findable, Accessible, Interoperable, Reusable) [2] have become a global consensus for scientific data management. However, current FAIR practices focus primarily on “whether data exists and is describable”. Particularly in cross-disciplinary, cross-institutional, and crossdata-center scenarios, data accessibility and interoperability still have significant room for improvement. Emerging research paradigms like AI4Science increasingly depend on joint training and inference over large-scale, multimodal data [3]. However, in practice, scientific data remains dispersed across fragmented systems characterized by heterogeneous formats, interfaces, and permission models, leading to severe data silos. This landscape of dispersion, heterogeneity, and siloization [4] [5] has emerged as a critical bottleneck hindering the advancement of these new paradigms.

A. Problem Statement Although a variety of general-purpose data transfer and access protocols have been proposed and deployed, their direct reuse for scientific data still presents obvious shortcomings: First, the lack of a unified abstraction exacerbates the problem of “Data Silo” [6]. Different data sources are usually exposed as independent service interfaces with inconsistent access methods and naming rules. This leads to the need for extensive customized development at the application layer to achieve cross-source data composition. Second, legacy file-centric models incur severe “Read Amplification” [7]. Researchers often require only partial fields, columns, or sub-regions of a dataset. However, existing protocols usually require downloading the complete file, bringing significant network and storage overheads. Third, data and metadata are fragmented in the access path. Information such as data types, structures, and semantics cannot be transmitted synchronously with the data stream, increasing the complexity of cross-system interoperability. Finally, existing protocols mainly focus on “how to transmit data” while ignoring “where and how data should be processed”. In large-scale data scenarios, data must be fully downloaded to compute nodes, making it difficult to support in-situ computing [8] and task-driven data supply. B. Solution Approach and Contributions To address the aforementioned challenges, the main contributions of this paper are as follows: We proposed DACP, a scientific data access and collaboration protocol. This approach provides a unified, task-oriented, and computation-centric abstraction for scientific data access, enabling efficient data collaboration across Wide Area Networks (WANs). • We presented faird, a scalable reference implementation based on the DACP protocol.The source code is publicly available at https://github.com/rdcn-link/faird. •

The remainder of this paper is organized as follows: Section II analyzes existing limitations. Sections III and IV detail the DACP design and faird implementation. Section V evaluates the performance, followed by application scenarios in Section VI and conclusions in Section VII.

TABLE I C OMPARISON OF C OMPUTER N ETWORK , WWW, L INKED DATA , AND R ESEARCH DATA N ETWORK (A DAPTED FROM [5]) Network Type

Primary Objective

Core Element

Core Technology

Computer Network [9]

Shield heterogeneity of computing devices; provide interoperability between devices

Devices

IP, TCP, etc.

World Wide Web (WWW) [10]

Shield heterogeneity of documents; provide interoperability between documents

Documents

URL, HTTP, HTML, Hyperlink, etc.

Linked Data [11]

Shield heterogeneity of data (information); provide interoperability between data

Resource Description

URI, HTTP, RDF, RDF link, etc.

Research Data Network [5]

Shield heterogeneity of research data; provide interoperability between data and services

Research Data

Identification, Addressing, Access, and Operation Protocols for research data resources

II. A NALYSIS OF L IMITATIONS IN E XISTING P ROTOCOL S TACKS A Research Data Network [5] constitutes a unified and standardized infrastructure designed for distributed scientific data. As illustrated in Table 1, significant distinctions exist between research data networks and the World Wide Web (WWW) regarding their objectives, core elements, and underlying technologies. This section provides a critical analysis of typical protocol stacks. A. REST/JSON REST/JSON represents the most ubiquitous paradigm for data access and exchange [12]. Consequently, it imposes significant limitations within the context of Research Data Networks: • Serialization Bottlenecks: Scientific data typically consists of dense numerical matrices. As a text-based protocol, JSON incurs substantial CPU overhead due to extensive serialization and deserialization operations [13]. • Type System Deficiency: The JSON type system supports only a generic Number type and lacks the capability to explicitly distinguish between critical scientific types such as int8, uint64, float32, and float16. • Lack of Native Slicing: Accessing high-dimensional data often requires extracting specific columns or sub-regions; however, standard REST APIs typically mandate the transmission of the complete JSON structure. B. FTP/GridFTP FTP [14] and GridFTP [15] were primarily designed for high-throughput bulk data migration. Their architectures are inherently constrained by the file as the fundamental unit of transfer. Under emerging research paradigms characterized by online exploration and on-demand analysis, their limitations are becoming increasingly pronounced: • Read Amplification: Scientific tasks often require only specific columns, rows, or slices within a dataset. Yet, FTP transmission mandates the download of entire CSV or HDF5 [16] files, which leads to the read amplification problem [7]. • Opaque Internal Structure (Schema): The internal data structure (Schema) remains a “black box” to the transport

layer. Users are compelled to download the file header or even the entire file before they can extract metadata. In addition, other modern extension protocols, such as WebDAV [17] and QUIC [18], do not directly address the specific challenges of scientific data access. WebDAV remains fundamentally a file-centric protocol. In large-scale data center scenarios, the overhead associated with parsing its XML-based metadata often becomes a significant performance bottleneck. QUIC is strictly positioned at the Transport Layer. As such, it is agnostic to data models and does not concern itself with application-level logic regarding where or how data should be computed or composed. In summary, current protocol stacks still exhibit certain deficiencies when applied to the accessibility and interoperability of scientific data. Scientists are compelled to spend a disproportionate amount of time locating, downloading, parsing, and organizing data, rather than focusing on the scientific problems themselves. III. DACP P ROTOCOL D ESIGN DACP is architected as an application-layer streaming protocol. Its objective is not to replace underlying transport technologies, but rather to build upon existing network infrastructures to empower scientific data with unified representation, programmable access, and cross-domain collaboration capabilities. This section elaborates on the design of DACP. A. Core Data Model: Streaming Data Frame (SDF) The data model of DACP is designed to abstract away the heterogeneity of underlying storage systems, providing a unified, computation-oriented abstraction. DACP establishes the Streaming Data Frame (SDF) as its fundamental data primitive. The SDF abstracts scientific data content into a logical, streaming two-dimensional tabular structure, exposing a unified access interface. Formally, an SDF is represented as a tuple: D = ⟨S, F⟩

(1)

1) Schema (S): defines the topological structure and column type constraints of the 2D data frame: S = {(attr1 , τ1 ), . . . , (attrm , τm )}

(2)

GET

Iterator<Row>

COOK

PUT ......

SDF Operations

next()

META

ID

0001

Schema

[<Id, Int>, <Path, String>, <Format, String>, <Content, Blob>]

id

Name

00011

img_01. tiff

00012

data.csv

00013

log.txt

Format

Blob

dacp://...

<0xA2...>

dacp://...

text

<0x2D...>

ITERATOR

FILTSR

SELECT

ID

00012

Schema

[<Id, Int>, <ShapeLen, Float>, <fromNode, Int>, <toNode, Int>]

URIRef

tiff <0xFA...> Current Row csv

LENGTH

...

next()

id

ShapeLen

fromNode

toNode

000121

1693.675

1

2

000122

3174.518

3

Current Row

7

000123

2200.843

6

11

LIMIT

...

ID

0002

Schema

[<Id, Int>, <Species, String>, <Year, String>, <Station, String>]

id

Species

Year

Station

00021

Sheepgrass

2009

Inner Mongolia Station

00022

Large Needlegrass

00023

Crested Wheatgrass

Inner Mongolia

2009Row Current Station 2009

Inner Mongolia Station

next()

Fig. 1. The unified logical abstraction and interaction architecture of the Streaming Data Frame (SDF). The figure illustrates the core protocol methods (GET/PUT/COOK), the streaming retrieval mechanism via GET, and the File List Framing strategy. Users can perform coarse-grained filtering on the file list (e.g., Format=’csv’) or recursively ’drill down’ and retrieve specific file contents using the same set of operation primitives.

Here, attri denotes the column name, and τi represents the data type (supporting various scientific primitives such as int, float, and binary). The existence of the Schema eliminates parsing ambiguities inherent in heterogeneous data, ensuring the self-describing nature of the SDF and end-to-end type safety. 2) Frame Stream (F ): The data entity is materialized as an ordered streaming sequence, denoted as F . Each element βk represents a Record Batch, containing a finite set of rows strictly conforming to the constraints of S (i.e., Rows ⊂ βk ). βk serves as the atomic unit for transport and internally adopts a columnar memory layout. This design is critical for enabling Zero-Copy transmission between the network layer and application memory. The SDF is designed to natively support streaming computation. Logically, it exposes an Iterator<Row> semantic to ensure the flexibility of row-level processing. Physically, however, multiple logical rows are encapsulated into βk to amortize network I/O overhead through batching. Once a βk arrives at the client, computation logic is triggered immediately without waiting for subsequent batches βk+1 , thereby accommodating the requirements for high-throughput streaming data processing in WAN environments. An SDF can be mapped to either a single physical file or a list of files. A file list is mapped into a structured SDF, where file metadata is mapped to standard columns, and file content is mapped to a Binary-typed blob column. This blob column possesses an Expandable Capability, allowing unstructured binary content to be read in real-time as a new SDF. Conversely, for structured physical files (e.g., CSV, Parquet), DACP directly encapsulates their internal rows and

columns into a single SDF. B. Core Methods: GET, PUT, and COOK To enable standardized operations on scientific data, DACP defines three fundamental methods: GET, PUT, and COOK. • GET: The method is employed to request existing, static data resources. The client provides the target resource URI, and upon successful authentication, the server returns the content in the form of a Streaming Data Frame (SDF). Crucially, GET supports predicate pushdown, enabling filtering logic to be executed on the server side, thereby circumventing the movement of massive datasets across the network. • PUT: The method is designed to upload and persist client-side data streams to the server. The input is similarly formatted as an SDF, facilitating the efficient, streaming ingestion of data into data centers. • COOK: The method serves as the gateway interface for computational task offloading. Unlike static data retrieval, the COOK method accepts a Directed Acyclic Graph (DAG) as its functional payload. A task is formally represented as G = (V, E), where each vertex v ∈ V denotes a standardized Operator (e.g., Filter, Select), and each edge e ∈ E represents the streaming flow of SDFs. By adopting a non-blocking mechanism, COOK allows clients to consume or visualize the initial batch of results in real-time. C. Resource Addressing DACP adopts a unified resource addressing scheme alongside a phased interaction model to facilitate the precise lo-

calization of scientific resources within heterogeneous environments. A standard URI is defined to uniquely identify an SDF across Wide Area Networks (WANs), with the naming convention specified as follows: dacp://⟨host⟩ : ⟨port⟩/[⟨dataset name⟩]/⟨path⟩

dacp: The protocol scheme identifier. • host:port: The network address and service port of the DACP server. • dataset name: An optional dataset name. A Dataset serves as a logical collection unit for SDFs. It supports the definition of shared metadata or permission policies at the collection level, enabling all enclosed SDFs to automatically inherit this contextual information. • path: The specific resource path of the SDF. This can point to a concrete physical file path or a logical dataset directory. DACP Request SDF URI &Token Flow DAG

SDF

Row1

Global DAG

4. Stream

Row2 ...... Rown

SDF

1. Submit Job(cook) Server A 3. Dispatch & Run

User

Server C dacp://10.0.0.2/data3

Fig. 3. A cross-domain collaborative analysis scenario, where the global logical DAG is decomposed into physical sub-tasks for in-situ execution and synchronized streaming.

which recursively pulls data from upstream dependency nodes; only at this point are the upstream operators activated for execution. Building on this mechanism, data is no longer loaded as monolithic files but flows between operators in units of fine-grained logical rows. This design significantly enhances the processing efficiency for large-scale scientific data. IV. FAIRD : T HE R EFERENCE S ERVER I MPLEMENTATION OF DACP

1. GET 2. Stream

2. Dispatch & Run Server B

(3)

•

DACP Client

dacp://10.0.0.1/data(1&2)

DACP Server

Schema Header RecordBatch 1

Wide-Area Network

... RecordBatch N

Fig. 2. The Client-Server interaction mechanism of DACP.

The DACP interaction process initiates with the client establishing a connection and exchanging credentials to obtain a short-lived access Token. Subsequently, the client issues data requests to server. For static retrieval via GET, the client consumes the SDF through a row-based iterator, strictly adhering to lazy evaluation principles to minimize latency. D. Cross-Domain Collaboration DACP enables cross-domain data collaboration by orchestrating the DAG defined in the COOK method across distributed environments. During the scheduling phase, first DACP provides end-toend task planning, which includes data discovery, acquisition, processing, analysis, and result integration, and maps them to multiple physical sub-tasks. Second, it ensures task execution reliability through DAG-based sub-task scheduling and transaction control. Finally, it establishes mechanisms for task scheduling, monitoring, and fault handling. Regarding security, DACP establishes a token-based data flow mechanism. Downstream nodes must present a short-lived token, issued during the scheduling phase, to pull data from upstream nodes. In the task execution phase, DACP employs a lazy loading and streaming execution paradigm. Upon the client offloading a task to the server, full-scale computation is not triggered immediately. Instead, the actual computation process is initiated by the outermost node of the DAG (i.e., the result consumer),

This paper presents the design and implementation of faird, the reference server for the DACP protocol. In terms of architectural layering, Apache Arrow Flight [19] serves as the underlying Transport Layer, responsible for providing a highperformance, zero-copy data communication channel. Faird is constructed atop this foundation as the Application Layer. This section elaborates on the core components of faird following a bottom-up approach. A. Multimodal Data Source To abstract away the heterogeneity inherent in scientific data formats (e.g., CSV, NetCDF, TIFF, and arbitrary binary formats), faird implements a Multimodal Data Source framework. This abstraction is responsible for mapping specific physical files into logical SDF representations. For structured data, the source leverages memory mapping (mmap) to directly project data into the Apache Arrow format, thereby achieving zero-copy read efficiency. Conversely, for semi-structured or unstructured data, the source adheres to the “File List Framing” strategy described previously, effectively integrating unstructured content into the unified logical SDF view. B. SDF Engine The SDF Engine serves as the core kernel of faird, governing the lifecycle management and logical abstraction of in-memory SDF data. It is responsible not only for mapping physical file types to standard Arrow types but also for defining and maintaining column definitions to facilitate schema-aware computation for downstream operators. Notably, the initialization of an SDF does not trigger immediate data loading; physical data is materialized into memory as Arrow structures only during the DAG execution phase. Additionally, the engine incorporates a built-in library of fundamental operators (e.g.,

consists of a random distribution of 1 large file (1GB), 10 medium files (100MB each), and 10,000 small files (10KB each), aiming to verify the comprehensive transmission efficiency under hybrid loads.

Map, Filter, Select). These operators are capable of executing high-performance batch computing directly on the Arrow columnar in-memory data layout. C. Ecosystem Integration

V. P ERFORMANCE E VALUATION To validate the performance efficiency of the DACP protocol in real-world data transmission scenarios—specifically contrasting it with the traditional FTP protocol—we conducted a series of comparative experiments on a high-performance computing node. A. Experimental Setup The experiments were conducted on a high-performance computing node equipped with an Intel Xeon Gold CPU and 376GB RAM. To minimize storage bottlenecks, the system utilizes a disk array with a sequential throughput of 1202 MB/s. The network environment provides 3.45 Gbps bandwidth with negligible latency (0.2 ms). B. Datasets To comprehensively evaluate DACP across heterogeneous scientific data, we constructed a mixed testbed: • Structured Data: We utilized the Yelp Open Dataset [20], selecting a subset of 1 million user review rows in JSON format. Each row maintains a uniform schema with five key-value pairs. • Unstructured Mixed Data: Based on the ImageNet [21] validation set, we designed a random mixed workload to simulate complex scientific repositories. This dataset

40 20 0

2

4

6

Batch Index

8

Mean Rate (104 rows/s)

Rate (104 rows/s)

60

(b) Download: Overall Performance

100

10

(c) Upload: Batch-wise Rate Mean Rate (104 rows/s)

35 30 25 20 15 10 5 0

2

4

6

Batch Index

8

10

80 60 40 20 0

35 30 25 20 15 10 5 0

DACP

Protocol

FTP

(d) Upload: Overall Performance

DACP

Protocol

FTP

Fig. 4. Performance evaluation on structured data.

(a) Download: Interleaved Files 1400 1200 1000 800 600 400 200 0

Overall Avg. Rate (Rows/s)

The faird client is a lightweight software development kit (SDK) implementation of the DACP protocol, designed to mask the complexity of underlying communications. It does not execute computations directly but provides a chainable API. Users can construct a logical DAG using built-in or user-defined operators. Upon triggering a computation request, the Client serializes the constructed DAG and submits it to the server. For structured data returned by server, the client leverages Apache Arrow’s zero-copy mechanism to convert it directly into native in-memory objects; for unstructured Blob columns, it reconstructs them into file objects or byte streams on the fly.

DACP FTP

80

Rate (104 rows/s)

D. faird Client

(a) Download: Batch-wise Rate 100

Overall Avg. Rate (Rows/s)

To achieve seamless interoperability with mainstream computing frameworks, faird provides native adapters for Apache Spark, PyTorch, and HuggingFace. In the Spark scenario, faird functions as a high-performance data source, automatically mapping read requests to the protocol’s underlying GET or COOK primitives. For AI training, faird feeds data streams directly into the PyTorch pipeline via a streaming adapter. Furthermore, it deeply integrates with HuggingFace, allowing users to load remote data directly via the dacp:// protocol within load_dataset. This capability efficiently supports the training and fine-tuning of large models.

DACP (BLOB)DACP (Binary)

Protocol

FTP

(b) Upload: Interleaved Files 1400 1200 1000 800 600 400 200 0

DACP

Protocol

FTP

Fig. 5. Performance evaluation on unstructured data(mixed).

The experimental results demonstrate that DACP significantly outperforms FTP in overall performance across most scenarios, achieving speedup ratios ranging from 1.21× to 5.3×. Specifically, for structured data, DACP surpasses FTP across all metrics, with speedups between 3.10× and 5.36×. For unstructured data, the performance of DACP (BLOB) is comparable to that of DACP (Binary) and slightly superior to FTP, yielding a speedup of 1.21×. Furthermore, DACP maintains symmetrical download and upload rates, whereas FTP exhibits a 13%–27% degradation in upload rates under random interleaved scenarios. These results indicate that DACP possesses a distinct performance advantage over FTP, particularly in complex workload scenarios. VI. A PPLICATION S CENARIOS A. Efficient and Unified Data Provisioning Addressing the fragmentation of data across relational databases, file systems, and object stores, DACP abstracts these heterogeneous sources into a unified SDF representation.

This provides a consistent access model, allowing scientists to utilize a single codebase and query interface for diverse data types without developing complex adapter layers. B. Interactive Data Exploration DACP significantly optimizes interactive exploration on PBscale datasets by mitigating the bottlenecks of traditional “download-then-process” paradigms. DACP allows users to submit a DAG containing filter logic, enabling the server to return only the requested slices as a Streaming Data Frame (SDF). This “seek-and-stream” capability achieves near realtime responsiveness, drastically improving exploration efficiency. C. Cross DataCenter Collaborative Workflows DACP optimizes cross-data center collaboration by replacing the costly replication of intermediate results with an in-situ computation strategy. Adhering to the principle of moving operators, not data, DACP employs lazy evaluation: upstream computation and Arrow Flight streaming are triggered only upon downstream demand. This ensures that only filtered, high-value data streams—rather than raw massive datasets—traverse the WAN. VII. C ONCLUSION In response to the limitations of legacy protocols (e.g., HTTP/FTP) in supporting data-intensive paradigms like AI4Science, this paper proposes DACP, a Data Access and Collaboration Protocol tailored for scientific discovery. DACP redefines the semantics of scientific data access. By establishing the Streaming Data Frame (SDF) as a standard unit, it effectively abstracts the heterogeneity of underlying storage systems. Through unified resource addressing, columnar streaming, and reverse provisioning mechanisms, DACP enables cross-domain collaboration via data discovery, insitu computation, and result streaming. We presented faird, a reference implementation that validates the feasibility of DACP in building collaborative infrastructures, with experiments demonstrating its superior performance across various scenarios. In the future, DACP aims to empower scientists to access and analyze remote data as transparently and efficiently as local resources, thereby fully unleashing the potential of data elements in scientific discovery. ACKNOWLEDGMENT This work is supported by the National Key R&D Program of China (Grant No. 2021YFF0704200), and the Informatization Plan of Chinese Academy of Sciences (Grant No. CASWX2022GC-02). R EFERENCES [1] T. Hey, S. Tansley, K. M. Tolle et al., The fourth paradigm: dataintensive scientific discovery. Microsoft research Redmond, WA, 2009, vol. 1. [2] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne et al., “The fair guiding principles for scientific data management and stewardship,” Scientific data, vol. 3, no. 1, pp. 1–9, 2016.

[3] H. Wang, T. Fu, Y. Du, W. Gao, K. Huang, Z. Liu, P. Chandak, S. Liu, P. Van Katwyk, A. Deac et al., “Scientific discovery in the age of artificial intelligence,” Nature, vol. 620, no. 7972, pp. 47–60, 2023. [4] A. Halevy, A. Rajaraman, and J. Ordille, “Data integration: The teenage years,” in Proceedings of the 32nd international conference on Very large data bases, 2006, pp. 9–16. [5] S. Zhihong, Z. Xiaojie, W. Huajin, T. Jizhou, G. Xuebing, W. Hui, M. Yufang, and W. Linhuan, “Research data network: Concept, systems and applications,” Frontiers of Data and Computing, vol. 6, no. 4, pp. 3–21, 2024. [6] C. P. Chen and C.-Y. Zhang, “Data-intensive applications, challenges, techniques and technologies: A survey on big data,” Information sciences, vol. 275, pp. 314–347, 2014. [7] R. P. Abernathey, T. Augspurger, A. Banihirwe, C. C. Blackmon-Luca, T. J. Crone, C. L. Gentemann, J. J. Hamman, N. Henderson, C. Lepore, T. A. McCaie et al., “Cloud-native repositories for big scientific data,” Computing in Science & Engineering, vol. 23, no. 2, pp. 26–35, 2021. [8] R. Balasubramonian, J. Chang, T. Manning, J. H. Moreno, R. Murphy, R. Nair, and S. Swanson, “Near-data processing: Insights from a micro46 workshop,” IEEE Micro, vol. 34, no. 4, pp. 36–42, 2014. [9] V. Cerf and R. Kahn, “A protocol for packet network intercommunication,” IEEE Transactions on communications, vol. 22, no. 5, pp. 637– 648, 1974. [10] T. Berners-Lee, R. Cailliau, A. Luotonen, H. F. Nielsen, and A. Secret, “The world-wide web,” Communications of the ACM, vol. 37, no. 8, pp. 76–82, 1994. [11] N. Shadbolt, T. Berners-Lee, and W. Hall, “The semantic web revisited,” IEEE intelligent systems, vol. 21, no. 3, pp. 96–101, 2006. [12] R. T. Fielding and R. N. Taylor, “Principled design of the modern web architecture,” ACM Transactions on Internet Technology (TOIT), vol. 2, no. 2, pp. 115–150, 2002. [13] K. Maeda, “Performance evaluation of object serialization libraries in xml, json and binary formats,” in 2012 Second International Conference on Digital Information and Communication Technology and it’s Applications (DICTAP). IEEE, 2012, pp. 177–182. [14] M. Gien, “A file transfer protocol (ftp),” Computer Networks (1976), vol. 2, no. 4-5, pp. 312–319, 1978. [15] W. Allcock, J. Bresnahan, R. Kettimuthu, and M. Link, “The globus striped gridftp framework and server,” in SC’05: Proceedings of the 2005 ACM/IEEE conference on Supercomputing. IEEE, 2005, pp. 54– 54. [16] S. Koranne, “Hierarchical data format 5: Hdf5,” in Handbook of open source tools. Springer, 2010, pp. 191–200. [17] L. Dusseault, “Http extensions for web distributed authoring and versioning (webdav),” Tech. Rep., 2007. [18] A. Langley, A. Riddoch, A. Wilk, A. Vicente, C. Krasic, D. Zhang, F. Yang, F. Kouranov, I. Swett, J. Iyengar et al., “The quic transport protocol: Design and internet-scale deployment,” in Proceedings of the conference of the ACM special interest group on data communication, 2017, pp. 183–196. [19] The Apache Software Foundation, “Apache arrow: A cross-language development platform for in-memory data,” 2024, accessed: 2025-12-20. [Online]. Available: https://arrow.apache.org/ [20] Yelp Inc., “Yelp Open Dataset,” 2025, accessed: 2025-12-29. [Online]. Available: https://www.yelp.com/dataset [21] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009, pp. 248–255.

Record · ID 155183 · SHA-256 ca5162ccbac4fb56
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.