ConceptioArchivearXiv CS
arXiv CSopen access

RAIDS: Rethinking Data Systems as Responsible Intelligent Infrastructure

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
databasesdatamanagementsqlstorage
databases, sql, data management, storage

RAIDS: Rethinking Data Systems as Responsible Intelligent Infrastructure

arXiv:2606.21831v1 [cs.DB] 20 Jun 2026

Zhengyi Yang

Wenke Yang

Guanfeng Liu

Lu Qin

University of Sydney University of New South Wales [email protected] [email protected]

Macquarie University University of Technology Sydney [email protected] [email protected]

Abstract—Data systems are evolving from information infrastructure into decision infrastructure. Yet responsibility mechanisms have not kept pace: an output can be accurate or efficient while still lacking sufficient support, satisfied constraints, and actionability for responsible use. We propose RAIDS (Responsible and Intelligent Data System), a vision for data systems as responsible intelligent infrastructure. RAIDS treats responsibility not as post-hoc metadata, but as execution semantics for holistic data-to-decision and data mining pipelines. Its core abstraction is an operator-level responsibility contract: each operator exposes an output together with support, constraint, and actionability state under an explicit responsibility context, and these contracts compose across pipelines. These states capture whether an output is grounded, whether execution satisfies relevant limits, and which action modes are permissible. We introduce responsibility preservation as the organizing systems objective: responsibility state should remain sufficient as execution proceeds, or the system should repair, replan, escalate, refuse, or otherwise change course. We outline a BlueSky research agenda for RAIDS, spanning responsibility-preserving execution, responsibility-aware optimization, provenance, oversight, and evaluation. Index Terms—responsible data intelligence, data systems, data mining, responsible AI, provenance, sustainability

The responsibility gap. Emerging efforts have raised responsibility-related concerns in data systems, including fairness and transparency [33], [34], governance and accountability [35], [36], privacy and data protection [37], [38], sustainability [39], safety, and human oversight [32], [40]. Existing systems also expose fragments of responsibility through mechanisms such as data validation [41], documentation [42], [43], provenance [44], [45], and auditing [36], [46], [47]. These mechanisms provide necessary visibility, but they often sit beside execution: their outputs are recorded, documented, or checked without participating in runtime decisions, control flow, or action protocols. As outputs move across operators, models, teams, and deployment boundaries, downstream actors must reconstruct trust, constraints, uncertainty, and appropriate use after the fact. Evidence may be lost, policy constraints may be dropped, environmental impacts may remain invisible, and human review may never be triggered. The responsibility gap is not the absence of another checklist or governance component; it is a systems gap: responsibility is not yet preserved and acted on throughout the data-to-decision process.

I. I NTRODUCTION Data systems are evolving from information infrastructure into decision infrastructure. Beyond storing, querying, and analyzing data, they are increasingly used to support realworld decision making across critical domains including health care [1]–[3], finance [4]–[6], public services [7]–[11], scientific discovery [12]–[14], climate response [15]–[17], and AI-assisted work [18]–[20]. This transformation spans the entire data intelligence stack, from database systems [21]– [23] and data mining [24]–[27] to emerging LLM-powered data analytics and applications [28]–[31]. Capability requires responsibility. As data systems become embedded in real-world decision making, their outputs carry consequences beyond information access and analysis. In critical domains, system outputs can influence decisions about care, credit, public resources, scientific priorities, environmental response, and everyday work [32]. A pipeline can no longer be judged solely by efficiency, scalability, accuracy, or utility; it must also be judged by whether its outputs are appropriate for use in context. A technically strong result may still be unsuitable when it is unsupported by sufficient evidence, violates applicable constraints or regulations, imposes unnecessary environmental costs, or requires human judgment before action. Together, these developments point to a broader question for the community: how should systems determine whether an output is responsible enough to use?

Why now. Three shifts make this question timely. Data mining, AI analytics, and LLM-empowered pipelines now turn query results and mined patterns directly into claims, recommendations, and actions. At the same time, privacy, fairness, safety, sustainability, auditability, and human oversight increasingly determine whether an output can be used, not merely documented. Modern data infrastructure already exposes the needed hooks: metadata, validation, provenance, policy checks, monitoring, optimization, and control flow. The BlueSky question is whether these hooks can become execution semantics: can responsibility change what a pipeline computes, preserves, repairs, escalates, or refuses? The RAIDS idea. We propose RAIDS (Responsible and Intelligent Data System), a vision for responsibility-native data systems. The central claim is responsibility as execution semantics: responsibility should participate in execution rather than remain external to it. A data system should expose not only an output, but whether it has sufficient support, satisfies relevant constraints, and has appropriate actionability in context. RAIDS asks which pipelines are admissible, which plans are preferable, what evidence must be preserved, when outputs should be qualified, and when execution should repair, escalate, refuse, audit, invite review, or stop execution. Responsibility is not an after-the-fact assessment but a firstclass systems concern throughout the data-to-decision process.

Contributions. The paper makes three contributions. Together, they recast responsibility as a systems problem and define a research agenda for responsibility-native data systems. • Responsibility-native infrastructure. We argue that as data systems evolve from information infrastructure into decision infrastructure, responsibility must become a native systems concern. We recast modern data systems as a Responsible Data Intelligence Loop, in which responsibility is preserved throughout the data-to-decision process. • Executable responsibility contract. We introduce a responsibility contract as a systems primitive for responsibility as execution semantics. The contract exposes support, constraint, and actionability under an explicit responsibility context, enabling responsibility to be propagated, optimized, monitored, audited, and used to govern execution. • A BlueSky research agenda. We outline a research agenda organized around responsibility preservation: sufficient support, satisfied constraints, and appropriate actionability throughout a data-to-decision pipeline. The agenda connects execution, optimization, provenance, oversight, and evaluation within a unified systems perspective. II. L IMITS OF C URRENT A PPROACHES From visibility to execution. Prior work provides building blocks for RAIDS, but the dominant pattern remains visibility rather than execution. Capability-oriented systems make data infrastructure more automated, adaptive, and model-mediated: deployed data and AI pipelines expose lifecycle debt [48], [49]; self-driving databases automate system decisions [50]; and LLM-facing interfaces [28], semantic operators [29], and vector substrates [23] move retrieval, interpretation, and generation into data management. These advances expand capability; RAIDS asks what responsibility state must accompany outputs as they become inputs to action? Foundational but fragmented. Responsibility-oriented work exposes information for responsible use. Responsible data management frames fairness, transparency, governance, and accountability as data-system concerns [33]; data governance identifies decision rights, responsibilities, and contextdependent coordination mechanisms [35], [51]–[53]. Provenance captures derivations and traces [44], [45], [54]; auditing and influence analysis expose accountability gaps and input effects [36], [55]; documentation records intended use and limitations [42], [43]; and machine-readable metadata supports responsible-use fields [56], [57]. Fairness-aware repair [58], integration [59], and ranking [60], [61] expose social constraints; sustainability work exposes environmental costs [39]; validation and repair expose quality failures [41], [62], [63]; label-quality analysis exposes annotation failures [64]; and retrieval-augmented generation (RAG) evaluation exposes evidence-use failures [65]. Together, these streams make responsibility visible, documented, explainable, and auditable, but they remain disconnected fragments. RAIDS asks when this information should become part of execution itself. The missing coordination. Current approaches lack endto-end coordination, not responsibility information. Prove-

nance can explain data origins; fairness and sustainability mechanisms can reveal constraints; documentation can expose assumptions; and RAG evaluation can reveal unsupported generation. Yet these artifacts often remain external to plan selection, retrieval steering, ranking objectives, action protocols, escalation policies, and refusal decisions. RAIDS asks when responsibility information becomes execution state that invalidates plans, requires repair, changes rankings, triggers review, qualifies outputs, escalates decisions, or stops automation. III. T HE R ESPONSIBLE DATA I NTELLIGENCE L OOP From pipeline to loop. Conventional data pipelines assume that execution ends when an output is produced, so responsibility is assessed at the point of output. Modern data systems increasingly operate inside decision cycles: outputs become inputs to action, and consequences return through feedback to influence future data, policies, and decisions. When support is insufficient, constraints are violated, or actionability is uncertain, execution should not simply continue; systems may repair evidence, replan, request review, escalate, or refuse unsafe action before re-entering the loop. Responsibility therefore does not terminate at output; the needed abstraction is a loop, not a one-way pipeline. RAIDS views data systems through the Responsible Data Intelligence Loop in Fig. 1, where responsibility must be preserved across processing, understanding, accountable action, and feedback. Responsibility as loop state. Responsibility is not a layer outside the loop; it is loop state: what supports the output, which constraints apply, and which action should follow. If this state is insufficient, systems should repair it, qualify the result, escalate, refuse, or record audit conditions. Responsibility preservation means preserving this state when sufficient, repairing it when degraded, and using it to change course when insufficient. This is the visual claim in Fig. 1: RAIDS puts responsibility in the loop; the contract links loop stages, the solid flow routes understanding through accountable action, and the dashed curve marks feedback: monitor, contest, or update outcome state before re-entry. Loop phases. In processing, systems identify relevant data, preserved quality evidence, affected groups, applicable policies, and costs. In understanding, systems turn data into claims, uncertainty, explanations, and provenance. In accountable action, systems decide whether the result should be used, qualified, reviewed, refused, audited, or improved. Responsibility state couples these phases: weak processing can make claims unsupported, weak understanding can hide uncertainty, and weak action protocols can make an output irresponsible. Scope of the vision. RAIDS is neither post-training responsible AI nor a governance layer. Failures arise before model invocation; post-hoc governance cannot change execution. Nor is RAIDS reducible to provenance, fairness, sustainability, or explainability, each of which captures a fragment. The problem is responsibility coordination: how should infrastructure preserve, repair, and operationalize responsibility state as pipelines move from data to decision and feedback?

RAIDS: Responsible Data Intelligence Loop

Conventional Pipeline

1. Processing Data Responsibility Contract

Processing Outcome

(Support, Constraint, Actionability)

2. Understanding

Modeling & Analysis continue / repair / escalate / refuse

Output 3. Accountable Action Post-hoc Responsibility Feedback: monitor / contest / update

Fig. 1. Conventional pipelines attach responsibility after output. RAIDS makes the responsibility contract part of the loop: responsibility state (support, constraint, and actionability) travels through processing, understanding, accountable action, and feedback before returning to outcome.

The BlueSky opportunity. The opportunity is to make responsibility executable. Database research has turned broad goals into systems abstractions: consistency into transaction semantics, performance into optimization, reliability into recovery, and scale into distributed execution. RAIDS asks for an analogous move in responsible data intelligence. The Responsible Data Intelligence Loop defines the system boundary; the responsibility contract defines the execution interface through which responsibility participates in computation.

privacy mechanisms and regulation shape disclosure and data use [37], [38], [75]; and policy frameworks emphasize human fallback, notice, explanation, safety, and discrimination protection [74]. In RAIDS, these obligations become executionchanging parameters. Executable states. The contract has three validity states. • Support provides epistemic validity: whether the output is sufficiently grounded by evidence, freshness, uncertainty, provenance, conflicts, quality, and explanation. • Constraint provides normative validity: whether the output IV. T HE R ESPONSIBILITY C ONTRACT and execution path satisfy policy, fairness, privacy, safety, Making responsibility native requires an infrastructure primsustainability, resource, and impact limits. itive: a responsibility contract. This is the core abstraction: the • Actionability provides operational validity: what the next interface through which responsibility participates in execution protocol should be, such as answer, qualify, warn, ask, defer, and tests whether operators preserve it when they compose. A escalate, refuse, audit, or report. conventional operator contract specifies inputs, outputs, cost, Together, support, constraint, and actionability correspond to and accuracy; a RAIDS contract exposes epistemic support, epistemic, normative, and operational validity. Constraint state normative constraints, and allowable action modes. is executable because it determines which outputs, operators, DataOperator(input, task context, responsibility context) and continuations remain admissible. For example, fairness → (output, support, constraint, actionability). criteria [69], [71], [76] and privacy requirements [37], [38], Context. Here input denotes input data and output the pro- [77] are not labels attached after execution; they can change duced result. The task context describes task type, intent, how data are selected, repaired, aggregated, anonymized, scope, target output, domain, time window, and pipeline role. ranked, withheld, qualified, or refused before use. The responsibility context specifies obligations and thresholds The states guide continuation. If insufficient, RAIDS may refor judging the output; the same mining, retrieval, ranking, or fresh evidence, preserve conflict, replan, reduce footprint, ask, decision-support task may be exploratory in one context and escalate, refuse, invite contestation, or produce an audit trail. The goal is to expose enough state for responsible use, not to regulated in another. RAIDS uses four compact dimensions: make every component solve every social objective: delayed • Evidence and quality: evidence, metadata, uncertainty, proveharms can arise when actions influence future data [10], [78]. nance, freshness, robustness, and explanation [66]–[68]. Systems implications. The contract implies four system behav• Policy and accountability: regulation [38], [40], governance decision rights [35], accountable roles, privacy duties, report- iors: preserve or repair responsibility state across operators; ing duties, and audit requirements that parameterize execution. use that state in optimization; keep it inspectable for evidence, policy, impact, review, and risk; and let weak support, vio• Social and environmental impact: affected groups and distributional effects [60], [69]–[72], sustainability budgets and lated constraints, or uncertain actionability alter control flow through qualification, clarification, escalation, or refusal. footprint [39], and social-benefit obligations [73]. • Decision and oversight: qualification rules, human-in-theResponsibility as execution semantics. Here responsibilloop triggers, review, escalation, refusal, contestability, and ity becomes execution semantics rather than decoration. A feedback [32], [74]. responsibility contract changes the set of valid pipelines, These dimensions keep responsibility contextual without turn- not just the explanation shown after execution. It can make ing RAIDS into a compliance framework. Data-management an optimizer prefer stronger evidence, refresh stale sources, work links transparency, fairness, and data protection [34]; preserve conflict, or require human review before action. In

traditional analytics, the system boundary ends at a query result, dashboard, or prediction; in responsible data intelligence, it extends to the conditions under which the output becomes actionable. A result that cannot be explained, audited, or safely used is incomplete, even if syntactically valid and statistically accurate; RAIDS treats refusal, qualification, escalation, contestation, and uncertainty exposure as system behaviors. V. A B LUE S KY R ESEARCH AGENDA The target is responsibility preservation: after composition, the system should still determine whether an output is supported, constraint-compliant, and actionable. If state degrades, execution should repair it or change course. Responsibility state remains visible, compositional, optimizable, and actionable. Responsibility state management. RAIDS needs a state model for support, constraint, and actionability. The challenge is to represent heterogeneous responsibility information in a form lightweight enough for execution yet structured enough to be propagated, composed, degraded, and repaired. The state should carry evidence, uncertainty, provenance, quality, policy obligations, fairness constraints, sustainability limits, review requirements, and unresolved risks. For example, in evidence-intensive pipelines such as RAG [79]–[81], retrieved passages, provenance, freshness, coverage, conflicts, claimevidence links, and verification results should become support state, not prompt-local artifacts. The question is how to define algebraic and operational rules across relational, graph, textual, model-based, and human-facing operators. Responsibility-preserving execution. Execution needs invariants, monitors, and repair protocols that preserve responsibility state while operators compose. It should preserve sufficient support, satisfied constraints, and actionability before an output informs action. For example, a pipeline may need to detect schema or data drift, validate or repair records, preserve conflicting evidence, or stop when constraints fail; production data/AI pipelines expose hidden dependencies and observability needs [48], [49], [82], while validation and repair systems show how support can be checked and improved during execution [41], [62], [63]. RAIDS turns these mechanisms from controls into pipeline-level invariants: if state degrades, execution should repair, qualify, replan, or stop. Responsibility-aware optimization. Optimization needs cost and admissibility models that include responsibility context. Query optimizers trade off latency, I/O, and memory; RAIDS asks when evidence coverage, provenance, privacy, fairness, compliance, carbon impact, social impact, or review burden makes a plan inadmissible or less preferred. For example, private query release may alter answers to limit disclosure [37], fair query processing and exposure-aware ranking may change result selection [60], [83], [84], graph ranking and sampling may trade utility against robustness, fairness, and throughput [85], [86], and sustainable query processing may make carbon impact part of the objective [39]. Goals aligned with United Nations Sustainable Development Goals (SDGs) [73] should become constraints and audit criteria in

socially beneficial analytics [87]. The optimization problem is admissible plan selection under responsibility context. Claim-level provenance and explanation. Responsible outputs need claim/evidence graphs with decision traces [80], [88], [89]. Reports and decision outputs are collections of claims, not atomic results; each claim should carry evidence, transformations or models, policies, approvals, and unresolved uncertainty. For example, a compliance summary or scientific report should expose which records, transformations, models, and assumptions support each claim. Provenance foundations and deployment tracing show how derivations and execution traces can be queried [44], [45], [54]; RAIDS extends this machinery from data items and jobs to claims, decisions, and action conditions. Explanation should tell the system and user what can be trusted, what remains unresolved, and what action is permitted [68]. Decision and oversight protocols. Actionability requires runtime policies for deciding what the system should do next. Human oversight must become a runtime protocol, not a passive interface after automation: support, constraint, and responsibility context should govern when to answer, qualify, ask, escalate, refuse, audit, or support contestation. For example, a public-service or health analytics system may answer when support is strong, qualify incomplete evidence, request review for high stakes, and refuse automation when safety or discrimination risks remain. Rights frameworks emphasize notice, explanation, human fallback, safety, and discrimination protection [74], while risk and AI governance frameworks connect such obligations to lifecycle management [32], [40], [90]. RAIDS asks how to translate obligations into runtime modes with clear triggers, handoffs, and audits. Evaluation and success criteria. Holistic evaluation shows why a single score is insufficient [18], [27]; evaluation should test responsibility preservation after composition, not component quality alone. Strong components can lose support, violate constraints, or choose the wrong action mode when composed. In RAIDS, correct behavior is not always to answer: a system may be correct because it qualifies, escalates, refuses, or requests review. Success is a class of guarantees, metrics, benchmarks, and execution protocols: support, constraints, and actionability should remain meaningful after composition; policy, privacy, fairness, safety, resource, and social or environmental limits should change plans or outputs when needed; and the system should explain what it produced, why it is supported, which limits apply, and what should happen next. Progress is visible when systems preserve responsibility state across heterogeneous operators and choose the continuation mode under evidence, constraints, and oversight requirements. VI. C ONCLUSION RAIDS makes responsibility part of execution semantics: evidence, uncertainty, provenance, policy, sustainability, social impact, and oversight become state that guides computation and action. Future data systems should not only compute results; they should preserve the conditions under which results can responsibly become action.

R EFERENCES [1] Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan, “Dissecting racial bias in an algorithm used to manage the health of populations,” Science, vol. 366, no. 6464, pp. 447–453, 2019. [2] K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, P. Payne, M. Seneviratne, P. Gamble, C. Kelly, N. Schärli, A. Chowdhery, P. Mansfield, D. Demner-Fushman, B. Agüera y Arcas, D. Webster, G. S. Corrado, Y. Matias, K. Chou, J. Gottweis, N. Tomasev, Y. Liu, A. Rajkomar, J. Barral, C. Semturs, A. Karthikesalingam, and V. Natarajan, “Large language models encode clinical knowledge,” Nature, vol. 620, pp. 172– 180, 2023. [3] Y. Yan, H. Li, H. He, G. Kai, Z. Yang, and G. Liu, “SALP-CG: Standardaligned LLM pipeline for classifying and grading large volumes of online conversational health data,” arXiv preprint arXiv:2601.09717, 2026. [4] S. Gu, B. Kelly, and D. Xiu, “Empirical asset pricing via machine learning,” Journal of Financial Economics, vol. 136, no. 1, pp. 222– 253, 2020. [5] W. Wang, J. Yu, Z. Yang, M. Ju, S. Yu, J. Wu, L. Liu, Y. Liu, J. Shepherd, and W. Zhang, “AEFA: An ensemble framework for fraud detection in the forex market,” in Advanced Data Mining and Applications. Springer Nature Singapore, 2025, pp. 34–49. [6] X. Shu, M. Ju, Z. Chen, Y. Ding, W. Zhang, D. Wen, and Z. Yang, “ForexAgent: Identifying trading strategies in forex markets with large language models,” in 2026 IEEE International Conference on Big Data and Smart Computing, ser. BigComp, 2026, pp. 55–62. [7] D. F. Engstrom, D. E. Ho, C. M. Sharkey, and M.-F. Cuéllar, “Government by algorithm: Artificial intelligence in federal administrative agencies,” Administrative Conference of the United States, Tech. Rep., 2020. [8] J. A. Kroll, J. Huey, S. Barocas, E. W. Felten, J. R. Reidenberg, D. G. Robinson, and H. Yu, “Accountable algorithms,” University of Pennsylvania Law Review, vol. 165, no. 3, pp. 633–705, 2017. [9] A. Chouldechova, “Fair prediction with disparate impact: A study of bias in recidivism prediction instruments,” Big Data, vol. 5, no. 2, pp. 153–163, 2017. [10] D. Ensign, S. A. Friedler, S. Neville, C. Scheidegger, and S. Venkatasubramanian, “Runaway feedback loops in predictive policing,” in Proceedings of the 1st Conference on Fairness, Accountability and Transparency, ser. FAT* ’18, 2018, pp. 160–171. [11] D. Yan, J. Liu, B. Han, Z. Yang, J. He, J. Xu, R. Y. Sunindijo, and C. C. Wang, “Improving access to building licensing information in Australia: Design and development of a graph-based retrieval-augmented generation artificial intelligence system,” Buildings, vol. 16, no. 6, p. 1224, 2026. [12] T. Hey, S. Tansley, and K. Tolle, Eds., The Fourth Paradigm: DataIntensive Scientific Discovery. Microsoft Research, 2009. [13] J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žı́dek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. RomeraParedes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, and D. Hassabis, “Highly accurate protein structure prediction with AlphaFold,” Nature, vol. 596, pp. 583–589, 2021. [14] A. Merchant, S. Batzner, S. S. Schoenholz, M. Aykol, G. Cheon, and E. D. Cubuk, “Scaling deep learning for materials discovery,” Nature, vol. 624, pp. 80–85, 2023. [15] K. Bi, L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian, “Accurate medium-range global weather forecasting with 3D neural networks,” Nature, vol. 619, pp. 533–538, 2023. [16] M. Hino, E. Benami, and N. Brooks, “Machine learning for environmental monitoring,” Nature Sustainability, vol. 1, pp. 583–588, 2018. [17] L. Chen, B. Han, X. Wang, J. Zhao, W. Yang, and Z. Yang, “Machine learning methods in weather and climate applications: A survey,” Applied Sciences, vol. 13, no. 21, p. 12019, 2023. [18] P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson,

Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda, “Holistic evaluation of language models,” Transactions on Machine Learning Research, 2023. [19] X. Tang, L. Chen, W. Yang, Z. Yang, M. Ju, X. Shu, Z. Yang, and Y. Tang, “Tabular-textual question answering: From parallel program generation to large language models,” World Wide Web, vol. 28, no. 4, p. 42, 2025. [20] W. Yang, Z. Yang, L. Chen, R. Yan, Z. Yang, L. Zhang, and Y. Tang, “Parallel program generation for hybrid tabular-textual question answering,” in Web and Big Data. Springer Nature Singapore, 2024, pp. 121–137. [21] H. V. Jagadish, J. Gehrke, A. Labrinidis, Y. Papakonstantinou, J. M. Patel, R. Ramakrishnan, and C. Shahabi, “Big data and its technical challenges,” Communications of the ACM, vol. 57, no. 7, pp. 86–94, 2014. [22] A. Ailamaki, S. Madden, D. Abadi, G. Alonso, S. Amer-Yahia, M. Balazinska, P. A. Bernstein, P. Boncz, M. Cafarella, S. Chaudhuri, S. Davidson, D. DeWitt, Y. Diao, X. L. Dong, M. Franklin, J. Freire, J. Gehrke, A. Halevy, J. M. Hellerstein, M. D. Hill, S. Idreos, Y. Ioannidis, C. Koch, D. Kossmann, T. Kraska, A. Kumar, G. Li, V. Markl, R. Miller, C. Mohan, T. Neumann, B. C. Ooi, F. Ozcan, A. Parameswaran, I. Pandis, J. M. Patel, A. Pavlo, D. Porobic, V. Sanca, M. Stonebraker, J. Stoyanovich, D. Suciu, W.-C. Tan, S. Venkataraman, M. Zaharia, and S. B. Zdonik, “The cambridge report on database research,” arXiv preprint arXiv:2504.11259, 2025. [23] J. J. Pan, J. Wang, and G. Li, “Survey of vector database management systems,” The VLDB Journal, vol. 33, no. 5, pp. 1591–1615, 2024. [24] M. I. Jordan and T. M. Mitchell, “Machine learning: Trends, perspectives, and prospects,” Science, vol. 349, no. 6245, pp. 255–260, 2015. [25] M. Mazumder, C. Banbury, X. Yao, B. Karlaš, W. G. Rojas, S. Diamos, G. Diamos, L. He, A. Parrish, H. R. Kirk, J. Quaye, C. Rastogi, D. Kiela, D. Jurado, D. Kanter, R. Mosquera, J. Ciro, L. Aroyo, B. Acun, L. Chen, M. S. Raje, M. Bartolo, S. Eyuboglu, A. Ghorbani, E. Goodman, O. Inel, T. Kane, C. R. Kirkpatrick, T.-S. Kuo, J. Mueller, T. Thrush, J. Vanschoren, M. Warren, A. Williams, S. Yeung, N. Ardalani, P. Paritosh, L. Bat-Leah, C. Zhang, J. Zou, C.-J. Wu, C. Coleman, A. Ng, P. Mattson, and V. J. Reddi, “DataPerf: Benchmarks for data-centric AI development,” in Advances in Neural Information Processing Systems, 2023, datasets and Benchmarks Track. [26] C. Cortes, C. Sanz, L. Etcheverry, and A. Marotta, “Data quality management for responsible AI in data lakes,” in VLDB Workshops, 2024. [27] K. Ning, Z. Pan, Y. Jiang, A. Schneider, Y. Nevmyvaka, and D. Song, “Towards interpretable and trustworthy time series reasoning: A BlueSky vision,” in 2025 IEEE International Conference on Data Mining Workshops (ICDMW), 2025, pp. 2497–2502. [28] D. Gao, H. Wang, Y. Li, X. Sun, Y. Qian, B. Ding, and J. Zhou, “Text-toSQL empowered by large language models: A benchmark evaluation,” Proceedings of the VLDB Endowment, vol. 17, no. 5, pp. 1132–1145, 2024. [29] L. Patel, S. Jha, M. Pan, H. Gupta, P. Asawa, C. Guestrin, and M. Zaharia, “Semantic operators and their optimization: Enabling LLMbased data processing with accuracy guarantees in LOTUS,” Proceedings of the VLDB Endowment, vol. 18, no. 11, pp. 4171–4184, 2025. [30] S. Shankar, T. Chambers, T. Shah, A. G. Parameswaran, and E. Wu, “DocETL: Agentic query rewriting and evaluation for complex document processing,” arXiv preprint arXiv:2410.12189, 2024. [31] L. Lai, C. Luo, Y. Lou, M. Ju, and Z. Yang, “Graphy’our data: Towards end-to-end modeling, exploring and generating report from raw data,” in Companion of the 2025 International Conference on Management of Data, ser. SIGMOD Companion ’25, 2025, pp. 147–150. [32] National Institute of Standards and Technology, “Artificial intelligence risk management framework (AI RMF 1.0),” National Institute of Standards and Technology, Tech. Rep. NIST AI 100-1, 2023. [33] J. Stoyanovich, S. Abiteboul, B. Howe, H. V. Jagadish, and S. Schelter, “Responsible data management,” Communications of the ACM, vol. 65, no. 6, pp. 64–74, 2022. [34] S. Abiteboul and J. Stoyanovich, “Transparency, fairness, data protection, neutrality: Data management challenges in the face of new regulation,” arXiv preprint arXiv:1903.03683, 2019. [35] V. Khatri and C. V. Brown, “Designing data governance,” Communications of the ACM, vol. 53, no. 1, pp. 148–152, 2010.

[36] I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes, “Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing,” in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, ser. FAT* ’20, 2020, pp. 33–44. [37] C. Dwork, F. McSherry, K. Nissim, and A. Smith, “Calibrating noise to sensitivity in private data analysis,” in Theory of Cryptography, ser. TCC ’06, 2006, pp. 265–284. [38] European Parliament and Council of the European Union, “Regulation (EU) 2016/679 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data (General Data Protection Regulation),” Official Journal of the European Union, OJ L 119, 4 May 2016, 2016. [39] M. Bachras and H.-A. Jacobsen, “Environmental footprints of query processing: A vision for sustainable database architectures,” Proceedings of the VLDB Endowment, vol. 18, no. 11, pp. 4064–4072, 2025. [40] European Parliament and Council of the European Union, “Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act),” Official Journal of the European Union, OJ L, 2024/1689, 12 July 2024, 2024. [41] E. Breck, N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich, “Data validation for machine learning,” in Proceedings of the 2nd Conference on Machine Learning and Systems, ser. MLSys, 2019, pp. 334–347. [42] M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, “Model cards for model reporting,” in Proceedings of the Conference on Fairness, Accountability, and Transparency, ser. FAT* ’19, 2019, pp. 220–229. [43] T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daumé III, and K. Crawford, “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021. [44] T. J. Green, G. Karvounarakis, and V. Tannen, “Provenance semirings,” in Proceedings of the Twenty-Sixth ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems, ser. PODS ’07, 2007, pp. 31–40. [45] F. Psallidas, A. Agrawal, C. Sugunan, K. Ibrahim, K. Karanasos, J. Camacho-Rodrı́guez, A. Floratou, C. Curino, and R. Ramakrishnan, “OneProvenance: Efficient extraction of dynamic coarse-grained provenance from database query event logs,” Proceedings of the VLDB Endowment, vol. 16, no. 12, pp. 3662–3675, 2023. [46] C. Sandvig, K. Hamilton, K. Karahalios, and C. Langbort, “Auditing algorithms: Research methods for detecting discrimination on internet platforms,” in Data and Discrimination: Converting Critical Concerns into Productive Inquiry, 2014. [47] I. D. Raji and J. Buolamwini, “Actionable auditing: Investigating the impact of publicly naming biased performance results of commercial AI products,” in Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, ser. AIES ’19, 2019, pp. 429–435. [48] N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich, “Data management challenges in production machine learning,” SIGMOD Record, vol. 46, no. 2, pp. 17–20, 2017. [49] D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison, “Hidden technical debt in machine learning systems,” in Advances in Neural Information Processing Systems, vol. 28, 2015, pp. 2503–2511. [50] A. Pavlo, M. Butrovich, L. Ma, P. Menon, W. S. Lim, D. Van Aken, and W. Zhang, “Make your database system dream of electric sheep: Towards self-driving operation,” Proceedings of the VLDB Endowment, vol. 14, no. 12, pp. 3211–3221, 2021. [51] K. Weber, B. Otto, and H. Österle, “One size does not fit all—a contingency approach to data governance,” ACM Journal of Data and Information Quality, vol. 1, no. 1, 2009. [52] R. Abraham, J. Schneider, and J. vom Brocke, “Data governance: A conceptual framework, structured review, and research agenda,” International Journal of Information Management, vol. 49, pp. 424–438, 2019. [53] M. Janssen, P. Brous, E. Estevez, L. S. Barbosa, and T. Janowski, “Data governance: Organizing data for trustworthy artificial intelligence,” Government Information Quarterly, vol. 37, no. 3, p. 101493, 2020. [54] A. Chapman, P. Missier, G. Simonelli, and R. Torlone, “Capturing and querying fine-grained provenance of preprocessing pipelines in data science,” Proceedings of the VLDB Endowment, vol. 14, no. 4, pp. 507– 520, 2021. [55] A. Datta, S. Sen, and Y. Zick, “Algorithmic transparency via quantitative input influence: Theory and experiments with learning systems,” in 2016 IEEE Symposium on Security and Privacy, 2016, pp. 598–617.

[56] M. Akhtar, O. Benjelloun, C. Conforti, L. Foschini, J. Giner-Miguelez, P. Gijsbers, S. Goswami, N. Jain, M. Karamousadakis, M. Kuchnik, S. Krishna, S. Lesage, Q. Lhoest, P. Marcenac, M. Maskey, P. Mattson, L. Oala, H. Oderinwale, P. Ruyssen, T. Santos, R. Shinde, E. Simperl, A. Suresh, G. Thomas, S. Tykhonov, J. Vanschoren, S. Varma, J. van der Velde, S. Vogler, C.-J. Wu, and L. Zhang, “Croissant: A metadata format for ML-ready datasets,” in Advances in Neural Information Processing Systems, 2024, datasets and Benchmarks Track. [57] N. Jain, M. Akhtar, J. Giner-Miguelez, R. Shinde, J. Vanschoren, S. Vogler, S. Goswami, Y. Rao, T. Santos, L. Oala, M. Karamousadakis, M. Maskey, P. Marcenac, C. Conforti, M. Kuchnik, L. Aroyo, O. Benjelloun, and E. Simperl, “A standardized machine-readable dataset documentation format for responsible AI,” arXiv preprint arXiv:2407.16883, 2024. [58] B. Salimi, L. Rodriguez, B. Howe, and D. Suciu, “Interventional fairness: Causal database repair for algorithmic fairness,” in Proceedings of the 2019 International Conference on Management of Data, ser. SIGMOD ’19, 2019, pp. 793–810. [59] F. Nargesian, A. Asudeh, and H. V. Jagadish, “Tailoring data source distributions for fairness-aware data integration,” Proceedings of the VLDB Endowment, vol. 14, no. 11, pp. 2519–2532, 2021. [60] A. Singh and T. Joachims, “Fairness of exposure in rankings,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ser. KDD ’18, 2018, pp. 2219– 2228. [61] A. J. Biega, K. P. Gummadi, and G. Weikum, “Equity of attention: Amortizing individual fairness in rankings,” in Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, ser. SIGIR ’18, 2018, pp. 405–414. [62] T. Rekatsinas, X. Chu, I. F. Ilyas, and C. Ré, “HoloClean: Holistic data repairs with probabilistic inference,” Proceedings of the VLDB Endowment, vol. 10, no. 11, pp. 1190–1201, 2017. [63] S. Krishnan, J. Wang, E. Wu, M. J. Franklin, and K. Goldberg, “ActiveClean: Interactive data cleaning while learning convex loss models,” in Proceedings of the 2016 International Conference on Management of Data, ser. SIGMOD ’16, 2016, pp. 948–959. [64] C. G. Northcutt, L. Jiang, and I. L. Chuang, “Confident learning: Estimating uncertainty in dataset labels,” Journal of Artificial Intelligence Research, vol. 70, pp. 1373–1411, 2021. [65] X. Yang, K. Sun, H. Xin, Y. Sun, N. Bhalla, X. Chen, S. Choudhary, R. D. Gui, Z. W. Jiang, Z. Jiang, L. Kong, B. Moran, J. Wang, Y. E. Xu, A. Yan, C. Yang, E. Yuan, H. Zha, N. Tang, L. Chen, N. Scheffer, Y. Liu, N. Shah, R. Wanga, A. Kumar, W.-t. Yih, and X. L. Dong, “CRAG: Comprehensive RAG benchmark,” in Advances in Neural Information Processing Systems, 2024, datasets and Benchmarks Track. [66] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, J. Bouwman, A. J. Brookes, T. Clark, M. Crosas, I. Dillo, O. Dumon, S. Edmunds, C. T. Evelo, R. Finkers, A. Gonzalez-Beltran, A. J. G. Gray, P. Groth, C. Goble, J. S. Grethe, J. Heringa, P. A. C. ’t Hoen, R. Hooft, T. Kuhn, R. Kok, J. Kok, S. J. Lusher, M. E. Martone, A. Mons, A. L. Packer, B. Persson, P. Rocca-Serra, M. Roos, R. van Schaik, S.-A. Sansone, E. Schultes, T. Sengstag, T. Slater, G. Strawn, M. A. Swertz, M. Thompson, J. van der Lei, E. van Mulligen, J. Velterop, A. Waagmeester, P. Wittenburg, K. Wolstencroft, J. Zhao, and B. Mons, “The FAIR guiding principles for scientific data management and stewardship,” Scientific Data, vol. 3, p. 160018, 2016. [67] M. Li, X. Zhang, H. Ying, Y. Li, X. Han, and D. Yu, “Data quality aware hierarchical federated reinforcement learning framework for dynamic treatment regimes,” in 2023 IEEE International Conference on Data Mining (ICDM), 2023, pp. 1103–1108. [68] R. Doddaiah, P. S. Parvatharaju, E. A. Rundensteiner, and T. Hartvigsen, “Class-specific explainability for deep time series classifiers,” in 2022 IEEE International Conference on Data Mining (ICDM), 2022, pp. 101– 110. [69] M. Hardt, E. Price, and N. Srebro, “Equality of opportunity in supervised learning,” in Advances in Neural Information Processing Systems, vol. 29, 2016, pp. 3315–3323. [70] Q. Hu and H. Rangwala, “Metric-free individual fairness with cooperative contextual bandits,” in 2020 IEEE International Conference on Data Mining (ICDM), 2020, pp. 182–191. [71] W. Zhang and J. C. Weiss, “Fair decision-making under uncertainty,” in 2021 IEEE International Conference on Data Mining (ICDM), 2021, pp. 886–895.

[72] X. Tang, Y. Ding, Z. Yang, Y. Chen, Y. Gu, W. Yang, M. Ju, X. Cao, Y. Liu, and W. Zhang, “Do they understand them? an updated evaluation on nonbinary pronoun handling in large language models,” in AI 2025: Advances in Artificial Intelligence. Springer Nature Singapore, 2025, pp. 204–219. [73] United Nations, “Transforming our world: The 2030 agenda for sustainable development,” United Nations, 2015. [74] White House Office of Science and Technology Policy, “Blueprint for an AI bill of rights: Making automated systems work for the american people,” The White House, 2022. [75] O. Feyisetan, T. Diethe, and T. Drake, “Leveraging hierarchical representations for preserving privacy and utility in text,” in 2019 IEEE International Conference on Data Mining (ICDM), 2019, pp. 210–219. [76] C. Dwork, M. Hardt, T. Pitassi, O. Reingold, and R. Zemel, “Fairness through awareness,” in Proceedings of the 3rd Innovations in Theoretical Computer Science Conference, ser. ITCS ’12, 2012, pp. 214–226. [77] D. Chen, Q. Zhang, L. M. Kaplan, A. Jøsang, D. H. Jeong, F. Chen, and J.-H. Cho, “fair-LDP: Uncertainty-guided fairness and privacy for federated healthcare learning,” in 2025 IEEE International Conference on Data Mining (ICDM), 2025, pp. 130–139. [78] L. T. Liu, S. Dean, E. Rolf, M. Simchowitz, and M. Hardt, “Delayed impact of fair machine learning,” in Proceedings of the 35th International Conference on Machine Learning, ser. ICML ’18, 2018, pp. 3150–3158. [79] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 9459–9474. [80] D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, and J. Larson, “From local to global: A graph RAG approach to queryfocused summarization,” arXiv preprint arXiv:2404.16130, 2024. [81] J. Liu, Z. Chen, S. Qiao, M. Ju, D. Zhang, B. Han, S. Yu, X. Shu, J. Wu, D. Wen, X. Cao, G. Liu, and Z. Yang, “A2RAG: Adaptive agentic

graph retrieval for cost-aware and reliable reasoning,” arXiv preprint arXiv:2601.21162, 2026. [82] S. Shankar and A. G. Parameswaran, “Towards observability for production machine learning pipelines,” Proceedings of the VLDB Endowment, vol. 15, no. 13, pp. 4015–4022, 2022. [83] S. Shetiya, I. P. Swift, A. Asudeh, and G. Das, “Fairness-aware range queries for selecting unbiased data,” in 2022 IEEE 38th International Conference on Data Engineering, ser. ICDE, 2022, pp. 1423–1436. [84] M. Zehlike, K. Yang, and J. Stoyanovich, “Fairness in ranking, part I: Score-based ranking,” ACM Computing Surveys, vol. 55, no. 6, pp. 1–36, 2022. [85] Y. Li, Y. Xu, X. Lin, W. Zhang, and Y. Zhang, “Ranking on dynamic graphs: An effective and robust band-pass disentangled approach,” in Proceedings of the ACM on Web Conference 2025, ser. WWW ’25, 2025, pp. 3918–3929. [86] F. Masrour, F. Santos, P.-N. Tan, and A.-H. Esfahanian, “Fairnessaware graph sampling for network analysis,” in 2022 IEEE International Conference on Data Mining (ICDM), 2022, pp. 1107–1112. [87] Y. Ding, X. Tang, Z. Yang, W. Zhang, S. Wu, Y. Huang, L. Lan, W. Li, Y. Chen, M. Ju, W. Yang, T. Hoang, M. Klymenko, X. Xu, and W. Zhang, “EulerESG: Automating ESG disclosure analysis with LLMs,” arXiv preprint arXiv:2511.21712, 2025. [88] J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia, “ARES: An automated evaluation framework for retrieval-augmented generation systems,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics, 2024. [89] R. Friel, M. Belyi, and A. Sanyal, “RAGBench: Explainable benchmark for retrieval-augmented generation systems,” arXiv preprint arXiv:2407.11005, 2024. [90] I. D. Raji, P. Xu, C. Honigsberg, and D. E. Ho, “Outsider oversight: Designing a third party audit ecosystem for AI governance,” arXiv preprint arXiv:2206.04737, 2022.

Related documents

Record · ID 300040 · SHA-256 e68138e51627c1da
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.