ConceptioArchivearXiv CS
arXiv CSopen access

Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems Prajjwal Gupta, Prasang Gupta, Vishal Bhutani, Apoorva Sharma, Sumanth Chundru, Waqar Sarguroh, Kevin Paul PricewaterhouseCoopers, U.S. Correspondence: [email protected]

Abstract I shipped an AI product, how do I track its performance ?

arXiv:2606.23403v1 [cs.AI] 22 Jun 2026

As agentic LLM systems move from prototypes to deployment across increasingly diverse domains, evaluating them has become both more important and more difficult. The challenge is not only that individual metrics may be unreliable, but that evaluation goals are often left implicit. Without a clear account of what a system is expected to do, how it can fail, and which failures matter, metric choices become difficult to justify, interpret, or validate. We present Litmus, a zero-label system that designs evaluation and monitoring metrics for AI pipelines by eliciting evaluation intent from source code and targeted interrogation. Instead of assuming that the evaluation target is already known, Litmus first identifies what must be measured and why, then converts those answers into constraints for constructing a justified, per-stage metric portfolio. We evaluate Litmus on three real, code-defined AI pipelines—financial account grouping, scientific QA, and inherent risk assessment—against AutoMetrics and three DynamicRubric baselines. Litmus achieves the broadest or tied-broadest concern coverage, spans more pipeline stages, produces a near-zero-redundancy portfolio, and ranks first in validity against per-row quality labels on all three pipelines—decisively on scientific QA (Spearman ρ = 0.72 vs. less than 0.47 for every baseline), and within overlapping confidence intervals in relation to two components of the audit framework despite using no labels during metric design. Our results support a shift from automatic metric implementation to automatic metric specification: before asking which metric to compute, evaluation systems should ask what must be measured and why.

1

Clarification question

Codebase

Confirmed facts

Adversarial validity question

Should the fallback path fire often ?

Does mean latency hide tail spikes that hurt ? constraints

Metrics

Rarely, only on parse errors

Justified metric portfolio Retrieval: context coverage AI core: p95 latency

Figure 1: Litmus designs metrics by asking questions first. From source code and a practitioner’s goal alone, it interrogates the pipeline through practitioner-facing clarification questions and internal adversarial validity questions. The answers become facts that act as constraints, yielding a justified, per-stage metric portfolio.

engineering, healthcare administration, and enterprise knowledge management. Many of these systems are no longer single prompt-response models: they are agentic or pipeline-based applications that combine retrieval, tool use, routing logic, structured outputs, business rules, fallbacks, confidence estimates, and human review. This shift changes the role of evaluation. In research settings, evaluation is often used to compare model outputs; in production settings, it must also support debugging, monitoring, auditability, and risk management. This distinction matters because production failures are rarely explained by a single final-output score. A financial-document pipeline may fail because retrieval returned the wrong context, because a rule-based override was skipped, because a fallback path fired too often, because confidence was miscalibrated, or because a downstream formatter masked an upstream error. Conversely, a behavior

Introduction

LLM-based systems are increasingly moving from demonstrations to production workflows in domains such as finance, customer support, software 1

that appears undesirable in isolation may be correct under the system design: frequent fallbacks may indicate model weakness in one deployment but appropriate caution in another. For industry practitioners, the practical question is therefore not only whether an output is good, but which component produced the behavior, whether that behavior was expected, and which operational risk should be monitored. Current evaluation practice is poorly matched to this need. Automatic metrics have become the default instrument for evaluating natural language generation (NLG) systems: a recent survey of 110 ACL and INLG papers finds that 94% report at least one automatic metric (Schmidtová et al., 2024). Yet the same survey finds that 76.9% of reported metric usages provide no rationale, and that unclear evaluation goals contribute to a “kitchen-sink” style of reporting many weakly motivated scores (Schmidtová et al., 2024; Zhou et al., 2022). This is downstream of a more basic issue: a metric cannot be justified, interpreted, operationalized, or monitored against an evaluation goal that has never been made explicit. For LLM pipelines, the practical process of translating source-code structure and practitioner intent into concrete evaluation and monitoring metrics remains under-specified. We argue that metric design should begin with an elicitation step that makes the evaluation goal explicit. The question that precedes which metric? is: what is this component supposed to do, how can it fail silently, and what is therefore worth measuring? This is fundamentally an inquiry problem—the same uncertainty-resolution view that underlies clarification-question generation (Rao and Daumé III, 2018). We adopt this view for AI-system evaluation. Rather than treating metrics as generic objects to be selected after the fact, we treat them as measurement commitments whose validity depends on unresolved facts about system design, deployment intent, and failure semantics. We instantiate this view in Litmus (Figure 1), a zero-label system that designs evaluation and monitoring metrics for an AI pipeline directly from its source code. Litmus first builds a code-grounded model of the pipeline: its major components, their roles, the AI capabilities they use, and the failure surfaces they expose. It then maps these components to an evaluation-pattern taxonomy to identify plausible metric families. However, code can expose where measurement may be needed without fully specifying what should count as success, fail-

ure, or acceptable risk. Litmus therefore interrogates the pipeline to convert such ambiguities into metric-design constraints. Practitioner-facing clarification questions target under-specified decisions that affect metric design, such as which processing tier carries production traffic or whether a fallback path is expected to fire often or rarely. Internal adversarial validity questions challenge each candidate metric on pipeline fit, data assumptions, measurement validity, and direction of goodness. The resulting answers become confirmed facts: constraints that admit, reject, or re-scope metrics. Litmus therefore outputs not a single holistic score, but a justified portfolio of stage-specific evaluation and monitoring metrics and is currently deployed and being used to support governance of AI tooling across multiple client projects. Our contributions are threefold: 1. We reframe automatic metric design as goal elicitation by interrogation, connecting metric validity concerns in NLG evaluation to clarification-question generation and inquirybased uncertainty resolution. 2. We introduce Litmus (Figure 2), a zero-label system that derives a justified, per-stage evaluation and monitoring portfolio from source code, with practitioner answers and internal validity checks represented as explicit metricdesign constraints. 3. We evaluate Litmus on three real, codedefined AI pipelines spanning distinct domains—financial account grouping, scientific QA, and inherent risk assessment—against AutoMetrics and DynamicRubric baselines. To assess metric-design quality beyond label agreement, we additionally introduce portfolio-level axes—coverage, grounding, and redundancy—that characterize what an automatic metric-design system produces.

2

Related Work

Evaluation in deployed AI systems serves a broader role than benchmark comparison: it must support debugging, regression testing, monitoring, auditability, and operational decision-making. Prior work on production machine learning emphasizes that deployed ML systems accumulate hidden technical debt through data dependencies, configuration assumptions, feedback loops, and boundary failures (Sculley et al., 2015). Breck et al. (2017) 2

argue for systematic tests that go beyond aggregate model quality, including data validation, infrastructure checks, and monitoring; Amershi et al. (2019) similarly show that engineering AI-enabled systems requires connecting model behavior to data, code, and deployment context. Documentation and auditing frameworks make a related point: intended use, limitations, assumptions, and evaluation evidence should be explicit in deployed AI systems (Mitchell et al., 2019; Gebru et al., 2021; Raji et al., 2020). Litmus follows this production-oriented view: it treats metric design as a system-level task in which metrics should be grounded in components, failure surfaces, and monitoring needs rather than only in final outputs. A growing body of work scrutinizes how metrics are used and reported. Schmidtová et al. (2024) survey current NLG evaluation practice and find pervasive missing rationales, missing implementation details, and missing correlations with human judgement; Zhou et al. (2022) trace the “kitchensink” tendency to unclear evaluation goals. Validity critiques of overlap metrics (Reiter, 2018; Papineni et al., 2002; Lin, 2004) motivate the move toward LLM-as-judge evaluation (Liu et al., 2023; Fu et al., 2023). However, LLM-as-judge methods do not remove the need for specification: a judge prompt still encodes assumptions about what matters, what evidence should be considered, and how scores should be interpreted. Prior work also shows that LLM judges may exhibit systematic biases and model-family alignment effects (Zheng et al., 2023). Litmus targets the specification gap these works identify: it produces metrics whose rationale is a first-class artifact. Recent benchmarks measure general agent capability across reasoning-and-acting, tool use, web navigation, and software-engineering tasks (Yao et al., 2023; Qin et al., 2024; Liu et al., 2024; Zhou et al., 2024; Jimenez et al., 2024). Deployed industry pipelines, however, contain domain-specific rules, retrieval layers, routing logic, fallbacks, confidence thresholds, and compliance constraints. Litmus addresses a complementary problem: designing metrics for a particular code-defined pipeline by identifying what should be measured, at which stage, and why. The closest conceptual neighbour to our framing is clarification-question generation. Rao and Daumé III (2018) formalize “a good question is one whose expected answer is most useful” via the expected value of perfect information, Rao and

Daumé III (2019) generate rather than merely rank such questions, and Majumder et al. (2021) identify information “essential to accomplish an underlying goal but currently missing from the context.” Litmus adopts this usefulness-of-answer stance: every question it surfaces is annotated with the specific metric decision its answer would resolve. In our setting, the missing information is not needed to answer a user query, but to determine whether a candidate metric is appropriate, what it should measure, and how it should be interpreted. We distinguish this from mainstream question generation (QG/NQG), which generates answerable comprehension questions from a passage (Mulla and Gharpure, 2023; Guo et al., 2024; Flor, 2025)— a different speech act from resolving evaluationintent ambiguity. Recent systems automate parts of evaluation construction. AutoMetrics takes outputs plus labels and fits a single runnable holistic judge – a metric implementation system (Ryan et al., 2025). DynamicRubric (Wang and Blanco, 2026) has the judge generate its own rubric and then score, in per-instance (-I NST) and dataset-wide (-DS) variants, with a DPO-fine-tuned generator variant (F INE T UNED). Both assume the evaluation target is given and focus on the final output. Litmus instead elicits the target first and designs a per-stage portfolio from source code with no labels. This distinction is especially important for production AI pipelines, where failures may arise in intermediate components and where monitoring requires metrics that are traceable to system behavior.

3

The Litmus System

Because code structure alone cannot determine intended behavior, acceptable failure modes, the available evidence for measurement, or even the direction in which a metric should improve, Litmus treats metric design as a constrained synthesis problem, organized into the stages below (Figure 2). 3.1

Architecture Reconstruction

Litmus begins by constructing an evidencegrounded representation of the repository. A deterministic scanner parses the codebase, builds import and symbol/call graphs, computes centrality scores, and selects salient files for LLM analysis. An LLM then synthesizes an 8–20 node component graph describing the system’s major modules, their roles, AI capabilities, and associated risk surfaces, and an 3

architecture + classification ↦ condition design

1

Architecture Reconstruction

2

Pattern Classification

3

Interrogation

CORE

4

Metric Design

5

Traceability & Export

Channel A — Clarification Qs

Source code repository

Static scanner

Classify components · agent

import graph · centrality · salient files tree-sitter · symbol / call graph

→ evaluation-pattern taxonomy

Pipeline stages

LLM synthesis · agent

→ self · 4 axes fit · validity · assumptions · direction may reject a contradicting metric

the only input

Cross-cutting strategies

Adversarial critic · agent

RAG · agentic → reference metric families

prune phantom · recover missed

Deterministic step

LLM agent step

Constraint gate

Map evaluation → monitoring

admit / reject / re-scope

evaluation metrics → monitoring counterparts

Channel B — Adversarial validity Qs

acquisition · knowledge base AI core · assurance

8–20 node component graph roles · capabilities · risk surfaces

→ practitioner each tagged with an explicit metric-impact statement

Metric kinds ≤ a few high-signal per component deterministic checks LLM-as-judge (5-pt CoT rubric) · composite

Export runnable artifacts per-stage, justified metric portfolio

Confirmed facts = hard constraints

Practitioner input

Self / adversarial check

Confirmed facts (constraints)

Output artifact

Figure 2: The Litmus pipeline. Static analysis and LLM synthesis reconstruct a validated component architecture (Phases 1–2); the system then interrogates the artifact and the practitioner (Phase 3), and the resulting confirmed facts become hard constraints that admit, reject, or re-scope each candidate metric (Phase 4) before export (Phase 5).

adversarial critic checks it against code evidence, pruning phantom components and flagging missed ones, to reduce over-interpretation. The result is a repository-level architecture that is both structural (rooted in dependency and call relationships) and semantic (components are labeled with their functional roles); full prompts are given in Appendix D.1, with architecture-reconstruction detail in Appendix C. 3.2

firmed facts, the designer emits a small number of high-signal metrics per component: deterministic checks, LLM-as-judge metrics (scored on a fivelevel rubric with an evidence-backed rationale), and composite metrics; each specifying its evaluated component, purpose, required data, computation, direction of goodness, and validity conditions. The result is a justified specification rather than a list of names: every metric is linked to the evidence and confirmed facts that support it, so the portfolio can be audited. Litmus exports runnable definitions with this traceability and, where possible, maps each evaluation metric to a production monitoring counterpart, carrying the same rationale from offline design to deployment.

From Patterns to Constraints

Litmus next maps each component to an evaluation taxonomy of pipeline-stage patterns (data acquisition, knowledge base, AI core, assurance) and cross-cutting strategies (e.g., retrieval-augmented generation, agentic orchestration). This does not select final metrics; it narrows the design space by flagging which metric families are plausibly relevant per component, e.g., a retrieval component activates coverage, grounding, and latency families, while leaving the underlying assumptions unresolved. Litmus settles those assumptions through interrogation: rather than silently assuming answers, it converts each metric-affecting uncertainty into an explicit question whose answer is recorded as a confirmed fact: a user- or critic-validated constraint that downstream metric design must satisfy. Confirmed facts may determine whether a metric is included, how it is computed and thresholded, what data must be instrumented, how multiple signals are aggregated, or which direction indicates improvement. 3.3

4

Experimental Setup

4.1

Evaluation Domains

Because Litmus is a source-code-based metricdesign system while the baselines construct finaloutput judges or rubrics, we interpret the results as trade-offs under a zero-label design-time constraint rather than universal superiority claims. We evaluated three real, code-defined AI pipelines from different domains, testing whether the benefits of eliciting evaluation targets generalize across production-style systems with different tasks, data, and failure modes rather than being specific to one artifact. Account grouping This audit task classifies individual client general-ledger accounts into standardized, aggregated financial-statement categories. For each account the pipeline emits an account_group, a confidence score, and a naturallanguage reason, via multiple decision paths (rule-

Metric Design and Export

Only after interrogation does Litmus finalize the portfolio. Conditioned on the reconstructed architecture, the component classifications, and the con4

System

Code

Labels

Scope

construction. We adopt two measures from the measurement-validity framework of Ryan et al. (2025): validity, their criterion validity, measuring agreement between system scores and independent per-row quality labels; and degradation sensitivity, their robustness sensitivity measure, testing whether a metric penalizes synthetically degraded outputs. We also report three portfolio-level design-quality axes—coverage, grounding, and redundancy—that characterize what a metric-design system emits. Because final-output baselines were not built to optimize these axes, we treat them as descriptive diagnostics rather than a like-for-like contest. Each axis is defined below; full scoring rubrics are in Appendix B.

Mon.

Litmus yes no per-stage yes AutoMetrics no yes final no DynRub-I NST no no instance no DynRub-DS no no dataset no DynRub-FT no feedback final no Table 1: Comparison protocol. Litmus is source-code grounded and produces a per-stage metric portfolio, while the baselines are final-output evaluation or rubric systems. Mon. = monitoring metrics; FT = fine-tuned.

based overrides, disaggregation logic, and retrievalaugmented mapping). Scientific QA We use the established PaperQA2/LitQA2 (Skarlinski et al., 2024) questionanswering setting and evaluate the answersynthesis stage on a key-passage corpus, isolating answer quality given relevant evidence.

Validity This is the criterion validity of Ryan et al. (2025): the association (Kendall’s τ , Spearman’s ρ) between a system’s score on uncorrupted outputs and an independent per-row quality label. Scope-matched judges are scored against the matching label; for aggregate reporting we compute a stitched per-row score (Section 5).

Inherent risk assessment (IRA) This is a core audit-planning task: auditors must flag high-risk areas needing scrutiny without over-auditing lowrisk ones. It is organized around inherent risk factors (IRFs), a standardized set of risk categories indicating where material misstatement or audit complexity may arise. For each (account group, FSLI, IRF) tuple, the pipeline assigns a risk level and produces an evidence-grounded justification from retrieved client documents. 4.2

Coverage It is the fraction of a fixed, domainspecific list of reference failure concerns addressed by at least one metric (ten for account grouping, six each for scientific QA and inherent risk assessment), capturing whether a portfolio spans the failure surface rather than emitting one generic score. Grounding This measures how concretely a metric refers to implementation artifacts, named modules, fields, rules, branches, or source-code behavior, rather than generic evaluation language, on a [0, 1] scale (higher is stronger).

Systems Compared

We compare Litmus with four baseline configurations that automate final-output evaluation or rubric construction. Table 1 summarizes the role and inputs of each system. Litmus emits per-stage portfolios (14 metrics for account grouping, 13 for scientific QA, 14 for inherent risk assessment) that include monitoring metrics, scope-matched LLM-as-judge metrics, and code-evaluated checks. The baselines (Section 2) instead score the final output and, unlike Litmus, produce no per-stage monitoring metrics; AutoMetrics is additionally trained on a separate client/train split, so its validity is out-of-distribution. 4.3

Redundancy This is the fraction of metric pairs judged to measure substantially the same behavior (lower is better), penalizing portfolios that look broad but repeat one property. It is meaningful only for multi-metric portfolios; single-judge baselines (the DynamicRubric variants) are trivially non-redundant, which we mark explicitly. Degradation sensitivity This is the sensitivity measure of AutoMetric’s construct-validity robustness check: whether a metric assigns lower scores to synthetically degraded outputs (we corrupt 10 rows over five severity levels per domain).

Evaluation Axes

We score every system under a single protocol applied uniformly across all three domains, computing each axis identically for all systems, including ours, so no axis advantages a design by

5

Results

Because every Litmus metric is conditioned on the recovered component graph, we check that 5

Multi-axis (single-pass, temp. 0) System

Single-vector validity (B3)

Cov. ↑ Ground. ↑ Redund. ↓ Valid. τ ↑ Valid. ρ ↑ Deg. ind. ↑

P1: Account grouping (financial), n=112 AutoMetrics 0.30 0.41 0.14 DynRub-I NST 0.40 0.15 0.00 DynRub-DS 0.40 0.35 0.00 DynRub-FT 0.40 0.35 0.00 Litmus (ours) 0.80 0.45 0.02

0.30 0.40 0.25 0.41 0.49

0.34 0.45 0.27 0.47 0.51

0.54 0.71 0.52 0.68 0.77

P2: Scientific QA (LitQA2), n=58 AutoMetrics 0.67 0.19 DynRub-I NST 0.83 0.15 DynRub-DS 0.67 0.35 DynRub-FT 1.00 0.35 Litmus (ours) 1.00 0.32

0.14 0.00 0.00 0.00 0.01

0.68 −0.06 −0.12 0.27 0.71

0.69 −0.06 −0.12 0.27 0.72

0.92 0.95 0.95 0.95 0.94

P3: Inherent risk assessment (IRF), n=112 AutoMetrics 0.83 0.16 0.07 DynRub-I NST 0.83 0.15 0.00 DynRub-DS 0.00 0.35 0.00 DynRub-FT 0.83 0.35 0.00 Litmus (ours) 0.83 0.42 0.05

0.25 0.09 0.25 0.21 0.27

0.31 0.11 0.28 0.25 0.32

0.71 0.65 0.49 0.82 0.84

ρ [95% CI]

τ -b

p

0.45 [0.29, 0.59] 0.45 [0.29, 0.58] 0.27 [0.08, 0.42] 0.47 [0.32, 0.60] 0.53 [0.37, 0.67]

0.35 0.40 0.25 0.41 0.49

< 0.001 < 0.001 0.005 < 0.001 < 0.001

0.47 [0.24, 0.63] 0.42 < 0.001 −0.06 [−0.12, −0.04] −0.06 1.00 −0.12 [−0.23, −0.09] −0.12 1.00 0.27 [−0.07, 0.71] 0.27 0.040 0.72 [0.43, 0.93] 0.71 < 0.001 0.36 [0.18, 0.51] 0.11 [−0.08, 0.28] 0.28 [0.10, 0.45] 0.25 [0.05, 0.42] 0.39 [0.22, 0.52]

0.27 0.09 0.25 0.21 0.28

< 0.001 0.245 0.003 0.008 < 0.001

Table 2: Multi-axis comparison and single-vector validity (B3) across three pipelines. Bold marks the best value per panel; arrows give the preferred direction. Cov. = coverage, Ground. = grounding, Redund. = redundancy, Valid. τ /ρ = validity vs. the per-row quality label, Deg. ind. = degradation sensitivity (single-pass, temperature 0; the first three are LLM-assessed, see Section 7). B3 is the stitched single-vector score vs. the same label, with bootstrap 95% CIs and permutation p-values; its ρ is a single coefficient and differs from the per-metric Valid. ρ.

graph rather than assume it: for the scientific-QA pipeline (PaperQA2), Litmus recovers a 30-node graph that is 65% edge-verified against the treesitter call graph, maps every recovered module to real code (coverage F1 1.00, zero phantoms), and earns an adversarial-critic confidence of 0.78 (Table 4, Appendix C)—so its metrics rest on a graph grounded in resolvable code rather than assumed.

mus’s judges are also among the most sensitive to corrupted outputs—top or tied on degradation on P1 and P3, and on par with the rubric systems on P2 (0.94 vs. 0.95)—so it is the only portfolio that is strong on both validity and degradation sensitivity across all three pipelines.

On this footing, the pattern across the three pipelines is consistent: Litmus is the broadest and least redundant portfolio with the strongest label validity. It is top or tied on coverage across all pipelines and uniquely produces per-stage and operational metrics rather than scoring the final output alone; its redundancy is near zero throughout, well below AutoMetrics; and on grounding it leads on P1 and P3, trailing only slightly on P2 (Table 2). Despite using zero labels at design time, Litmus attains the strongest label agreement on all three pipelines (ρ = 0.51, 0.72, 0.32), ahead of every baseline including the label-fitted AutoMetrics; the stitched single-vector statistic preserves this ranking, with Litmus first and individually significant throughout (p < 0.001), though on P1 and P3 its interval overlaps the strongest baseline, so we claim a favorable ranking under the zero-label constraint rather than a uniformly significant margin. Lit-

We reframed automatic metric design as goal elicitation by interrogation: rather than selecting metrics first, Litmus interrogates a pipeline’s source code and its practitioner to make the evaluation target explicit, then admits, rejects, or re-scopes each candidate against the resolved constraints. Across three domains—account grouping, scientific QA, and inherent risk assessment—this zero-label procedure consistently yields broader coverage, nearzero redundancy, and the strongest label validity, ahead of label-fitted and rubric-generating baselines; that it ranks first on validity everywhere while baselines swing across domains points to designtime elicitation, not artifact-specific tuning. Key next steps are multi-judge-family replication to isolate the method from assessor bias, a human study of whether practitioner answers measurably change the resulting metrics, and broader pipelines to test generalization.

6

6

Conclusion and Future Work

7

Limitations

tonomously. Two of our three evaluation pipelines (account grouping and inherent risk assessment) are real audit workflows operating over proprietary client financial data; we do not release that data, identifiers are sanitized, and only aggregate measurements are reported. Because these are highstakes audit and financial settings, automatically designed metrics should support—never replace— professional judgment and existing review controls. Designed metrics, particularly LLM-as-judge metrics, can encode the biases of the underlying model and should be reviewed and calibrated by practitioners before deployment. Finally, because metrics influence which system behaviours are surfaced and acted upon, over-trust in any automatically designed metric—including ours—risks masking failure modes it does not cover; the coverage axis is intended to make such gaps visible rather than to certify completeness.

First, scope. Our empirical evaluation spans three real, code-defined pipelines from different domains—account grouping, scientific QA, and inherent risk assessment—but all are scored by a single judge-model family. The quantitative results should therefore be read as evidence across three deployments within one assessor family; Litmus itself is pipeline-agnostic, with no assumptions specific to any one domain, and broader multi-judge evaluation remains future work. Second, the systems are not strictly commensurable: Litmus is a zero-label metric-design system, whereas AutoMetrics is label-fitted and DynamicRubric is rubricgenerating, and both score only the final output. AutoMetrics is fit on a separate client dataset and evaluated on the held-out validity rows, so its reported figure is already an out-of-distribution estimate rather than an optimistic in-distribution one. We treat this role difference as intentional rather than a confound, but it means absolute numbers should be read as characterizing what each system produces; the defensible claim is the ranking under a shared zero-label, design-time constraint. Third, three of our axes (coverage, grounding, and redundancy) are themselves LLM-assessed by the same model family that powers Litmus, risking judge–system alignment bias (Zheng et al., 2023; Panickssery et al., 2024); we mitigate this with a broad, fixed set of reference concerns per pipeline, but the absolute scores inherit the assessor’s biases and the ranking is again the more defensible claim. Fourth, coverage is monotone in portfolio size—Litmus emits a substantially larger portfolio than the baselines (14 metrics versus 1–8 on account grouping: AutoMetrics emits 8 and DynamicRubric a single holistic rubric score)—and the reference concern lists were authored by us, so the coverage gap should be read as an upper bound on the size-controlled advantage. Finally, the clarification-question channel is a design affordance whose end-to-end benefit (practitioner answers measurably changing the resulting metrics) is established here at the specification level and warrants a dedicated human study.

8

References Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. 2019. Software engineering for machine learning: A case study. In International Conference on Software Engineering: Software Engineering in Practice. Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. 2017. The ml test score: A rubric for ml production readiness and technical debt reduction. In IEEE International Conference on Big Data. Michael Flor. 2025. Question generation with large language models and generative ai. In Automatic Question Generation, Synthesis Lectures on Human Language Technologies. Springer, Cham. Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire. In Proceedings of EMNLP. Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Communications of the ACM. Shasha Guo, Lizi Liao, Cuiping Li, and Tat-Seng Chua. 2024. A survey on neural question generation: Methods, applications, and prospects. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI), pages 8038–8047.

Ethical Considerations

Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations.

Litmus analyzes source code and produces evaluation specifications; it does not make end-userfacing decisions, and the designed metrics are intended to inform practitioners rather than to act au7

Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81. Association for Computational Linguistics.

algorithmic auditing. In Proceedings of the Conference on Fairness, Accountability, and Transparency. Sudha Rao and Hal Daumé III. 2018. Learning to ask good questions: Ranking clarification questions using neural expected value of perfect information. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2737–2746, Melbourne, Australia. Association for Computational Linguistics.

Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, and 3 others. 2024. Agentbench: Evaluating llms as agents. In International Conference on Learning Representations.

Sudha Rao and Hal Daumé III. 2019. Answer-based adversarial training for generating clarification questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 143–155. Association for Computational Linguistics.

Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: Nlg evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522. Association for Computational Linguistics.

Ehud Reiter. 2018. A structured review of the validity of BLEU. Computational Linguistics, 44(3):393–401.

Bodhisattwa Prasad Majumder, Sudha Rao, Michel Galley, and Julian McAuley. 2021. Ask what’s missing and what’s useful: Improving clarification question generation using global knowledge. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4300–4312. Association for Computational Linguistics.

Michael J. Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu, Di Xu, and Diyi Yang. 2025. AutoMetrics: Approximate human judgements with automatically generated evaluators. Preprint, arXiv:2512.17267. Patrícia Schmidtová, Saad Mahamood, Simone Balloccu, Ondřej Dušek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondřej Plátek, and Adarsa Sivaprasad. 2024. Automatic metrics in natural language generation: A survey of current evaluation practices. Preprint, arXiv:2408.09169.

Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency.

D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. 2015. Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems.

Nikahat Mulla and Prachi Gharpure. 2023. Automatic question generation: a review of methodologies, datasets, evaluation metrics, and applications. Progress in Artificial Intelligence, 12(1):1–32.

Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques, and Andrew D. White. 2024. Language agents achieve superhuman synthesis of scientific knowledge. Preprint, arXiv:2409.13740.

Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024. LLM evaluators recognize and favor their own generations. Preprint, arXiv:2404.13076. Kishore Papineni, Salim Roukos, Todd Ward, and WeiJing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318.

Zijie Wang and Eduardo Blanco. 2026. Generating and refining dynamic evaluation rubrics for LLM-as-ajudge. Preprint, arXiv:2605.30568. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.

Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024. Toolllm: Facilitating large language models to master 16000+ real-world apis. In International Conference on Learning Representations.

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-Bench and chatbot arena. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023) Datasets and Benchmarks Track.

Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the ai accountability gap: Defining an end-to-end framework for internal

8

ours—so that no axis advantages a particular design by construction. The axes serve two distinct purposes. Validity is our label-grounded comparison: it tests whether a system’s scores agree with an external per-row quality signal, the standard notion of whether an automatically produced metric or judge is trustworthy. Coverage, grounding, and redundancy are portfolio-level design-quality axes that we introduce to characterize what an automatic metric-design system emits; they are not objectives the final-output baselines were built to optimize, so we report them as descriptive, diagnostic context rather than as a like-for-like contest. Degradation sensitivity is a complementary robustness check that we report alongside validity: it confirms a system reacts to corrupted outputs, which validity alone does not establish.

Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, and Alexandra Olteanu. 2022. Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 314–324. Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations.

A

Reference Failure-Concerns

The coverage axis for the account-grouping pipeline uses the following fixed list of ten failureconcerns: (1) account assigned to a semantically wrong group; (2) group inconsistent with / contradicting the stated FSLI; (3) business process context ignored; (4) high confidence on an incorrect assignment (miscalibration); (5) missing or empty reasoning; (6) grouping not grounded in retrieved evidence (hallucination); (7) wrong granularity of disaggregation; (8) override rules for account types not respected; (9) operational failure (exceptions/retries/degraded service); (10) retrieval quality issues (low recall / irrelevant chunks).

B

C.1.1

Validity measures the association between a system’s score on uncorrupted outputs and an independent per-row quality label, reported as Kendall’s τ and Spearman’s ρ. It is the only axis grounded in a target external to every system and is therefore our primary head-to-head comparison. The label is used only here: no system except AutoMetrics (which is label-fitted) consumes it, and Litmus uses none at design time. For branch- or stagescoped metrics, each judge is evaluated against the scope-matched label; for aggregate reporting we also compute a stitched per-row score, as described in Section 5.

Scoring and Implementation Details

All LLM-based scoring uses temperature 0 with up to three retries. The same judge-model family is used for metric scoring, LLM-derived proxy labels for account grouping and inherent risk assessment, and LLM-assessed design-level axes such as coverage, grounding, and redundancy. This keeps the scoring setup consistent across systems, but introduces possible judge-family alignment bias, which we discuss in Section 7. The coverage reference lists contain ten concerns for account grouping and six each for scientific QA and inherent risk assessment (the accountgrouping concerns are enumerated in Appendix A). For degradation sensitivity, each domain corrupts 10 rows across five cumulative severity levels.

C

Architecture Generation: Pipeline and Protocol

C.1

Evaluation Axes

Validity

C.1.2

Coverage

Coverage measures the fraction of a fixed, domainspecific list of reference failure concerns addressed by at least one metric: ten reference concerns for account grouping, and six each for scientific QA and inherent risk assessment. It asks whether a portfolio spans the relevant failure surface rather than producing a single generic score. Because the final-output baselines emit one judge or rubric, coverage characterizes a property of portfolio breadth they were not designed to provide, and we report it descriptively. C.1.3

Grounding

Grounding measures how concretely a metric refers to implementation artifacts—named modules, fields, rules, branches, or source-code behavior—rather than using generic evaluation language, scored on a [0, 1] scale with higher values

We score every system under a single protocol applied uniformly across all three domains, computing each axis identically for all systems—including 9

indicating stronger grounding. It reflects Litmus’s source-code-grounded design goal; baselines that score only the final output do not have access to such artifacts, so this axis describes artifact specificity rather than a deficiency the baselines could have avoided. C.1.4

# Component

Kind

Silent-fail

1 Trial-Balance Ingest entry compute low 2 FSLI Exact Matcher supporting compute medium 3 Account-Type Override Mapper ai_core prompt high 4 Prioritized-RAG Mapper ai_core compute high 5 Disaggregated-FSLI Mapper ai_core compute high 6 Azure Cognitive Search supporting vector_store medium 7 Orchestrator + Post-proc orchestrator compute high 8 Output Assembler output compute medium

Redundancy

Table 3: Components Litmus recovers for the accountgrouping pipeline. Roles, kinds, and silent-failure risk are emitted fields of the synthesizer schema.

Redundancy measures the fraction of metric pairs judged to measure substantially the same behavior, with lower being better. It penalizes portfolios that appear broad but repeatedly measure the same property. The axis is meaningful only for multimetric portfolios: single-judge baselines (the DynamicRubric variants) are trivially non-redundant, which we mark explicitly so the axis is not misread as an advantage. C.1.5

Role

allel mapping branches (account-type override, prioritized-RAG, disaggregated-FSLI), a shared Azure Cognitive Search client, an orchestrator that routes accounts and post-processes the result (resolving FSLI contradictions and tidying account_group), and an output assembler. The three branches it finds are the same ones the validity analysis is scoped to, even though that scoping came from the routing labels and not from the graph.

Degradation sensitivity

Degradation sensitivity measures how much a metric’s score drops when outputs are synthetically corrupted at increasing severity. A trustworthy metric must react to degraded outputs; we report degradation sensitivity alongside validity because the two are complementary—sensitivity confirms that a metric responds when outputs are broken, while validity confirms that it tracks genuine quality on real outputs.

How many edges are real. Each proposed edge cites caller→callee symbol pairs that a deterministic pass checks against the tree-sitter call graph: verified when the evidence resolves, inferred when none is cited or found, unsupported when the cited symbols are absent (and then discarded unless needed for connectivity). Here, of 49 edges, 32 are verified, 1 inferred, and 16 unsupported (0 discarded, all retained for connectivity)—so 65% of the claimed data flow points to a function call we can actually locate.

How Litmus reconstructs an architecture. Figure 3 shows the generation pipeline. A deterministic front end scans the repository, extracts tree-sitter digests, scores files by callee-weighted centrality to pick salient ones, and builds a symbol/call graph. An LLM then synthesizes an 8–20 node component graph (clustering files into subsystems, inflating each into nodes, and stitching cross-subsystem edges), with every edge required to cite the calls that justify it. Two grounding steps follow: a deterministic pass marks each edge verified, inferred, or unsupported by checking its citations against the call graph (dropping unsupported ones), and an adversarial LLM critic re-reads the source to score per-node confidence, prune phantoms, and flag misses. Only after these checks is the graph used for metric design, which is what lets us report fidelity as a measurement rather than an assumption.

What the critic caught. The critic scores its confidence in each node and proposes missing ones; nodes below 0.3 flagged “not found” are removed automatically. On this pipeline it returned overall confidence 0.78 (lowest node confidence 0.65), pruned 0 phantom nodes—no component’s claimed files were absent—and flagged 3 missed, e.g. Types & Data Models. The evidence behind three components. Identifiers are sanitized; the data is proprietary (Section 8). Override Mapper (3) maps to the overridemapper module plus override_rules.json (classes ASSET, LIABILITY, EQUITY, REVENUE , EXPENSE , NON - OPERATING , OTHER - OPERATING , NON - CONTROLLING - INTEREST ); verified edges include orchestrator:route_accounts → override_mapper:classify_batch and → llm_client:complete, which is why its

The graph recovered for account grouping. Running this pipeline on the audit pipeline yields the eight components in Table 3: a trial-balance ingest point, an FSLI exact-matcher, three par10

Figure 3: Litmus’s architecture-generation pipeline. Deterministic stages (blue) bracket the LLM stages (orange): static analysis and a symbol/call graph feed an LLM synthesizer, whose output is grounded by deterministic edge verification and an adversarial LLM critic before a validated component graph is emitted. Dashed green arrows are the grounding signals—synthesized edges are checked against the call graph, and the critic re-reads the source.

judge is scoped to “Account type override” rows. RAG Mapper (4) maps to the RAG module with verified edges rag_mapper:map → azure_search_client:search and → llm_client:complete, licensing a groundedness judge. Azure Search (6), detected as vector_store, has three verified inbound edges (nodes 2, 4, 5), so one dependency-health monitor covers all three callers.

stage?” FLAG: it manages prompt templates and configuration rather than generating text, so the {LLM} classification is questionable. • Pipeline fit. “Do production-monitoring metrics (Error Rate per Exception Class, Retry Success Rate, Dead-Letter Queue Rate) apply to the test-suite component?” FLAG: test suites are not production services, so exception/retry rates indicate test failures, not production health.

Self-review (adversarial validity) questions. After drafting candidate metrics, Litmus interrogates each one along four axes —pipeline fit, validity conditions, data assumptions, and direction of goodness —and emits a verdict of PASS, FLAG, or REJECT. On the scientific-QA pipeline this selfreview passed 36, flagged 10, and rejected 11 of the candidate metrics. Representative resolved questions:

• Redundancy. “Is Exception Recovery Rate distinct from Graceful Degradation Success Rate?” FLAG: both measure whether caught exceptions still yield usable responses; the distinction is thin. Practitioner clarification questions. Litmus also drafts clarification questions for the practitioner, each emitted with an explicit metric-impact statement of what its answer changes. The following are representative questions for the scientificQA components (the direction-of-goodness and scoping decisions they resolve are noted in italics):

• Pipeline fit. “Does RAG Context Precision measure the component it is attached to?” RE JECT : the component is a PDF reader/parser that extracts text; it performs no query-based retrieval, so a retrieval-quality metric cannot apply.

• Evidence Retrieval & Summarization: “Is this stage tuned for broad recall or precise grounding when selecting evidence passages?” Determines whether the retrieval judge emphasizes recall-oriented or precision-oriented families and sets its direction of goodness.

• Data assumptions. “Does Tool Selection Accuracy have the ground truth it needs?” FLAG: the computation requires a golden-tool mapping; it is unclear whether such labels exist in this codebase. • Data assumptions. “Can Retrieval Recall@K be computed?” FLAG: Recall@K needs known-relevant documents per query, and no labeled evaluation set was found.

• Answer Generation & Formatting: “Should an answer that abstains (“insufficient evidence”) count as a success or a failure?” Sets the direction of goodness for the abstention/refusal-rate metric.

• Validity conditions. “Is the Settings & Configuration tier really an LLM-generation

• Agent Execution Loop: “Is a high tool-retry / re-search rate expected thoroughness or a 11

Edge verification verified

inferred

unsupported

Metadata Clients

Testing Core PaperQA Functionality Tests

Metadata PostProcessors (Journal Quality & Retractions)

Metadata Client & Clinical Trials Tests

Metadata Query Orchestrator

CrossRef API Client

Semantic Scholar API Client

Configuration & Settings Validation Tests

OpenAlex & Unpaywall API Clients

Test Fixtures & VCR Configuration

Agent & CLI Integration Tests

Agent Orchestration CLI Entry & Dispatch

Tool Definitions & Execution

Agent Execution Loop

Search Index Management

Core QA Engine Vector Store (Embedding Index)

Settings & Configuration

Contributed Sources OpenReview Paper Discovery & Download

Clinical Trials Search & Ingestion

Evidence Retrieval & Summarization

Answer Generation & Formatting

Document Ingestion & Chunking

Zotero Library Retrieval

Nemotron Reader Test Suite & Fixtures

Call Nemotron Model API

Merge & Postprocess BBox Results

Orchestrate PDF Page Parsing

PDF Reader Plugins Docling PDF Parser

PyMuPDF PDF Parser

Image Bbox Clustering

PyPDF/PDFPlumber PDF Parser

Figure 4: Architecture graph Litmus reconstructs for the scientific-QA pipeline (PaperQA2): 30 component nodes spanning seven subsystems—PDF reader plugins (Docling, PyMuPDF, PyPDF/PDFPlumber), agent orchestration (CLI dispatch, execution loop, tool definitions, search index), metadata clients (CrossRef, Semantic Scholar, OpenAlex/Unpaywall, journal-quality and retraction post-processors), the core QA engine (ingestion & chunking, evidence retrieval & summarization, answer generation, vector store, settings), the Nemotron reader (page-parse orchestration, Nemotron API call, bbox merge), contributed sources (Zotero, OpenReview, clinical trials), and the test suite. Nodes and data-flow edges are emitted by the synthesizer and grounded by deterministic edge verification and the adversarial critic before any metric is designed.

sign of degradation?” Sets the direction of goodness and threshold for the tool-retry-rate metric.

is unavailable, is falling back to another source acceptable or should it be flagged?” Decides whether the cross-source fallbackrate is monitored as healthy behavior or as a failure signal.

• Metadata Query Orchestrator: “When a metadata source (CrossRef, Semantic Scholar) 12

Fidelity measure Component nodes recovered Subsystems spanned Ground-truth source modules Data-flow edges verified Critic overall confidence Critic lowest-node confidence Phantom nodes pruned Missed components flagged Coverage P / R / F1 (all) Coverage P / R / F1 (AI-logic) Per-module F1 (strict)

Reproducing the numbers. The graph and critic annotations are persisted per solution, so these require no new model calls:

Value 30 7 32 65% 0.78 0.65 0 3 1.00 / 1.00 / 1.00 1.00 / 1.00 / 1.00 0.64

# dump latest graph + critic rows sqlite3 litmus.db \ "SELECT graph_json FROM architecture_graphs ORDER BY version DESC LIMIT 1;" > graph.json sqlite3 litmus.db \ "SELECT critic_annotations FROM architecture_graphs ORDER BY version DESC LIMIT 1;" > critic.json

Table 4: Architecture-reconstruction fidelity for the scientific-QA pipeline (PaperQA2). The seven subsystems are PDF reader plugins, agent orchestration, metadata clients, the core QA engine, the Nemotron reader, contributed sources, and tests. Strict per-module F1 (0.64) is lower than coverage F1 because Litmus clusters the 32 source modules into 30 functional nodes as shown in Figure 4 rather than emitting one node per file. All values are reproducible from the persisted graph and critic annotations (Appendix C).

# edge verification counts jq '[.edges[].verification] | group_by(.) | map({(.[0]): length}) | add' graph.json # critic confidence + missing nodes jq '.overallConfidence, [.annotations[] | select(.confidence < 0.3)], .missingNodes' critic.json

D

• PDF reader plugins: “Which parser (Docling, PyMuPDF, PyPDF) carries production traffic, and which are fallbacks?” Scopes the parse-failure-rate metric to the production parser rather than penalizing fallbacks.

Deterministic Algorithms and LLM Prompts by Component

This appendix documents each major Litmus component at the implementation level. For every phase we separate the two kinds of stages Litmus interleaves: deterministic stages (static analysis, graph construction, verification, rule-based selection), and LLM stages (synthesis, classification, critique, design) which are given as the verbatim system and user prompts issued to the model. Prompts are reproduced as sent, with three cosmetic normalizations for typesetting only: (i) runtime-interpolated values appear as {placeholder}; (ii) non-ASCII decorations in the original strings (box-drawing rules, bullets, arrows, ≤/≥) are mapped to ASCII; and (iii) embedded JSON schemas, which are generated programmatically from Zod definitions, are summarized rather than reprinted. Unless noted otherwise, every LLM call targets a Claude Opusclass model through an OpenAI-compatible gateway and requests structured JSON output; per-call temperature and token budget are stated with each prompt.

• Nemotron reader: “Is the Nemotron API the default page parser or an optional highaccuracy path?” Determines whether its availability and latency metrics are treated as production-critical. • Vector Store (Embedding Index): “What index staleness is acceptable before results are considered degraded?” Sets the threshold for the index-freshness monitor. How we built the ground truth. Without looking at Litmus’s output, we listed components from three sources we already have: the rule classes in override_rules.json; one component per distinct routing label (override, disaggregation, RAG, orchestrator, retrieval, operational, general); and the infrastructure dependencies named in the source—then one practitioner adjusted the list and confirmed boundaries. A node is a true positive if it matches a ground-truth component at the same granularity; spurious nodes are false positives, unmatched components false negatives. The AI-logic variant keeps only components with model calls or routing logic.

D.1

Architecture Reconstruction

A deterministic front end scans the repository, extracts tree-sitter structure, builds a symbol/call graph, and computes centrality. An LLM then synthesizes the component graph (single-pass, or the multi-pass cluster/inflate/stitch variant); deterministic passes verify every edge against the call graph, 13

assign node kinds and edge types, and repair connectivity; finally an adversarial LLM critic re-reads the source.

- operational: broader system health, reliability, observability, or governance surfaces

LLM prompts (architecture). The synthesizer, the multi-pass passes, and the critic share the analyst persona below. All synthesis calls use temperature 0.1; the critic uses temperature 0. Token budget is large (∼128k) to fit whole-repository context.

METRIC PRIORITY: - primary: components that directly shape semantic behavior or user-visible AI quality - secondary: components that materially support, constrain, or distort AI quality - tertiary: operational components whose metrics are useful but not first-order for AI evaluation - Do not use framework alignment to determine metric priority.

Shared

analyst

system

(buildArchitectureSystemPrompt);

prompt used

by

the

FRAMEWORK OVERLAY: - After reconstructing the real graph, annotate each node with whether it is strongly aligned to a framework stage, partially aligned, or adjacent. - A node is adjacent if it materially affects AI behavior, quality, latency, reliability , or observability, even if it is not a canonical framework component. - Do not force nodes into the framework based on naming alone.

synthesizer and the inflate pass. You are an AI systems analyst for evaluation engineering. Your task is to reconstruct the true execution architecture of this codebase from code evidence, then annotate that architecture so an AI engineer can understand: 1. how the system actually works, 2. which components are AI-core versus supporting or operational, 3. which components should receive metrics first.

EVALUATION OVERLAY: For each node, determine: - whether it is offline-evaluable - whether it is runtime-observable - whether quality can be measured directly, requires judge-style evaluation, or only has runtime proxies - what user-facing failure happens if the node is wrong

PRIORITY ORDER: 1. Execution truth 2. AI semantics 3. Evaluability and observability 4. Framework alignment EXECUTION TRUTH: - Infer the real runtime/dataflow from code, not from names. - Identify triggers, transformations, external calls, storage boundaries, branching logic, outputs, and feedback paths. - Prefer fewer truthful nodes over many shallow nodes. - Collapse helpers into parent nodes unless a helper has a distinct evaluation or observability surface.

NODE RULES: - Each node must map to real files or functions. - Each node must describe its primary runtime responsibility and downstream effect. - Keep descriptions concise and return compact JSON. - Prefer evaluability-relevant decomposition over generic software decomposition. - Preserve adjacent framework components if they still matter for evaluation.

AI SEMANTICS: Treat the following as first-class architectural behavior when present: - prompt_construction - model_call - retrieval - reranking - tool_use - memory - guardrail - judge - fallback - post_processing

EDGE RULES: - Edges should represent actual execution or control transitions. - Label edges when the transition matters for quality, failure, fallback, or observability. - Include feedback loops when retries, evaluations, or guardrails influence later behavior. DO NOT: - optimize for a pretty diagram over a truthful one - flatten AI behavior into generic "processing " - exclude adjacent components that still need evaluation - assume the framework is the system; it is only an overlay - mark a node as strongly aligned without code

SYSTEM ROLE: - ai_core: directly shapes semantic behavior, ranking, judgments, tool choice, or final answer quality - supporting: materially supports or constrains AI quality, freshness, latency, or traceability

14

evidence

files call across subsystem boundaries, then connect the architecture nodes that own those files. 5. CONNECTIVITY IS THE TOP PRIORITY: Every subsystem MUST have at least one edge connecting it to another subsystem. No subsystem may be isolated. 6. If symbol evidence is absent for a subsystem, infer edges from data flow patterns (e.g., orchestration calls processing, entry points feed pipelines, data subsystems serve compute subsystems). 7. Generate enough edges to make the graph navigable -- aim for at least one edge per subsystem pair that has a logical data flow relationship.

Cluster-pass system prompt (multi-pass synthesis, temperature 0.1). You are a codebase clustering engine. Group source files into 3-8 named subsystems based on their function call relationships and framework roles. RULES: 1. Every source file must belong to exactly one subsystem. 2. Minimize cross-subsystem call edges (high cohesion, low coupling). 3. Name subsystems by what they DO, not what they ARE (e.g. "Document Retrieval" not " Module A"). 4. Infrastructure nodes (databases, caches, queues, vector stores) form their own subsystem only if they have 3+ files; otherwise attach to their primary consumer.

Return ONLY valid JSON matching the schema. No markdown. {JSON schema}

Architecture critic system prompt (temperature 0).

5. Entry points (API routes, Lambda handlers, CLI) group together unless they serve clearly different domains. 6. Keep DISTINCT business/domain modules separate. If two groups of files implement different domain pipelines -- different domain vocabulary, inputs/outputs, or endto-end flow -- they MUST be separate subsystems. Never collapse them into one bucket. 7. Do NOT emit catch-all subsystems with vague names like "Domain Logic", "Core Logic", "Business Logic", "Processing", or " Miscellaneous" that lump unrelated domains together. Split such a group into its constituent domains, each named for what it does. 8. Choose the number of subsystems (3-8) that reflects the codebase's actual distinct domains and infrastructure -- do not overmerge to hit a smaller count.

You are an architecture critic. Your job is to review an architecture graph produced by another AI agent and verify it against the actual source code. CHECK FOR: 1. PHANTOM NODES: Nodes that claim source files that don't exist 2. MISSED NODES: High-importance files not represented by any node 3. WRONG EDGES: Claimed data flow that doesn't match import structure 4. MISCLASSIFIED ROLES: Nodes labeled as " ai_core" that don't use AI libraries 5. DUPLICATE NODES: Two nodes representing the same component For each node, provide: - confidence (0-1): how confident you are this node is accurate - issues: list of problems found (empty if none) - suggestions: list of improvements

Return ONLY valid JSON matching the schema. No markdown. {JSON schema}

Be adversarial. Assume the synthesizer made mistakes. Verify claims against the code.

Stitch-pass system prompt (connects nodes across subsystems).

Return ONLY valid JSON matching the schema. { JSON schema}

You are a cross-subsystem edge generator. Connect architecture nodes across different subsystems to ensure the graph is fully connected.

Architecture critic user prompt.

RULES: 1. Only create edges between nodes in DIFFERENT subsystems. 2. For every edge, populate evidenceCalls with caller->callee symbol pairs when symbol evidence is available. If no symbol evidence exists, use descriptive evidence like "data_flow: subsystemA.output -> subsystemB.input". 3. Do NOT duplicate edges that already exist within subsystems. 4. Use the CROSS-SUBSYSTEM SYMBOL EDGES section (if present) to identify which

Review this architecture graph for accuracy: GRAPH: {graph as JSON} SOURCE CODE EXCERPTS: {for each node, up to 2 source files (<=4000 chars each), or FILE NOT FOUND} FILE MANIFEST (for detecting missed nodes): {path [centrality=score] per file}

15

D.2

- Components that compare/match evidence against criteria - Matching or reconciliation logic between two data sources - Verification of assertions against supporting documents

Pattern-Based Metric Search Space

Classification narrows the metric search space by mapping each component to an evaluationpattern taxonomy: six pipeline stages (data acquisition, data transformation, knowledge base, AI core, assurance, continuous improvement) and five cross-cutting strategies (RAG, evidence matching, prompt-chaining, agentic, LLM-as-judge). A holistic LLM classifier sees the whole graph and emits per-component stage, cross-cutting strategies, and a single dominant pipeline pattern; an adversarial LLM validator checks the result; deterministic helpers provide a rule-based fallback.

LLM-as-Judge (judge): - Component evaluates, scores, ranks, or compares ANOTHER LLM's or AI model's OUTPUT (not domain objects like risks, accounts, or documents) - Rubric-based scoring with explicit criteria and a numeric scale applied to model-generated text or structured output - Pairwise preference / A-vs-B comparison of two model responses - Calibration / agreement testing against human-labeled data ON MODEL OUTPUTS - Outputs structured verdicts (pass/fail, score, reasoning) about another component's AI-generated output

LLM prompts (classification). The classifier runs at temperature 0.1; the validator at temperature 0. Both receive the taxonomy catalog (stage definitions and detection heuristics) inlined as JSON in place of {catalog}.

NOT judge if: - The component scores, classifies, or assesses DOMAIN OBJECTS (risks, accounts, documents, transactions) even if it uses an LLM to do so - The component is a domain classifier, ranker, or assessor that happens to produce scores -- that is "prompt_chain" or "rag", not "judge"

Holistic classifier system prompt. You are a holistic architecture classifier for AI evaluation engineering. Your task is to analyze the FULL architecture graph and: 1. Classify each component into a pattern stage from the framework catalog. 2. Detect cross-cutting strategies that span multiple components.

PRIMARY PIPELINE PATTERN (pipelinePattern): For EACH component, you MUST also emit a single dominant pipelinePattern label from this fixed set: rag | agentic | judge | prompt_chain | evidence_match | none

Unlike a per-component classifier, you can see the entire graph topology and detect graph-level patterns.

Decision rules (apply in order -- pick the FIRST that matches): 1. If the component is itself an LLM-as-Judge / scorer / grader THAT EVALUATES ANOTHER AI MODEL'S OUTPUT ( not domain objects) -> "judge" 2. Else if the component is part of an agent loop (tool-calling, ReAct, planning) -> "agentic" 3. Else if the component retrieves external context and feeds it to an LLM (or is the retriever in a clear retriever-> generator pipeline) -> "rag" 4. Else if the component compares/reconciles evidence between two sources -> "evidence_match" 5. Else if the component is one step in a sequential multi-LLM chain -> "prompt_chain" 6. Else (deterministic ETL, schema validation, infra glue, plain inference with no retrieval, etc.) -> "none"

{catalog: PATTERN STAGES and CROSS-CUTTING STRATEGIES as JSON} CROSS-CUTTING DETECTION RULES: RAG (Retrieval-Augmented Generation): - Retriever component feeding into a generator component with shared context - Vector store query followed by LLM call - Context injection into prompt from retrieved documents Prompt Chaining: - Sequential LLM calls where output of one feeds into input of next - Multi-step pipeline with intermediate LLM processing - Chain or workflow orchestration across multiple prompts Agentic: - Tool-use loops and decision nodes in the graph - ReAct loops or iterative decision-making - Autonomous multi-step execution with branching

A component's pipelinePattern is INDEPENDENT of its patternStage -- e.g. a RAG retriever has pipelinePattern="rag" AND patternStage="knowledge_base". Do NOT return "none" just because the component is non-LLM; only pick "none" when no

Evidence Matching:

16

LLM-pipeline pattern applies.

Adversarial metric critic system prompt.

CLASSIFICATION RULES: - Set patternStage to null for components that do not map to any framework stage. - pipelinePattern is REQUIRED for every component (use "none" if no pattern fits). - confidence should reflect how certain you are (0.0 to 1.0). - crossCuttingStrategies should list all strategies the component participates in ( multi-select; can include "judge"). - reasoning should be a concise explanation of your classification decision. - detectedPatterns should list all graph-level patterns you detect with the involved component IDs.

You are an adversarial metric critic for AI evaluation engineering. You review ALL metrics across ALL components at once to enforce quality, consistency, and deduplication.

Return ONLY valid JSON matching this schema. No markdown, no explanation. {JSON schema}

1. PIPELINE FIT -- does this metric make logical sense for where this component sits? - What does this component actually consume from upstream, and what does it emit downstream? - Can this component genuinely influence what the metric measures, or does the real signal live in a different component entirely? ( e.g. a retriever being scored on generation quality) - Is the measurement captured at a point where the data it needs is already present ? - Does the metric generalise to ANY component of the same type, or does it bite into what THIS component uniquely does? Generic-fit metrics are a REJECT, not a FLAG.

Your job is to issue a verdict for EVERY metric: PASS, REJECT, or FLAG -- with a structured, multi-axis review that tells the metric designer exactly how to fix it. ======================= FOUR-AXIS REVIEW -- run this mentally for every metric before picking a verdict. =======================

Classification validator user prompt (source digests loaded only for items with confidence < 0.8). Validate the following classifications against the architecture graph. CLASSIFICATIONS: {classifications as JSON} DETECTED PATTERNS: {detectedPatterns as JSON} ARCHITECTURE GRAPH: {graph as JSON} SOURCE CODE DIGESTS (for low-confidence items)

D.3

2. VALIDITY CONDITIONS -- under what specific conditions does this metric produce signal ? - Enumerate the runtime conditions that must hold for the number to be meaningful (e.g. "non-empty retrieval result", " English query", "user session has prior turn", "source doc contains the entity being cited"). - Identify cases where the metric will return a value but MEAN NOTHING (cache hits, fallback paths, empty inputs falling through, early exits). If those failure modes dominate real traffic, this is a REJECT. - Are the valid conditions realistic given how the component is exercised in production, or does the metric only fire on a narrow sliver of requests?

Interrogation: Turning Assumptions into Constraints

Interrogation runs two channels. Clarification questions to the practitioner are generated deterministically from AST features and detected code patterns; each is emitted with a metric-impact statement and capped at three per component. Validity questions are posed by an adversarial LLM critic that issues a per-metric verdict (pass / reject / flag) after a structured four-axis review (pipeline fit, validity conditions, data & signal assumptions, direction of goodness). Confirmed answers become hard constraints: a deterministic policy step rejects metrics whose threshold contradicts a confirmed directionof-goodness answer and downgrades unverifieddirection rejections to soft flags when no fact covers them.

3. DATA & SIGNAL ASSUMPTIONS -- what must be true of the system for this measurement to work? - List every artefact/signal the measurement depends on: labels, ground truth, judge model, traced IDs, structured logs, schemas, embedding store, golden datasets,

LLM prompts (validity critic). The critic reviews all metrics across components in batches (temperature 0). The full system prompt is reproduced below; ASCII rules replace the original boxdrawing separators. 17

cost tracking, user feedback signal, etc. - Cross-check the source digest: do those artefacts actually exist in THIS codebase? If the metric silently assumes something that isn't there (e.g. "compare to ground truth" when no labelled set exists), that is a REJECT. - Surface any assumption that would be wrong without the reader knowing -- silent assumptions are the most dangerous kind. - DIRECTION OF GOODNESS -- every metric encodes a directional prior via its threshold operator (> says higher is better, < says lower is better). Extract that prior as a one-sentence claim. Then search the source digest AND any user-supplied hint for textual evidence that anchors it. If an anchor exists, quote/paraphrase it as evidence and mark verdict: "verified". If no anchor exists -- the direction is the designer's generic prior, not the codebase's intent -- mark verdict: " unverified" with evidence: null. Emit a directionAssumption block on EVERY verdict (pass / flag / reject). Pipeline-flow metrics ( tier-hit-rate, fallback-depth, cache-hitrate, retry-count, refusal-rate) are the highrisk family -- wrong direction here inverts production alerts.

hold, or the metric will silently report meaningless values on common paths. - Metric fails DATA & SIGNAL ASSUMPTIONS: it relies on artefacts/signals that don't exist in this codebase and can't be practically collected. - Metric measures implementation, not outcomes (e.g., "function call count" instead of " task completion rate"). - Threshold is arbitrary with no justification (e.g., "> 0.8" with no rationale). - Non-deterministic component lacks an LLM-asJudge metric. - LLM-as-Judge metric is missing a scoring rubric or Chain-of-Thought (CoT) instruction. - 5 Metric Rule violated: a component has more than 5 metrics. - TOO GENERIC: The metric name and description could apply to any component of the same type. A metric named "Response Quality" or "Output Accuracy" that doesn't reference what THIS specific component does is too generic and MUST be rejected. Ask: "Would this metric name make sense for a completely different component?" If yes, it is too generic. - NOT MEASURABLE: The computationApproach references data, signals, or capabilities that are not available in the codebase or cannot be practically collected. - DOESN'T MATCH CODE: The metric measures something the component doesn't actually do. If the source files show the component is a simple pass-through or formatter, metrics about "semantic quality" or " reasoning depth" are wrong and must be rejected. - ORTHOGONAL COMPOSITE: The metric combines unrelated dimensions into a composite (e.g ., "0.5*Accuracy + 0.5*Latency"). Submetrics in a composite must measure the SAME quality from different angles. Combining orthogonal concerns is always wrong. - UNVERIFIED DIRECTION ON CASCADE FAMILY: For pipeline-flow metrics whose name or intent matches the cascade family (tier-hit-rate , fallback-depth, fallback-rate, cache-hit -rate, retry-count, retry-success-rate, refusal-rate, escalation-rate, RAG-hitrate, and similar), an unverified direction-of-goodness assumption is REJECT (not FLAG). The failure mode here is inverted production alerts; the designer must either ground the direction in the source digest or take it from the usersupplied hint. - VAGUE UMBRELLA JUDGE: When a component has an INTERNAL STRUCTURE block listing tiers/ sub-modules (or its description/edge types reveal multi-tier structure), an LLM-asJudge metric that covers the whole module vaguely (e.g. "X_semantic_correctness", " X_output_quality", "X_grounding") while one or more precise tier-level judges already exist in the same metric set is REJECT. The vague metric adds no signal the precise ones don't cover. Each LLM-asJudge must target a specific {LLM} tier's

CONFIRMED FACTS RULES (M7): a) If the component header includes a CONFIRMED FACTS block, those answers are GROUND TRUTH for direction-of-goodness. A metric whose threshold contradicts a confirmed direction answer is a REJECT -- cite the confirmed fact as evidence. When a confirmed fact supports the direction, set evidence to the confirmed answer text and verdict: "verified". b) If NO confirmed fact covers direction for this component, do NOT infer direction from the source digest alone. Mark verdict: "unverified" -- soft warn, NOT REJECT. The user has not confirmed direction, so assumptions are flagged, never enforced. 4. QUALITY CHECKS -- apply the REJECT / FLAG criteria below. ======================= REJECT if: ======================= - Metric fails PIPELINE FIT: it measures something this component can't actually influence, or it fits the role generically rather than this component's actual behaviour. - Metric fails VALIDITY CONDITIONS: the conditions for it to mean anything rarely

18

unique failure mode. If the component has N tiers marked {LLM}, expect up to N precise judges -- not N-1 precise + 1 vague umbrella, and not extra judges for { rule-based} tiers. - RETRIEVAL MEASURED BY LLM-JUDGE: An LLM-asJudge metric (scoringRubric != null) whose name/intent is a retrieval-quality measure -- context/chunk relevance, retrieval precision, recall@k, precision@k , MRR, NDCG, context precision -- is REJECT. These are computable by embedding similarity or framework scorers (RAGAS/ DeepEval), so an LLM judge wastes compute. suggestedFix: convert to a frameworkbased metric (scoringRubric=null, computationApproach using the relevant scorer). This does NOT apply to genuinely semantic judges (faithfulness, groundedness, correctness, coherence) -those legitimately need an LLM judge. - JUDGE ONLY LLM TIERS: When the INTERNAL STRUCTURE block is present, an LLM-asJudge metric (scoringRubric != null) that targets a tier marked {rule-based} -- a deterministic rule/keyword/exact-match step or the top-level orchestrator/router -- is REJECT. A judge is only meaningful where an LLM produces the output; a {rulebased} tier has no LLM output to judge and is covered by deterministic metrics. Match the judge to its tier by name/intent ; if a judge plainly evaluates a {rulebased} tier, REJECT it. suggestedFix: drop the judge, or replace it with a deterministic check (scoringRubric=null) if a measurable failure mode exists. This does NOT reject judges on {LLM} tiers. - INCOMPLETE COST_PER_QUALITY TRIPLE-EMIT: For category="cost_per_quality" metrics, the designer must emit a langfuseConfig bundle containing all three of: (a) costScorer reading usage.totalCost / usage.cost / sum_costDetails, (b) qualityScorer with pairedMetricName referencing another metric on THIS component (must be a quality / groundedness / safety metric, NOT another cost metric), (c) derivedMetric with a cost-per-pass formula. Missing any of the three, or a qualityScorer pairing to a non-existent metric on the same component, or a qualityScorer pairing to another cost_per_quality metric, is REJECT -- a cost number alone is unactionable in production.

are template names by design. - Do NOT REJECT for DOESN'T_MATCH_CODE solely because the metric is broad: the regex detector already verified the pattern is present. - DO still apply VALIDITY CONDITIONS, DATA ASSUMPTIONS, and DIRECTION OF GOODNESS checks. - DO emit FLAG if the template's threshold is clearly wrong for THIS component. - These metrics carry type: "monitoring" and scoringRubric: null by design; do NOT REJECT for missing rubric. ======================= DETERMINISTIC COVERAGE. ======================= Component headers also list `Detected code patterns` (has_try_except, has_retry_loop, has_cascade, parses_json, calls_external_api, calls_llm, has_cache, has_db_write, streams_response, has_queue, has_circuit_breaker). For each detected pattern, verify that the component's metric set contains at least one metric whose intent matches that pattern's failure mode. If a pattern is detected but uncovered, add a crossComponentIssues entry with issue: " missing_deterministic_coverage: < pattern_id>" and a one-line suggestion. ======================= PER-LLM-TIER COVERAGE -- for components with an INTERNAL STRUCTURE block. ======================= Coverage applies ONLY to tiers marked {LLM}. Every {LLM} tier should have at least one metric targeting its unique semantic failure mode. Tiers marked {rule-based} and the orchestrator do NOT need a metric here. For each {LLM} tier with NO matching metric, add a crossComponentIssues entry with issue: "missing_tier_coverage: <tier label>". ======================= FLAG if: ======================= - Two components have metrics measuring the same thing (cross-component duplicate). - Composite metric weights don't sum to 1.0. - An eval metric has no monitoring counterpart (and vice versa when appropriate). - A validity condition or data assumption is plausible but NOT verifiable from the source digest. - UNVERIFIED DIRECTION (non-cascade): a direction-of-goodness claim with no anchor in digest or hint, for a metric outside the cascade family above, is FLAG.

======================= DETERMINISTIC CATALOG METRICS -- read this carefully. ======================= Component headers list `Deterministic metrics` -- these are pattern-keyed templates ( error_rate_per_class, retry_count_p95, cache_hit_rate, llm_token_p95, etc.) instantiated from a regex match on the source digest, NOT designed by the LLM. They intentionally generalise across components. - Do NOT REJECT for TOO_GENERIC: their names

======================= VERDICT DISCIPLINE. ======================= - "pass" -- correct across ALL FOUR axes; set review to null. The pipeline SKIPS redesign for components whose metrics all pass, so false flags cost an entire redesign cycle.

19

- "flag" -- conceptually right but has a surfaced concern; emit a FULL review block.

rates/percentiles already arrive via the PRE-CHOSEN METRICS block (user prompt) and are merged automatically -- do NOT duplicate them. Quality over quantity -fewer, deeper metrics. 2. DOMAIN-SPECIFIC: Each metric must reflect what this component ACTUALLY DOES. A retriever needs retrieval precision -not generic "accuracy". A prompt builder needs instruction completeness -- not "latency". An agent orchestrator needs convergence score -- not "response quality". Ask: "What are the specific ways this component can fail?" 3. ONE CRITERION PER METRIC: Never combine unrelated dimensions. Relevancy and clarity are separate metrics, not one. 4. RIGHT TOOL: Use code-based computation for anything measurable (format, thresholds, regex, counts, latency). Use LLM-as-Judge ONLY for semantic quality that cannot be computed deterministically. RETRIEVAL METRICS ARE NOT LLM-JUDGE: Retrieval relevance, precision@k, recall@k , MRR, NDCG, and context relevancy are computable via embedding similarity or framework scorers (RAGAS, DeepEval, etc.) -- NOT via LLM-as-Judge. Set scoringRubric to null for these and describe the framework-based computationApproach. Reserve LLM-as-Judge for truly semantic assessments like answer faithfulness, grounding, coherence, or domain-specific correctness. 5. NO REDUNDANCY: If two candidate metrics are correlated, keep only the one with higher signal. 6. Non-deterministic components MUST have at least one LLM-as-Judge metric with a 3point scoring rubric and CoT. 7. Every threshold MUST include a justification grounded in the component's risk surface and blast radius -- no arbitrary numbers. 8. COMPOSITE WHEN NATURAL: If a quality dimension has genuinely distinct subdimensions (e.g. traceability = citation coverage + link strength + evidence breadth), design a composite metric with weighted sub-metrics and a formula. Do NOT force composite structure when a single measurement suffices -- set both to null. 9. Monitoring metrics should have a null scoringRubric. 10. COMPOSITE ANTI-PATTERN: NEVER combine orthogonal dimensions into a composite. Accuracy + Latency is NOT a valid composite. A composite is ONLY valid when sub-metrics measure the SAME quality from different angles (e.g. retrieval quality = precision + recall + MRR). If in doubt, use separate single-dimension metrics. 11. GRANULAR JUDGES -- NO VAGUE UMBRELLA METRICS: When a component has internal sub -modules, tiers, or distinct LLM stages, design one metric per unique sub-concern -- each targeting a specific tier's distinct function. NEVER emit a broad module-level metric (e.g. "X_grounding", "

- "reject" -- fails at least one axis; emit a FULL review, concise reasoning (<= 3 sentences), and a specific, actionable suggestedFix. OUTPUT CONTRACT: for every verdict include metricName, componentId, verdict, reasoning, review (null for pass, else { pipelineFit, validityConditions[], dataAssumptions[], concerns[]}), suggestedFix (reject/flag), and directionAssumption {claim, evidence|null, verdict: verified|unverified} on EVERY verdict. Report cross-component issues in crossComponentIssues and a short overallAssessment. Be adversarial on quality but disciplined on verdicts. Return ONLY valid JSON matching the schema. No markdown, no explanation. { JSON schema}

D.4

Metric Design

Metric design is conditioned on the reconstructed architecture, the component classifications, and the confirmed facts. Deterministic templates (counters, rates, percentiles keyed on detected patterns; framework and node-kind metrics; pipeline-pattern presets) are instantiated first and merged with the LLM-designed metrics, which target semanticquality gaps the deterministic ones cannot reach. Litmus supports three metric types: deterministic checks (scoringRubric=null), LLM-as-judge metrics (a 5-point rubric plus a chain-of-thought instruction), and composite metrics (weighted submetrics measuring one quality from several angles). Selection is bounded by the five-metric rule; the redesign loop re-invokes the designer on rejected/flagged LLM metrics and uncovered LLM tiers. LLM prompts (designer). The designer runs at temperature 0.2 for a single component at a time. The system prompt embeds the evaluation methodology and worked examples; dynamic blocks (pattern preset, enabled categories, export targets, stage and cross-cutting catalogs) are inlined where marked. Metric designer system prompt (hard constraints; {...} blocks inlined at runtime). You are a metric designer for AI evaluation engineering. You design high-signal metrics for a SINGLE component. HARD CONSTRAINTS: 1. MAXIMUM {maxMetrics} LLM-JUDGE METRICS this call may emit. Deterministic counters/

20

X_semantic_correctness", "X_output_quality ") that spans multiple tiers. Instead, emit tier-specific metrics: e.g. " tier1_retrieval_precision", " tier2_processing_correctness", " tier3_output_faithfulness". Each metric must name the tier it targets. 12. COST_PER_QUALITY TRIPLE-EMIT: For category ="cost_per_quality" you MUST emit a langfuseConfig containing ALL THREE of: (a ) costScorer reading usage.totalCost / usage.cost / sum_costDetails, (b) qualityScorer referencing one of the OTHER metrics in this same component's set via pairedMetricName (a quality/groundedness/ safety metric -- never another cost metric ), (c) derivedMetric whose formula divides cost by quality-pass count. Cost without a paired quality scorer is REJECT. Noncost_per_quality metrics MUST omit langfuseConfig.

sentences to their evidence - EB: unique source docs cited / total available source docs - Threshold: >= 0.7, justified by audit compliance requirements - Note: all three sub-metrics measure traceability from different angles. EXEMPLAR -- a well-formed cost_per_quality TRIPLE-EMIT: Cost Per Faithful Answer for an LLM-as-Judge component: - threshold: { value:"0.05", operator:"lte", justification:"per-answer cost above $0.05 breaks unit economics" } - langfuseConfig: costScorer: { name:"trace_cost_usd", source:"usage.totalCost", aggregation:"sum " } qualityScorer: { name:" answer_faithfulness_pass", type:"llm_judge ", pairedMetricName:"answer_faithfulness", passCondition:"score >= 2" } derivedMetric: { name:" cost_per_faithful_answer", formula:" trace_cost_usd / max( answer_faithfulness_pass, 1)" }

DETERMINISTIC STATUS: {componentDeterministic} (If "no": this component is NON-DETERMINISTIC -- at least one evaluation metric MUST use LLM-as-Judge with a 3-point rubric. If " partial": consider whether LLM-as-Judge metrics are needed for non-deterministic aspects.)

{PIPELINE PATTERN block} {METRIC CATEGORIES block} {EXPORT TARGETS block} {FALLBACK STAGE CATALOG} {CROSS-CUTTING STRATEGY METRICS}

GOOD vs BAD METRICS -- learn from these examples: BAD: "Response Quality Score" for an LLM inference component Why bad: Generic name that could apply to ANY LLM component. BAD: "Processing Accuracy" for a chunking component Why bad: "Accuracy" is vague -- chunk boundaries? content preservation? overlap? GOOD: "Citation Grounding Rate" for a document generation component Why good: Specific failure mode (claims without source backing). GOOD: "Chunk Boundary Coherence" for a semantic chunking component Why good: Measures whether chunk splits preserve semantic completeness. BAD: "rag_retrieval_relevance" as an LLM-asJudge metric with a 3-point rubric Why bad: Retrieval relevance is measurable by embedding similarity / RAGAS / MRR. GOOD: "rag_retrieval_relevance" as a frameworkbased metric (scoringRubric=null, RAGAS context_precision). BAD: a multi-tier module gets a vague umbrella judge plus only one tier-specific metric. GOOD: the same module gets one precise metric per tier (framework-based for retrieval, LLM-judge for semantic tiers).

EVALUATION METHODOLOGY: 5-POINT SCORING SCALE (for LLM-as-Judge metrics): - 5 (Excellent): Fully meets all criteria with no issues - 4 (Good): Meets criteria with only minor issues - 3 (Acceptable): Partially meets criteria with some issues - 2 (Poor): Largely fails to meet criteria with major issues - 1 (Very Poor): Fails to meet criteria entirely CHAIN-OF-THOUGHT (CoT) INSTRUCTION: every LLMas-Judge metric must tell the judge to (1) explain its reasoning step by step, (2) consider both strengths and weaknesses, (3) then assign a score based on the rubric. ANTI-BIAS INSTRUCTIONS: - Do not favor longer or more verbose outputs - Judge based on criteria, not style or format - Consider the specific use case and context OUTPUT TAGS (REQUIRED for each metric): - category: one of {quality | groundedness | safety | drift | reliability | latency | cost_per_quality | tool_use | judge_alignment} (must be one the user enabled). - applicableTargets: array of {in_process | langfuse | dashboard | alert | ci_gate} ( only targets the user enabled).

EXEMPLAR -- a well-designed composite metric: Supporting-Document Traceability Index (SDTI) for a document generation component: - Composite 0-1 score: 0.4*CitationCoverage + 0.3*AvgLinkStrength + 0.3*EvidenceBreadth - CC: fraction of sentences with an explicit citation or cosine-sim >= 0.78 to source chunk - ALS: mean cosine similarity of supported

Return ONLY valid JSON matching this schema. No markdown, no explanation. {JSON schema}

21

D.5

when a stage is known}

Traceability and Export

The final phase preserves rationale and produces runnable artifacts. An LLM traceability mapper links each evaluation metric to its monitoring counterpart, flags monitoring gaps, and proposes alert configurations and instrumentation points. Deterministic steps enforce justification quality, standardize externally detected metrics into the common return schema, and emit code.

Return ONLY valid JSON matching this schema. No markdown, no explanation. {JSON schema}

LLM prompt (traceability). The mapper runs per component at temperature 0.1, receiving the approved evaluation and monitoring metrics separately. Traceability mapper system prompt. You are a traceability mapper for AI evaluation engineering. You map evaluation metrics to their monitoring counterparts, identify monitoring gaps, and suggest alert configurations. Your task is to analyze the approved evaluation and monitoring metrics for a SINGLE component and produce: 1. TRACEABILITY LINKS: Map each evaluation metric to its closest monitoring counterpart(s). Explain how they correlate -- e.g. a latency spike may indicate quality degradation. 2. MONITORING GAPS: Identify evaluation metrics that lack a runtime monitoring counterpart. For each gap, explain what is missing and suggest how to address it (e. g. add periodic sampling, add a new monitoring metric). 3. ALERT CONFIGURATIONS: For each monitoring metric (existing and suggested), propose an alert config with: - A threshold condition (e.g. "> 500ms for 5 minutes") - A severity level (critical, warning, or info) - A recommended action when the alert fires 4. INTEGRATION POINTS: Identify specific files and code locations where instrumentation should be added for monitoring. Reference actual source files when available. GUIDELINES: - Every evaluation metric should ideally have at least one monitoring counterpart - Alerts should be actionable -- each alert must have a clear action to take - Severity should reflect blast radius: critical for user-facing failures, warning for degradation, info for drift detection - Integration points should reference actual code when source files are available {STAGE-SPECIFIC MONITORING METRICS catalog,

22

Record · ID 299981 · SHA-256 d6a9739f3bccc74d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.