Preprint
D EPTH B ENCH CAD: W HEN D OES D EEPER AUDITING Y IELD M ORE R ELIABLE C ONCLUSIONS ? Zhihao Xie Independent Researcher [email protected]
arXiv:2609.15122v1 [cs.SE] 14 Sep 2026
Hongye Yang College of Computing Georgia Institute of Technology [email protected]
Boxiao Huang College of Computing Georgia Institute of Technology [email protected]
Shengjun Xiong Independent Researcher [email protected]
A BSTRACT Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations that can be evaluated, which can ultimately make model-level estimates less accurate. We study this phenomenon and the conditions under which it arises. We decompose behavioral evaluation into three evidence levels—task templates, stochastic generations, and within-program edits—define an average failure risk that is invariant to audit depth, and combine three-level variance with measured execution costs to analyze the tradeoff between deeper edit auditing and broader independent coverage. Experiments across two CAD environments and five generation systems show that the value of deeper auditing depends on where evaluation uncertainty originates. When template heterogeneity or generation stochasticity dominates, additional edit checks can increase total estimation error; when within-program state variation is large and generation is expensive, deeper auditing is more valuable. Variance and cost estimates from calibration predict the direction of this change and provide a diagnostic basis for allocating evidence on held-out tasks. These results show that the thoroughness of program inspection can diverge from the reliability of model evaluation, and they help determine whether the next unit of budget should be spent on a new task, a new generation, or additional edit checks. Keywords: generative CAD; behavioral evaluation; multistage sampling; finite population; costaware evaluation
1
I NTRODUCTION
Outputs from parametric CAD models must generate correct geometry at default parameters and remain executable while preserving geometric constraints and design intent under subsequent edits. DeepCAD represents the CAD modeling process as an operation sequence (Wu et al., 2021), while Fusion 360 Gallery provides real design histories and programmatic construction data (Willis et al., 2021). Text2CAD maps natural language to parametric CAD sequences (Khan et al., 2024), and CAD-Recode uses language models to recover executable CAD code from geometric input (Rukhovich et al., 2025). A program may appear correct in its initial state yet fail to build, selfintersect, invalidate parameters, or violate constraints after changes to thickness, hole diameter, spacing, or array count. Static geometric metrics therefore do not fully capture the editability or engineering utility of CAD programs. 1
Preprint
Behavioral evaluation of generative CAD contains three levels of evidence: task templates cover different part families and design difficulty; independent generations from the same template capture model stochasticity; and counterfactual edit audits characterize behavior across valid parameter states within a program. BenchCAD incorporates execution verification, parameter reasoning, and code editing into a programmatic CAD benchmark (Zhang et al., 2026). Text2CAD-Bench further provides systematic evaluation of text-to-parametric-CAD generation (Wang et al., 2026). Increasing evidence at any of the three levels can reduce evaluation uncertainty, but the statistical role and execution cost of each level differ. Under a fixed budget, excessive audit depth reduces task coverage, whereas shallow auditing may miss conditional failures. The numbers of templates, generations per template, and audits per program must therefore be chosen jointly. To ensure that different audit depths evaluate the same capability, we freeze the valid edit states for each template in advance and define the model-level estimand as the average probability of behavioral failure when randomly sampling a template, one model generation, and one edit state. Audit depth changes only the precision with which this estimand is measured, allowing budget allocations to be compared directly. Supplementary motivation and protocol explanation appear in Appendix J. Our main contributions are: (1) We identify and empirically characterize a separation between the thoroughness of program inspection and the accuracy of model-level estimation, showing how deeper auditing under a fixed budget can increase estimation error by reducing task coverage or independent generations. (2) We explain this effect through three-level variance and execution cost and test, on held-out data, whether these quantities predict the direction of the benefit from deeper auditing. (3) We evaluate how evidence allocation affects model-evaluation reliability through risk estimation, interval coverage, and judgment error, while explicitly reporting the associated calibration cost.
2
R ELATED W ORK
Executable CAD generation retains design histories and construction logic through operation sequences, programmatic data, and recovered code (Wu et al., 2021; Willis et al., 2021; Khan et al., 2024; Rukhovich et al., 2025). BenchCAD and Text2CAD-Bench extend evaluation toward execution, parameter reasoning, and editing (Zhang et al., 2026; Wang et al., 2026). Our analysis concerns how to allocate evaluation evidence under a fixed budget. It builds on multistage sampling (Neyman, 1934; Horvitz & Thompson, 1952), clustered observations (Liang & Zeger, 1986), and work on stochastic evaluation and selection bias (Henderson et al., 2018; Reimers & Gurevych, 2017; Dietterich, 1998; Varma & Simon, 2006). Detailed CAD, representation-learning, judge-validation, and sampling literature appears in Appendix H.
3
M ETHOD
To characterize how much independent evidence each evaluation record actually provides, we model generative-CAD behavioral evaluation as a three-level nested sampling process over templates, generations, and edits. This section defines the estimand, sampling estimator, variance decomposition, cost constraint, and calibration strategy. Letting the design structure determine effective sample size follows basic principles of design-based sampling and finite-population inference (Neyman, 1934; Horvitz & Thompson, 1952). 3.1
E VALUATION TARGET AND F INITE E DIT P OPULATION (s)
Let Ytim ∈ {0, 1} denote the outcome for system s on template t for generation i evaluated at state (s) m, where Ytim = 1 indicates an execution error, timeout, invalid topology, geometric-constraint violation, or parameter response inconsistent with the task specification. The failure risk of a single program over the complete edit population is 2
Preprint
M
µs (t, i) =
1 X (s) Y . M m=1 tim
(1)
The model-level estimand is defined as i h Rs = ET EI (s) |T µs T, I (s) .
(2)
Rs is the average probability of behavioral failure when randomly sampling a task template, a model generation, and a valid edit state. Actual audit depth affects estimation cost and variance but does not change this estimand. Single-program certification, all-state pass rate, and worst-case risk are different evaluation targets and are outside the scope of this paper. 3.2
T HREE -L EVEL S AMPLING AND R ISK E STIMATION
Given a formal evaluation budget, we first sample q templates from the available pool; then independently generate g programs for each template; finally, from each program's M frozen states, we sample k states without replacement. Let Sti denote the set of sampled states for program (t, i); the model-risk estimator is q g X 1 X 1 X (s) bs (q, g, k) = 1 R Ytim . q t=1 g i=1 k
(3)
m∈Sti
The three sampling levels serve different roles: q controls task coverage; g controls generation uncertainty within a task; k controls measurement error in an individual program's risk. Multiple generations from the same template share task difficulty, so the template is the outermost unit for variance estimation, data splitting, and bootstrap. Keeping the template as the outermost cluster also avoids treating correlated repeated observations as independent samples (Liang & Zeger, 1986). When comparing multiple systems, all systems use the same templates and edit-state indices to reduce comparison noise from task difficulty. A template and all of its generated programs are always kept as an intact cluster and never split across calibration, validation, or test sets. 3.3
T HREE -L EVEL VARIANCE D ECOMPOSITION 2 Let σT2 = V arT EI|T µs (T, I) denote between-template variance, σG = ET V arI|T µs (T, I) h i (s) 2 2 within-template generation variance, and σE = ET,I SE Ytim within-program edit-state variance, respectively. Under a balanced sampling design, the superpopulation variance of Equation (3) is σ2 2 σG 1 1 1 2 T b + + − σE . V ar Rs = q qg qg k M
(4)
Equation (4) shows that increasing audit depth k can reduce only edit-state sampling error; betweentemplate variation and generation stochasticity still require increasing q and g. For a finite reference pool with Q templates, G complete generations per template, and M states per program, resampling replay uses the finite-population variance in Equation (5). Estimation without replacement and inclusion-probability methods provide the classical basis for this resampling design on a frozen reference pool (Horvitz & Thompson, 1952). 1 1 1 1 1 1 1 1 2 2 bs = V arref R − ST2 + − SG + − SE . q Q q g G qg k M
(5)
Equation (5) is used to validate the method on a finite test pool. Confidence intervals for the future task distribution use templates as bootstrap clusters and do not treat the finite template set as the 3
Preprint
complete task population. We use the bootstrap to estimate uncertainty for the future task distribution; its statistical basis and confidence-interval properties are well established (Efron, 1979; Efron & Tibshirani, 1986). 3.4
C OST M ODEL AND J OINT A LLOCATION
Let cT be the preparation and scheduling cost of introducing a new template, cI be the cost of one model generation, parsing, and initial build, and cE be the cost of one parameter injection, program re-execution, and automated judgment. The total cost of a balanced design is C(q, g, k) = qcT + qg (cI + kcE ) .
(6)
Because systems have different generation and execution costs, our shared strategy freezes the number of generations per template g and audit depth per program k, together with the budget-mapping rule, and measures the cost parameters for system s before formal evaluation. The maximum number of templates available to that system within the budget is qs (g, k) = min Q,
C0 cT,s + g (cI,s + kcE,s )
.
(7)
Thus, different systems may use different numbers of templates qs while sharing the same g, k, and budget-allocation rule. Normalizing by the cost of one edit execution cE and defining rT = cT /cE , rI = cI /cE , the budget constraint becomes qrT + qg (rI + k) ≤ C0 .
(8)
We use discrete search to jointly select the three sampling parameters: ars (q, g, k), arg min Vd q,g,k∈N
s.t.
C(q, g, k) ≤ C0 .
(9)
The search range satisfies 1 ≤ q ≤ Q,1 ≤ g ≤ G and 1 ≤ k ≤ M . For each (g, k), the algorithm chooses the largest number of templates q allowed by the budget and then compares all feasible integer solutions. This procedure handles template saturation, finite-population corrections, and residual integer budget. 3.5
C ALIBRATION C OST AND U SAGE S TRATEGY
Calibration uses independent templates to estimate variance and cost. A pooled (g, k) configuration requires only target-system cost measurement; system-level calibration diagnoses deviations and reuse scenarios. Appendices A.3, F.4, and J give the allocation and calibration details.
4
E XPERIMENTAL D ESIGN
4.1
DATA AND E VALUATION E NVIRONMENTS
We compare five generation systems, S1–S5, on DepthBenchCAD. All five use a common CADgeneration interface and differ only in the underlying language model; each system receives the same task specification and parameter requirements and outputs an executable CadQuery program. System mappings, fixed runtime configurations, random-seed rules, and cost-measurement protocols are provided in Appendix D. The experiments use a primary environment A and a cross-task-family transfer environment B. Environment A contains eight task families and is used to test whether calibration parameters predict the benefit of deeper auditing on held-out templates. Environment B contains six task families disjoint from A and is used to test cross-family transfer of allocation strategies. The environments share 4
Preprint
the same risk definition and four edit protocols; complete task-family compositions and execution settings appear in Appendices B and D. Table 1: Experimental data matrix Env.
Tpl.
Fam.
Cal./Test
Gen./tpl.
States/prog. Sys.
Records
DepthBenchCAD-A DepthBenchCAD-B
72 48
8 6
24/48 12/36
5 4
16 16
28,800 15,360
5 5
Using one edit execution as the cost unit, the template cost is 2. Generation costs for S1–S5 are 1, 3, 8, 16, and 32 in environment A, and 2, 6, 16, 32, and 64 in environment B. The main reported budget is C0 =1024, with total cost q[2+g(rI+k)]. The main experimental results use a frozen, real, fully audited pool. Failure labels come from actual builds, counterfactual edits, and automated judgments. Alternative evidence allocations are evaluated only by sampling without replacement from observed records and by Monte Carlo replay. Complete execution, cost, and replay protocols are provided in Appendices C, D, and F. 4.2
C OUNTERFACTUAL E DIT S TATES AND AUTOMATED J UDGE
Each template contains 16 pre-frozen edit states, evenly divided among four categories. Table 2: Composition of counterfactual edit states Type
States
Intervention target
Primary checks
Local edit
4
Boundary edit Linked edit
4
Semantic edit
4
Routine change to one parameter Parameter values near the valid-range boundary Simultaneous changes to multiple related parameters High-level design-specification changes
Parameter binding, geometric response, and main-body connectivity Self-intersection, zero volume, minimum spacing, and manufacturing constraints Symmetry, spacing, arrays, and dimensional relations Coordinated response of related parameters, geometry, and design semantics
4
The state generator reads only the task specification and never the candidate program implementation. Each state records parameter values, units, valid ranges, expected geometric relations, and judgment tolerances. Complete definitions of the 16 states are provided in Appendix B. The automated judge checks, in order, successful build completion, valid entities and connected components, topology and hole/circular features, parameter response, dimensional/spacing/symmetry/containment relations, and high-level semantic constraints. The full contract for numerical tolerances, gray-zone routing, isolated execution, and geometry probes appears in Appendix C. If the initial state fails to build or violates nominal constraints, all 16 edit states for that program are counted as failures in the primary analysis. We also report conditional risk among initially valid programs to distinguish initial-generation failure from failures introduced during editing. Expert validation. The automated judge is validated on 800 stratified records reviewed by two experts. The experts label each record independently and blindly before a consensus label is formed; overall risk and confusion metrics are weighted by inverse sampling probability. Appendix E gives the full sampling, adjudication, agreement, and weighted-estimation protocol. 4.3
DATA S PLITS AND C ALIBRATION S ETUP
Splits are stratified by family at the template level: A uses three calibration and six test templates per family (24/48 total), and B uses two and six (12/36). All generations and states stay with their 5
Preprint
template. Full calibration audits all 16 states; smaller pilots test shrinkage and reuse. Test data never select variance parameters, shrinkage weights, depth, or costs; the full test pool supplies only reference risk and replay outcomes (Varma & Simon, 2006). Appendices D.3, F.1, and J retain the extended protocol. 4.4
C OMPARISON M ETHODS AND B UDGETS
We compare fixed audit depths, a two-level degenerate model, pooled and leave-one-system-out pooling, full system-level calibration, small calibration with shrinkage, edit-type stratification, and paired allocation for model pairs. Main-text results focus on fixed depth, pooled allocation, and system-level/stratified candidate configurations. Formal selection rules, budget mappings, and paired definitions are provided in Appendix F. 4.5
E VALUATION M ETRICS AND S TATISTICAL I NFERENCE
Risk-Estimation Efficiency We measure estimation error at both the program and model levels. Program-level MSE is the mean squared error between the failure proportion estimated from k sampled states and the failure proportion over all 16 states for that program, averaged across programs. Model-level MSE uses the mean failure risk of the full test pool as the reference and reports J=C0 ×MSE. For each fixed audit depth, only the calibration set is used to choose g, and the largest feasible q is determined by the cost constraint. Program-level MSE is estimated by state-subsampling replay, while model-level MSE is computed from the design variance of the finite test pool. Test-pool results are used only for post hoc validation and never for configuration selection. Benefit prediction is preregistered for two comparisons: k=4→8 and k=8→16. We define ∆J=Jdeep −Jshallow , where a negative value indicates that deeper auditing is beneficial. We compare the sign predicted during calibration with the actual ∆J in the test pool. Wrong-recommendation loss is the excess J of the recommended configuration over the smaller J of the two alternatives. In addition to J and benefit direction, we report regret relative to the oracle, 5% robust coverage, 95% interval coverage, automated-judge validity, and paired model decisions. Definitions, repetition counts, and statistical inference for these auxiliary metrics are given in Appendices E and F. Interval experiments use q=16, g=3, k=8 throughout and run 5,000 replays without replacement.
5
E XPERIMENTAL R ESULTS
This section addresses three questions: whether ignoring evidence hierarchy produces overconfident model conclusions; whether three-level variance and cost predict the direction of benefit from deeper auditing; and when recalibration is worth its additional cost. 5.1
E VIDENCE H IERARCHY AND E VALUATION R ELIABILITY
When generated programs are treated as independent units, mean interval coverage across the five systems is 91.81%; using the three-level design variance increases it to 95.26%. The difference is largest for S1 and S2, which exhibit stronger template-level correlation: coverage rises from 89.06% and 86.82% to 94.28% and 94.42%, respectively. Differences for S3–S5 are smaller, consistent with their weaker template-level heterogeneity. These results describe interval performance on the finite test pool and do not change the common risk point estimate used by both methods. Figure 1 compares program-level and model-level estimation errors at different audit depths in the primary environment. More thorough program inspection does not improve model-level estimation accuracy for every system. When S1's audit depth increases from 4 to 16, program-level MSE falls from 0.014284 to zero, but the total number of generated programs drops from 184 to 48 and model-level J rises from 0.0974 to 0.1649, an increase of about 69.3%. Full-state auditing eliminates edit-sampling error for these programs but cannot compensate for the loss of independent generations. 6
Preprint
Figure 1: Audit depth and estimation error in the primary environment. (a) Program-level MSE. (b) Model-level cost-scaled error, J=C0 ×MSE. The formal budget is C0 =1024; the numbers of templates q and generations g vary by configuration. Colors and markers denote S1–S5, and lines connect the three reported depths. At k=16, program-level MSE is zero, while model-level error may still increase or decrease. Exact values are reported in Table 18.
S4 shows the opposite pattern. Increasing depth from 4 to 16 reduces the number of templates from 46 to 30, yet J falls from 1.2881 to 0.5378, a decrease of about 58.3%. Here, the information gained from additional edit states is sufficient to offset the loss of coverage. S2 changes only slightly, indicating that some systems are nearly indifferent across audit depths. Across the ten predefined comparisons in environment A, calibration predicts the correct direction in nine. The sole error is S2 for k=4→8: predicted ∆J=−0.0121, while actual ∆J=0.0039. Mean wrong-recommendation loss across the ten comparisons is 0.000389. After local calibration in environment B, all ten directions are predicted correctly. These results support the three-level variance-and-cost explanation of audit benefit, although the twenty comparisons constitute only a limited mechanism validation. Figure 2 shows three classes of candidate configurations selected under a common calibration protocol. System-level and stratified configurations reduce estimation error for some systems, but neither pooled nor system-level configurations consistently outperform a strong fixed-depth baseline. The three-level model therefore provides a stable explanation of the source and direction of audit benefit, while the optimality of a particular configuration still depends on system variance structure, cost, and finite calibration data. A strong fixed depth remains an important baseline. Under the unified analysis protocol, fixed k = 16, the stratified configuration, the system-level configuration, and the cross-system pooled configuration have five-system mean J values of 0.471, 0.481, 0.502, and 0.533. We therefore do not interpret calibration-driven joint allocation as outperforming the best fixed depth on average. Its primary role is to explain why different systems prefer different audit depths and to predict the direction of benefit from increasing audit depth. 5.2
C ALIBRATION C OST AND T RANSFER S TABILITY
Calibration is substantially more expensive than a single formal evaluation: full calibration costs about 2.04–5.67 times the formal budget, while small calibration costs about 28.1%–76.6%. Any local gain from system-level allocation must therefore be weighed against calibration expenditure and the number of times calibration can be reused. Complete values are given in Table 13. Repeated template splits reveal finite-sample variability in audit-depth selection. Across random splits, median relative regret is 1.041 and mean relative regret is 1.098. Tail degradation is more 7
Preprint
Figure 2: Error ratios of three allocation strategies relative to the fixed k=16 baseline in the primary environment. Each point is the strategy's J divided by J for the same system at fixed k=16. The dashed line marks a ratio of 1; points to the left have lower error. Circles, squares, and diamonds denote pooled, system-level, and edit-stratified strategies, respectively. The formal budget is C0 =1024; calibration cost is accounted for separately. Exact values are reported in Table 19.
pronounced under leave-one-task-family-out validation: P90 relative regret reaches 1.709 for S1 and 1.608 for S2, showing that task-family coverage affects allocation transfer. Complete stability results appear in Table 14. The second CAD environment also exhibits different configuration preferences. Among fixed-depth baselines, k = 10 has the lowest mean J at 0.742. The pooled configuration transferred directly from the primary environment keeps g = 1, k = 12, with mean J = 0.756. Leave-one-system-out pooling has mean J of 0.790, while small calibration with four-way stratification achieves the lowest mean in the table, J = 0.667. Configurations are compared in Figure 3. Cross-task-family transfer shows that the primary-environment configuration can directly transfer its frozen g, k, but optimality is not guaranteed on a new task set. Changes in task composition alter both variance structure and execution cost, so transfer still requires remeasuring costs and comparing the transferred configuration against fixed-depth and calibrated strategies in the target environment.
5.3
J UDGE VALIDITY AND M ODEL D ECISIONS
Dual-expert validation shows that, under inverse-sampling-probability weighting of the expert sample, the automated judge and expert consensus preserve the same risk ranking across all five systems. The automated judge underestimates risk for S4 and S5 by about 1.97 and 3.65 percentage points, respectively, while S2's automated labels exactly match expert consensus. Complete weighted risks and classification metrics are reported in Table 12, with the validation protocol in Appendix E. Programs that fail the initial build contribute 16 failed states to the primary risk. We also report conditional edit risk restricted to initially valid programs. This dual reporting preserves end-to-end failure probability while enabling analysis of parameter editability after a program builds successfully. In model-pair experiments, paired allocation provides only a small improvement over a strong fixeddepth baseline and the benefit largely disappears at high budgets. Its value is concentrated in comparisons with small risk gaps, low paired variance, and reusable calibration cost. Complete correctdecision rates at three budgets are reported in Table 17. 8
Preprint
Figure 3: Mean cost-scaled error J for each strategy in cross-task-family environment B; lower is better. Blue highlights the best reported fixed depth, k=10; orange denotes direct transfer from A with g=1, k=12; green denotes small calibration with edit stratification; other strategies are gray. The formal budget is C0 =1024, with additional calibration cost reported separately. Exact values are reported in Table 20.
6
D ISCUSSION
Reliable evaluation requires matching independent evidence to the dominant uncertainty source. Template heterogeneity favors additional tasks, generation variability favors independent programs, and substantial edit-state variation with expensive generation favors deeper auditing. Joint allocation explains system-specific deviations and audit-benefit directions, but does not consistently outperform strong fixed-depth baselines. Calibration is useful only when its gains and reuse justify its additional cost. Transfer results also caution against extrapolating a configuration beyond its task composition (Torralba & Efros, 2011; Geirhos et al., 2020). The estimand covers 16 frozen states, not a continuous parameter space or open-ended interaction. The variance model assumes a prespecified sampling design and within-level exchangeability; correlated costs, nonrandom timeouts, or adaptive stopping require additional treatment. Small calibration sets introduce parameter uncertainty, and the judge has limited semantic coverage despite expert validation. The tested families, systems, and kernels also constrain generalization. Average failure risk does not establish engineering safety or worst-case correctness. Appendix I retains the full discussion of practical use, calibration, transfer, pairing, and limitations (Mitchell et al., 2019).
7
C ONCLUSION
Under a fixed budget, more thorough CAD-program auditing can reduce the reliability of model evaluation by sacrificing task coverage and independent generations. Three-level variance and measured execution costs explain this conflict and predict audit-benefit directions across the tested systems and environments. No depth is universally best, and joint allocation does not consistently outperform strong fixed-depth baselines. Its value lies in diagnosing the dominant uncertainty source and guiding budget adjustment, with calibration cost, task composition, and reuse determining practical efficiency. Appendix J retains the extended conclusion.
9
Preprint
AI U SE S TATEMENT ChatGPT and Codex were used to assist with literature search, language polishing, code editing, translation, and related writing tasks. All AI-assisted content, including factual statements, analyses, and conclusions, was independently reviewed and verified by the authors. The authors take full responsibility for the accuracy and integrity of the manuscript.
R EFERENCES R. Artstein and M. Poesio. Survey Article: Inter-Coder Agreement for Computational Linguistics. Computational Linguistics, 34(4):555–596, 2008. doi: 10.1162/coli.07-034-R2. URL https: //doi.org/10.1162/coli.07-034-R2. A. Avetisyan, M. Dahnert, A. Dai, M. Savva, A. X. Chang, and M. Niessner. Scan2CAD: Learning CAD Model Alignment in RGB-D Scans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. doi: 10.1109/CVPR.2019.00272. URL https://doi.org/10.1109/CVPR.2019.00272. D. Card, P. Henderson, U. Khandelwal, R. Jia, K. Mahowald, and D. Jurafsky. With Little Power Comes Great Responsibility. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020. doi: 10.18653/v1/2020.emnlp-main.745. URL https://doi.org/10.18653/v1/2020.emnlp-main.745. Z. Chen, A. Tagliasacchi, and H. Zhang. BSP-Net: Generating Compact Meshes via Binary Space Partitioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. doi: 10.1109/CVPR42600.2020.00012. URL https: //doi.org/10.1109/CVPR42600.2020.00012. J. Cohen. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement, 20(1):37–46, 1960. doi: 10.1177/001316446002000104. URL https: //doi.org/10.1177/001316446002000104. B. Deng, K. Genova, S. Yazdani, S. Bouaziz, G. Hinton, and A. Tagliasacchi. CvxNet: Learnable Convex Decomposition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. doi: 10.1109/CVPR42600.2020.00011. URL https: //doi.org/10.1109/CVPR42600.2020.00011. T. G. Dietterich. Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation, 10(7):1895–1923, 1998. doi: 10.1162/089976698300017197. URL https://doi.org/10.1162/089976698300017197. B. Efron. Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics, 7(1):1–26, 1979. doi: 10.1214/aos/1176344552. URL https://doi.org/10.1214/aos/117634 4552. B. Efron and R. Tibshirani. Bootstrap Methods for Standard Errors, Confidence Intervals, and Other Measures of Statistical Accuracy. Statistical Science, 1(1):54–75, 1986. doi: 10.1214/ss/11770 13815. URL https://doi.org/10.1214/ss/1177013815. R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann. Shortcut Learning in Deep Neural Networks. Nature Machine Intelligence, 2:665–673, 2020. doi: 10.1038/s42256-020-00257-z. URL https://doi.org/10.1038/s42256-020-002 57-z. H. Guo, S. Liu, H. Pan, Y. Liu, X. Tong, and B. Guo. ComplexGen: CAD Reconstruction by B-Rep Chain Complex Generation. ACM Transactions on Graphics, 41(4):129, 2022a. doi: 10.1145/3528223.3530078. URL https://doi.org/10.1145/3528223.3530078. H.-X. Guo, Y. Liu, H. Pan, and B. Guo. Implicit Conversion of Manifold B-Rep Solids by Neural Halfspace Representation. ACM Transactions on Graphics, 41(6):276, 2022b. doi: 10.1145/35 50454.3555502. URL https://doi.org/10.1145/3550454.3555502. 10
Preprint
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger. Deep Reinforcement Learning That Matters. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 2018. doi: 10.1609/aaai.v32i1.11694. URL https://doi.org/10.1609/aaai.v32i1 .11694. D. G. Horvitz and D. J. Thompson. A Generalization of Sampling Without Replacement from a Finite Universe. Journal of the American Statistical Association, 47(260):663–685, 1952. doi: 10.1080/01621459.1952.10483446. URL https://doi.org/10.1080/01621459.195 2.10483446. J. Huang, Y. Zhang, and M. Sun. PrimitiveNet: Primitive Instance Segmentation with Local Primitive Embedding under Adversarial Metric. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. doi: 10.1109/ICCV48922.2021.01506. URL https://doi.org/10.1109/ICCV48922.2021.01506. V. Ishimtsev, A. Bokhovkin, A. Artemov, S. Ignatyev, M. Niessner, D. Zorin, and E. Burnaev. CADDeform: Deformable Fitting of CAD Models to 3D Scans. In Computer Vision – ECCV 2020, 2020. doi: 10.1007/978-3-030-58601-0_36. URL https://doi.org/10.1007/978-3 -030-58601-0_36. P. K. Jayaraman, A. Sanghi, J. G. Lambourne, K. D. D. Willis, T. Davies, H. Shayani, and N. Morris. UV-Net: Learning from Boundary Representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. doi: 10.1109/CVPR46437.2021.0 1153. URL https://doi.org/10.1109/CVPR46437.2021.01153. B. Jones, D. Hildreth, D. Chen, I. Baran, V. G. Kim, and A. Schulz. AutoMate: A Dataset and Learning Approach for Automatic Mating of CAD Assemblies. ACM Transactions on Graphics, 40(6):227, 2021. doi: 10.1145/3478513.3480562. URL https://doi.org/10.1145/34 78513.3480562. B. T. Jones, M. Hu, M. Kodnongbua, V. G. Kim, and A. Schulz. Self-Supervised Representation Learning for CAD. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. doi: 10.1109/CVPR52729.2023.02043. URL https://doi.or g/10.1109/CVPR52729.2023.02043. R. K. Jones, T. Barton, X. Xu, K. Wang, E. Jiang, P. Guerrero, N. J. Mitra, and D. Ritchie. ShapeAssembly: Learning to Generate Programs for 3D Shape Structure Synthesis. ACM Transactions on Graphics, 39(6):234, 2020. doi: 10.1145/3414685.3417812. URL https: //doi.org/10.1145/3414685.3417812. M. S. Khan, S. Sinha, T. U. Sheikh, D. Stricker, S. A. Ali, and M. Z. Afzal. Text2CAD: Generating Sequential CAD Designs from Beginner-to-Expert Level Text Prompts. Advances in Neural Information Processing Systems 37 (NeurIPS 2024), 2024. doi: 10.52202/079017-0242. URL https://doi.org/10.52202/079017-0242. S. Koch, A. Matveev, Z. Jiang, F. Williams, A. Artemov, E. Burnaev, M. Alexa, D. Zorin, and D. Panozzo. ABC: A Big CAD Model Dataset for Geometric Deep Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. doi: 10.1109/CVPR.2019.00983. URL https://doi.org/10.1109/CVPR.2019.00983. C. Li, H. Pan, A. Bousseau, and N. J. Mitra. Sketch2CAD: Sequential CAD Modeling by Sketching in Context. ACM Transactions on Graphics, 39(6):164, 2020. doi: 10.1145/3414685.3417807. URL https://doi.org/10.1145/3414685.3417807. C. Li, H. Pan, A. Bousseau, and N. J. Mitra. Free2CAD: Parsing Freehand Drawings into CAD Commands. ACM Transactions on Graphics, 41(4):93, 2022. doi: 10.1145/3528223.3530133. URL https://doi.org/10.1145/3528223.3530133. L. Li, M. Sung, A. Dubrovina, L. Yi, and L. J. Guibas. Supervised Fitting of Geometric Primitives to 3D Point Clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. doi: 10.1109/CVPR.2019.00276. URL https://doi.org/10 .1109/CVPR.2019.00276. 11
Preprint
P. Li, J. Guo, X. Zhang, and D.-M. Yan. SECAD-Net: Self-Supervised CAD Reconstruction by Learning Sketch-Extrude Operations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. doi: 10.1109/CVPR52729.2023.01613. URL https://doi.org/10.1109/CVPR52729.2023.01613. K.-Y. Liang and S. L. Zeger. Longitudinal Data Analysis Using Generalized Linear Models. Biometrika, 73(1):13–22, 1986. doi: 10.1093/biomet/73.1.13. URL https://doi.or g/10.1093/biomet/73.1.13. M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT*), 2019. doi: 10.1145/3287560.3287596. URL https: //doi.org/10.1145/3287560.3287596. K. Mo, S. Zhu, A. X. Chang, L. Yi, S. Tripathi, L. J. Guibas, and H. Su. PartNet: A LargeScale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. doi: 10.1109/CVPR.2019.00100. URL https://doi.org/10.1109/CVPR.201 9.00100. J. Neyman. On the Two Different Aspects of the Representative Method: The Method of Stratified Sampling and the Method of Purposive Selection. Journal of the Royal Statistical Society, 97(4): 558–606, 1934. doi: 10.1111/j.2397-2335.1934.tb04184.x. URL https://doi.org/10.1 111/j.2397-2335.1934.tb04184.x. D. Paschalidou, A. O. Ulusoy, and A. Geiger. Superquadrics Revisited: Learning 3D Shape Parsing Beyond Cuboids. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. doi: 10.1109/CVPR.2019.01059. URL https://doi.org/10 .1109/CVPR.2019.01059. D. Paschalidou, A. Katharopoulos, A. Geiger, and S. Fidler. Neural Parts: Learning Expressive 3D Shape Abstractions with Invertible Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. doi: 10.1109/CVPR46437.20 21.00322. URL https://doi.org/10.1109/CVPR46437.2021.00322. N. Reimers and I. Gurevych. Reporting Score Distributions Makes a Difference: Performance Study of LSTM-Networks for Sequence Tagging. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2017. doi: 10.18653/v1/D17-1035. URL https://doi.org/10.18653/v1/D17-1035. D. Rukhovich, E. Dupont, D. Mallis, K. Cherenkova, A. Kacem, and D. Aouada. CAD-Recode: Reverse Engineering CAD Code from Point Clouds. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025. doi: 10.1109/ICCV51701.2025.00914. URL https://doi.org/10.1109/ICCV51701.2025.00914. R. Schnabel, R. Wahl, and R. Klein. Efficient RANSAC for Point-Cloud Shape Detection. Computer Graphics Forum, 26(2):214–226, 2007. doi: 10.1111/j.1467-8659.2007.01016.x. URL https: //doi.org/10.1111/j.1467-8659.2007.01016.x. G. Sharma, R. Goyal, D. Liu, E. Kalogerakis, and S. Maji. CSGNet: Neural Shape Parser for Constructive Solid Geometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. doi: 10.1109/CVPR.2018.00578. URL https: //doi.org/10.1109/CVPR.2018.00578. G. Sharma, D. Liu, S. Maji, E. Kalogerakis, S. Chaudhuri, and R. Měch. ParSeNet: A Parametric Surface Fitting Network for 3D Point Clouds. In Computer Vision – ECCV 2020, 2020. doi: 10.1007/978-3-030-58571-6_16. URL https://doi.org/10.1007/978-3-030-585 71-6_16. C. Sun, Q.-F. Zou, X. Tong, and Y. Liu. Learning Adaptive Hierarchical Cuboid Abstractions of 3D Shape Collections. ACM Transactions on Graphics, 38(6):241, 2019. doi: 10.1145/3355089.33 56529. URL https://doi.org/10.1145/3355089.3356529. 12
Preprint
A. Torralba and A. A. Efros. Unbiased Look at Dataset Bias. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2011. doi: 10.1109/CVPR.2011.5995347. URL https://doi.org/10.1109/CVPR.2011.5995347. S. Tulsiani, H. Su, L. J. Guibas, A. A. Efros, and J. Malik. Learning Shape Abstractions by Assembling Volumetric Primitives. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. doi: 10.1109/CVPR.2017.160. URL https: //doi.org/10.1109/CVPR.2017.160. M. A. Uy, Y.-Y. Chang, M. Sung, P. Goel, J. G. Lambourne, T. Birdal, and L. J. Guibas. Point2Cyl: Reverse Engineering 3D Objects from Point Clouds to Extrusion Cylinders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. doi: 10.1109/CVPR52688.2022.01155. URL https://doi.org/10.1109/CVPR52688.20 22.01155. S. Varma and R. Simon. Bias in Error Estimation When Using Cross-Validation for Model Selection. BMC Bioinformatics, 7:91, 2006. doi: 10.1186/1471-2105-7-91. URL https://doi.org/ 10.1186/1471-2105-7-91. L. Wang, H. Meng, Z. Xiang, J. Liu, P. Zhou, L. Chen, and Y. Tang. Text2CAD-Bench: A Benchmark for LLM-Based Text-to-Parametric CAD Generation. arXiv preprint arXiv:2605.18430, 2026. doi: 10.48550/arXiv.2605.18430. URL https://doi.org/10.48550/arXiv.2 605.18430. K. D. D. Willis, Y. Pu, J. Luo, H. Chu, T. Du, J. G. Lambourne, A. Solar-Lezama, and W. Matusik. Fusion 360 Gallery: A Dataset and Environment for Programmatic CAD Construction from Human Design Sequences. ACM Transactions on Graphics, 40(4):54, 2021. doi: 10.1145/3450626.3459818. URL https://doi.org/10.1145/3450626.3459818. K. D. D. Willis, P. K. Jayaraman, H. Chu, Y. Tian, Y. Li, D. Grandi, A. Sanghi, L. Tran, J. G. Lambourne, A. Solar-Lezama, and W. Matusik. JoinABLe: Learning Bottom-Up Assembly of Parametric CAD Joints. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. doi: 10.1109/CVPR52688.2022.01539. URL https: //doi.org/10.1109/CVPR52688.2022.01539. R. Wu, C. Xiao, and C. Zheng. DeepCAD: A Deep Generative Network for Computer-Aided Design Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. doi: 10.1109/ICCV48922.2021.00670. URL https://doi.org/10.1109/ICCV48 922.2021.00670. X. Xu, J. G. Lambourne, P. K. Jayaraman, Z. Wang, K. D. D. Willis, and Y. Furukawa. BrepGen: A B-Rep Generative Diffusion Model with Structured Latent Geometry. ACM Transactions on Graphics, 43(4):119, 2024. doi: 10.1145/3658129. URL https://doi.org/10.1145/36 58129. S. Yan, Z. Yang, C. Ma, H. Huang, E. Vouga, and Q. Huang. HPNet: Deep Primitive Segmentation Using Hybrid Representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. doi: 10.1109/ICCV48922.2021.00275. URL https://doi. org/10.1109/ICCV48922.2021.00275. F. Yu, Z. Chen, M. Li, A. Sanghi, H. Shayani, A. Mahdavi-Amiri, and H. Zhang. CAPRI-Net: Learning Compact CAD Shapes with Adaptive Primitive Assembly. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. doi: 10.1109/CVPR52 688.2022.01147. URL https://doi.org/10.1109/CVPR52688.2022.01147. H. Zhang, K. Liu, M. Chen, L. Li, S. Yang, C. Peng, and H. Chen. BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD. arXiv preprint arXiv:2605.10865, 2026. doi: 10.48550/arXiv.2605.10865. URL https://doi.org/10.48550/arXiv.2605.10 865.
13
Preprint
A
S TATISTICAL D ERIVATIONS AND C OST A LLOCATION
This appendix corresponds to Sections 3.3–3.5. The main text retains only the equations needed to understand the method; here we provide the design-based derivation, variance-component estimators, continuous approximation, and calibration break-even condition. A.1
D ESIGN -BASED D ERIVATION OF THE T HREE -L EVEL VARIANCE
Let Ytie ∈ {0, 1} denote the failure label for template t, independent generation i, and frozen edit state e. Each program has M pre-frozen states, from which k are sampled without replacement; each template has g independent generations, and a formal evaluation contains q templates. Program risk is the mean over the M states, and the model-risk estimator is the average of the sampled template, generation, and state means. For a fixed program (t,i), the conditional variance under simple random sampling without replacement is 2 Var(Ȳti (k) | t, i) = (1/k − 1/M )SE,ti .
(A.1)
Averaging g independent generations within the same template reduces the generation- and edit-level contributions by 1/g; averaging over q templates then gives the superpopulation design variance: b = σT2 /q + σI2 /(qg) + Var(R)
1 2 (1/k − 1/M )σE . qg
(A.2)
Here, σT2 denotes upper-level heterogeneity among template means, σI2 the generation heterogeneity 2 the mean finite-population variance among among full-program risks within a template, and σE frozen states within a program. The decomposition directly shows that increasing k affects only the third term; uncertainty at the template and generation levels can be reduced only by increasing q or g. When the validation target is a frozen reference pool containing Q templates and G complete generations per template, templates and generations are also sampled without replacement, yielding the finite-population form: 2 b = (1/q − 1/Q)σT2 + 1 (1/g − 1/G)σI2 + 1 (1/k − 1/M )σE VarF (R) . q qg
(A.3)
Equation (5) is used only for method validation on the frozen test pool. Intervals for the future task distribution still use templates as the outermost cluster and do not treat the finite test-template set as the complete task population. A.2
M ETHOD - OF -M OMENTS E STIMATION OF VARIANCE C OMPONENTS
Full calibration records contain all M=16 states for every program. Estimation follows the threelevel nested structure and explicitly subtracts lower-level measurement noise so that edit-sampling variation is not misattributed to generation heterogeneity, nor generation variation to template heterogeneity. Negative method-of-moments residuals can arise only from finite-sample noise. The executable protocol truncates them at zero, ensuring that all three variance components entering budget optimization are finite and nonnegative. Finite-population estimates on the frozen reference pool instead use complete template means, complete program means, and within-program state variances directly. A.3
C OST C ONSTRAINT, C ONTINUOUS A PPROXIMATION , AND I NTEGER S EARCH
Let cT , cI , and cE denote the costs of preparing a new template, generating/parsing/initially building one program, and executing/judging one edit, respectively. The formal cost of a balanced design is 14
Preprint
Component 2 σE
σI2 σT2
Table 3: Estimation order for the three variance components Calibration quantity Implementation detail Sample variance over the 16 states within each program, averaged across programs Variance of complete-program risk within a template Variance of complete template means
Preserves behavioral differences within a program If partial states are used, first subtract the 2 (1/k−1/M)SE measurement term 2 Subtract σI /g from finite generations per template; truncate negative residuals to 0
C(q, g, k) = q[cT + g(cI + kcE )].
(A.4)
Integer search is the final allocation rule used in this paper. It enumerates feasible (g,k), purchases the largest q allowed by the budget for each pair, and selects the minimum finite-population design variance. This naturally handles template saturation, finite-population corrections, and unused integer budget. for g = 1,. . . ,G: for k = 1,. . . ,M: d = cT + g(cI + k cE ) q = min(Q, floor(C0 / d)) b if q ≥ 1: evaluate VarF (R) return feasible (q,g,k) with minimum variance The continuous relaxation is used only to explain when increasing audit depth is worthwhile. Holding g fixed and defining d(k)=cT +g(cI +k cE ), substitution of q≈C0 /d(k) gives 2 d(k) σE V (k | g) ≈ , Hg + C0 gk σ2 σ2 Hg = σT2 + I − E . g gM
(A.5)
When Hg >0, the interior stationary point satisfies s k∗ =
2 (c + gc ) σE T I . g 2 cE Hg
(A.6)
2 This expression is used only for mechanism interpretation. Larger within-program variance σE or larger fixed generation cost favors deeper auditing, while larger template/generation variance or larger per-edit cost shifts budget toward new templates or generations. All reported results use the integer search.
A.4
C ALIBRATION C OST AND B REAK -E VEN C ONDITION
Let Jglobal and Jpilot denote the cost-scaled MSE of the shared rule and the system-level calibrated rule under formal budget C0 , and let Cpilot be the additional cost of one calibration. If the same calibration can be reused across multiple formal evaluations, the continuous approximation to the break-even number of evaluations is Nbreak ≈
Cpilot , C0 [1 − Jpilot /Jglobal ] 15
Jpilot < Jglobal .
(A.7)
Preprint
If Jpilot ≥Jglobal , no positive break-even count exists. This approximation is used only to assess whether separate calibration for one system is worthwhile. We therefore report calibration cost separately from the formal evaluation budget and treat full calibration as an analytical reference rather than a default deployment step.
B
TASK FAMILIES AND C ONSTRUCTION OF F ROZEN E DIT S TATES
This appendix corresponds to Sections 4.1–4.2. State construction reads only the task specification; candidate-program implementation, model identity, and runtime outcomes never enter the stategeneration process. B.1
F OURTEEN TASK FAMILIES AND T HEIR E DIT S EMANTICS Table 4: Task families, linked parameters, and high-level semantic contracts
Env.
Task families
Primary
Linked
A
mounting_b racket flange_pla te stepped_sh aft electronic s_enclosure
length
length + width
A
belt_pulley
outer diameter
A
ribbed_ang le bolt_patte rn_plate
length
A
pipe_clamp
outer diameter
B
gear_blank
outer diameter
B
hinge
leaf length
B
drawer_han dle bottle_cap
span
lattice_pa nel bearing_ho using
length
A A A
A
B B B
outer diameter shaft length length
length
outer diameter
outer diameter
Semantic contract
hole spacing; symmetric hole array; upright connected to base outer diameter + bolt circle; equally spaced holes; bolt circle central bore concentric shaft + shoulder shoulder length; wider shoulder; diameter concentric axial bore length + width wall; interior clearance follows the outer profile; preserve a single shell outer + hub hub length; concentric hub/rim; diameter through bore width + height rib count; uniformly distributed ribs; base/web connected spacing x + hole diameter; four-hole biaxial spacing y symmetry; holes remain within the plate outer + inner lug length; concentric pipe bore; diameter mirrored dual lugs outer + root teeth; radially uniform teeth; diameter concentric bore leaf width + barrel knuckle count; cover the hinge diameter axis; continuous pin bore span + mount grip height; symmetric mounts; spacing grip spans the two bosses outer diameter + rib count; cap remains hollow; height uniformly spaced exterior ribs length + width bar count; closed frame on all four sides; uniformly spaced bars base length + base bore diameter; concentric width bore/seat; connected base; four-hole symmetry
Environment A contains eight task families with nine variants each; environment B contains six disjoint task families with eight variants each. Frozen rules vary multiple dimensional ratios across variants, and variants are not allowed to degenerate into uniformly scaled copies of the same geometry. B.2
C ONSTRUCTION AND ACCEPTANCE C RITERIA FOR THE F OUR S TATE T YPES
16
Preprint
Table 5: Construction contract for frozen edit states Type
Construction target Acceptance criteria
At least one parameter changes; full vector is valid; normalized distance ≥0.10 Activates an engineering constraint; normalized constraint margin ≤0.05 At least two linked parameters linked Predefined change simultaneously; related-parameter complete dependency closure group semantic High-level semantic dependency closure ≥2; semantic contract exists; parameters and geometric signature must contract respond local
Routine single-parameter edit boundary Near a true feasible boundary
Primary stress Parameter binding and local geometric response Thin walls, spacing, containment, and degenerate geometry Dimensional relations, symmetry, and multi-parameter consistency Design semantics such as array count, shell structure, assembly, and concentricity
Each template contains exactly four states of each of the four types, for 16 complete parameter vectors. Construction rejects no-ops, out-of-range states, and duplicate vectors. Each state records units, valid ranges, changed parameters, expected relations, and tolerances. Formal evaluation samples only from these 16 frozen states without replacement and never adapts the state set after seeing the candidate program. B.3
S TATE -Q UALITY AND S PLIT-C ONSISTENCY C HECKS Table 6: Executable challenge-quality checks
Check
Criterion
Uniqueness and validity Edit magnitude Boundary stress
Each template has 16 unique parameter vectors; all parameters lie within frozen valid ranges Each state has normalized distance ≥0.10 from nominal Must target a true active constraint, with normalized engineering margin ≤0.05 Each semantic state links at least two parameters and carries a nonempty semantic contract Variants within a task family cannot all be uniform-scale copies Calibration/test are deterministically split within each family; calibration cannot be a simple variant-number prefix
Semantic closure Variant diversity Split integrity
These checks cover 120 templates and 1,920 frozen states. They ensure that deeper auditing adds evidence from the same target population rather than raising pass rates by lowering task difficulty or rewriting states for candidate programs.
C
I SOLATED E XECUTION AND AUTOMATED J UDGE
This appendix corresponds to Section 4.2. Candidate programs must expose build(params) and return a measurable CadQuery workplane or shape. The nominal state and every frozen state are executed in a fresh subprocess. C.1
I SOLATED W ORKER AND G EOMETRY P ROBES
17
Preprint
Table 7: Geometry probes produced by the isolated worker Category
Recorded fields
process
return code; timeout; wall-clock; failure attribution stdout/stderr; failure stage volume; surface area; is_valid; validity and topology solid/face/edge/vertex counts bounding-box bounds/dimensions; dimensions; containment; symmetry center of mass face types; circular-edge holes; axes; concentricity center/radius/axis hash of sorted face-geometry summaries edit-response detection
shape spatial features signature
Use
Infrastructure-level worker errors are retried at most once. Timeouts, model-code errors, and geometric build failures do not trigger model-level retries. If the nominal program is invalid because of model, code, or geometry failure, all 16 counterfactual states count as failures in the primary risk. Infrastructure errors retain a separate failure stage and are never rewritten as model failures. C.2
S IX -S TAGE AUTOMATED -J UDGMENT C ONTRACT Table 8: Ordered checks performed by the automated judge
Stage
Inputs
1 Build
process status; timeout
Rule
build finishes within limit; measurable shape returned 2 Entity is_valid; solid/component counts valid entities; required component count 3 Topology face/edge/circle probes feature counts, radii and locations match task contract 4 Response nominal/edit probes; geometry edited parameters cause the expected detectable signature response 5 Geometry bbox; centers; spacing; margins dimensions, symmetry, containment and clearance pass 6 Semantic expected relations; dependency high-level semantic contract remains satisfied closure
Length tolerance is max(0.05 mm, 0.001·Lref ), angular tolerance is 0.1◦ , and relative volume tolerance is 0.5%; discrete topology counts use exact rules. A 0.90T–1.10T gray zone is frozen around numerical thresholds. Gray-zone records are routed to manual review and excluded from analyses requiring a binary outcome until a review label is available. C.3
AUDIT-E XECUTION P SEUDOCODE run nominal build(params_nominal) in a fresh subprocess if infrastructure failure: emit infrastructure stage; do not relabel as model failure if nominal model/code/geometry failure: assign failure to all 16 frozen states else: cache nominal geometry probes for each selected frozen state: execute build(params_state) in a fresh subprocess extract probes; apply ordered judge stages 1. . . 6 emit state_id, edit_type, label, failure_stage, timing and diagnostics 18
Preprint
D
S YSTEM M APPING , RUNTIME E NVIRONMENT, AND F ROZEN P ROTOCOL
This appendix corresponds to Sections 4.1 and 4.3 and summarizes only the fixed configurations required to reproduce the experiments. External repository identifiers unrelated to the method or results are omitted from the anonymous manuscript. D.1
G ENERATION S YSTEMS Table 9: Five generation systems
System Underlying model and fixed Generation method version
Output interface
S1
GPT-4.1 mini
CadQuery program
S2 S3 S4 S5
GPT-5.5 Gemini 2.5 Flash Gemini 3.1 Pro Claude Sonnet 4.6
D.2
Independent generation under a common prompt Same as above Same as above Same as above Same as above
Same as above Same as above Same as above Same as above
E NVIRONMENTS AND DATA M ATRIX Table 10: Complete pools for the two frozen environments
Item
Environment A
Environment B
Task families / templates calibration / test Complete generations/template Frozen states/program state-level audit records
8 / 72 24 / 48 5
6 / 48 12 / 36 4
16 (4×4 categories) 28,800
16 (4×4 categories) 15,360
D.3
E XECUTION , C OST, AND R ANDOMNESS P ROTOCOL Table 11: Frozen execution protocol
Protocol item
Fixed value
Master seed Formal budget Template cost / edit cost A generation cost, S1–S5 B generation cost, S1–S5 Software environment Resources timeout Retries Fixed audit depths Preregistered benefit comparisons
20270901 C0 =1024; sensitivity budgets 512 / 1024 / 2048 2 / 1 (normalized to one edit execution) 1 / 3 / 8 / 16 / 32 2 / 6 / 16 / 32 / 64 Ubuntu 22.04; Python 3.11; CadQuery 2.5.x; OCCT 7.8.x 16 vCPU; 64 GB RAM; concurrency=1 generation 180 s; nominal build 60 s; edit 30 s infrastructure 1; model/code/geometry 0 1, 4, 8, 9, 10, 12, 16 4→8 and 8→16
Per-run wall-clock time is the raw cost measure. The main estimate uses a 10% trimmed mean, with the ordinary mean and median used for sensitivity checks. Calibration/test splitting always treats the template as the outermost unit; all generations and state records from the same template remain in the same split. 19
Preprint
E
E XPERT VALIDATION
This appendix corresponds to Sections 4.2 and 5.3. The expert sample is stratified by system × edit type × automatic label; each nonempty stratum contributes 20 records or all available records, for 800 total. Two experts with parametric-CAD experience independently inspect the task specification, execution log, and edited geometry without knowing system identity or automated label, and then form consensus through a predefined adjudication process. E.1
W EIGHTED M ETRICS AND E XPERT AGREEMENT
If stratum h contains Nh records in the target audit population and nh expert-sampled records, each record receives inverse-sampling-probability weight wh =Nh /nh . Expert-reference risk and automated-judgment risk are both estimated on the same expert sample using these weights; precision, recall, specificity, and accuracy use the same weighting. The two risks and weighted confusion metrics in Table 12 therefore share the same target population and weighting scheme. Cohen's κ between the two experts is computed before consensus; the 800 doubly labeled records yield κ≈0.797. Table 12: Expert-reference risk and weighted confusion metrics for the automated judge Sys.
Expert risk
Weighted auto. risk
Prec.
Recall
Spec.
Acc.
S1 S2 S3 S4 S5
0.0909 0.2151 0.3261 0.4060 0.4825
0.0957 0.2151 0.3188 0.3863 0.4460
0.950 1.000 1.000 1.000 1.000
1.000 1.000 0.977 0.952 0.924
0.995 1.000 1.000 1.000 1.000
0.995 1.000 0.993 0.980 0.963
The main bias of the automated judge comes from missed failures rather than additional false positives: precision remains 1.000 for S3–S5, while recall decreases as system difficulty increases. This supports using automated-judgment risk in the main text as a reviewable model-level estimate while reporting expert-reference risk alongside it.
F
S UPPLEMENTARY E XPERIMENTS , ROBUSTNESS , AND M ODEL C OMPARISON
This appendix corresponds to Sections 4.4–5.3. It contains cost, stability, and model-comparison results omitted from the main text for space, together with exact code-level definitions of each strategy. Exact values for the main-text result figures are collected in Appendix F.6. F.1
C ALIBRATION C OST Table 13: Calibration cost (formal budget C0 =1024)
System
Full-calibration cost
Small-calibration cost
Small calibration / formal budget
S1 S2 S3 S4 S5
2088 2328 2928 3888 5808
288 320 400 528 784
28.1% 31.3% 39.1% 51.6% 76.6%
Full calibration serves as a mechanism-analysis reference. Small calibration is used to study whether system-level configurations can amortize their additional cost when calibration is reused across multiple leaderboard rounds or model ablations. 20
Preprint
F.2
S PLIT S TABILITY AND TASK -FAMILY E XTRAPOLATION Table 14: Random template splits and leave-one-task-family-out validation
System Random-split 5% Random regret coverage median/P90
LOFO 5% coverage
LOFO regret median/P90
S1 S2 S3 S4 S5
25.0% 12.5% 75.0% 37.5% 62.5%
1.286 / 1.709 1.390 / 1.608 1.000 / 1.215 1.083 / 1.210 1.029 / 1.471
45.0% 65.0% 62.5% 57.5% 72.5%
1.059 / 1.444 1.028 / 1.374 1.042 / 1.161 1.038 / 1.347 1.033 / 1.229
Random-split stability is evaluated with 40 family-stratified splits in environment A: three templates from each task family are used for calibration and the remaining six for held-out evaluation, giving 200 system × split outcomes. LOFO calibrates on seven complete task families and evaluates on the held-out family, so it measures change in task composition rather than ordinary random-split noise. F.3
H IERARCHICAL I NTERVAL E XPERIMENT Table 15: The sole difference between the interval-coverage methods
Method
outer unit
Variance structure
Critical value / target
Three-level design interval
template
tq−1 ; covers mean risk of the full frozen test pool
programindependent Control
generation program
template + generation + edit; finite-population corrections at each level Treats programs as independent and does not explicitly retain the template level
program-level df; uses the same point estimates and samples as the three-level interval
The interval experiment fixes q=16, g=3, and k=8 and performs 5,000 replays without replacement for each system. Both methods use exactly the same sampled records and risk point estimates; coverage differs only in whether template-level correlation is retained. Section 5.1 reports the coverage results, so the same numbers are not repeated here. F.4
F ORMAL S ELECTION C ONTRACT FOR A LLOCATION S TRATEGIES
Fixed-depth and system-level joint search share the same variance estimates, cost model, budget constraint, and minimization objective. Fixed-depth search constrains k to a prespecified depth, while system-level search jointly searches g and k. Therefore, when system-level search selects k = k0 , its (q, g) must match the optimal feasible configuration under fixed k0 . The edit term under four-way stratification uses the finite stratified variance: VarE,strat =
1 X 2 Wh qg h
1 1 − kh Mh
2 σE,h ,
Wh =
1 , Mh = 4. 4
(F.1)
A B For the paired strategy, Ztie = Ytie − Ytie is defined on shared keys, and the same three-level variance and cost constraint are applied to Z. The oracle enumerates feasible configurations only post hoc on held-out data; relative regret=Jselected /Joracle , and regret≤1.05 defines membership in the 5% robust region. The oracle never participates in configuration selection. For cross-environment transfer, (g,k) calibrated in environment A is frozen and only qs is remapped using measured costs in environment B.
21
Preprint
Strategy
Table 16: Information boundaries of the allocation strategies Calibration Frozen/selected quantities Test stage information
fixed depth
Target-system calibration pooled Multiple calibration systems LOSO Calibration systems pooled excluding target system-level Full target-system calibration editFour state types from stratified target/pooled calibration paired Calibration keys shared by the model pair
F.5
Fix k; choose g on calibration Jointly select (g,k) Jointly select (g,k) Select target-system (g,k) kh ≥1 per type Estimate three-level variance using Z = YA − Y B
Purchase the maximum q under target-system cost Map qs separately from each system's cost Simulate a new system; target provides only cost Analytical reference, not a default deployment Equal-weight risk over four types; joint integer search Compare on the same template/generation/state keys
PAIRED M ODEL D ECISIONS Table 17: Mean correct-decision rate across ten model pairs
Method
C0 =512
C0 =1024
C0 =2048
Fixed k=8 Fixed k=9 Fixed k=10 Paired system-level allocation Paired system-level + four-way stratification
95.8% 95.8% 95.0% 95.2% 95.9%
97.9% 97.7% 98.0% 98.2% 98.4%
99.7% 99.6% 99.7% 99.6% 99.6%
The gain from paired allocation is at most sub-percentage-point under low and medium budgets and largely vanishes at high budget. We therefore position it in the main text as a focused review tool for closely matched systems rather than the default configuration for a standard leaderboard. F.6
E XACT DATA FOR M AIN -T EXT R ESULT F IGURES
This section reports the complete values, sampling configurations, and comparison conventions underlying the three result figures in the main text for ease of lookup and verification. Table 18: Estimation efficiency in the primary environment (corresponding to Figure 1) Sys.
k=4: q/g; J
k=8: q/g; J
k=16: q/g; J
Program-level MSE: k=4→8→16
S1 S2 S3 S4 S5
46/4; 0.0974 44/3; 0.3050 39/2; 0.7178 46/1; 1.2881 7/4; 2.4026
48/2; 0.1126 42/2; 0.3097 48/1; 0.6533 39/1; 0.7728 24/1; 1.3358
48/1; 0.1649 48/1; 0.3014 39/1; 0.5080 30/1; 0.5378 20/1; 0.8429
0.014284→0.004761→0 0.026178→0.008726→0 0.038301→0.012767→0 0.043678→0.014559→0 0.045749→0.015250→0
Note: The formal budget is C0 =1024; J is model-level cost-scaled mean squared error. See Figure 1.
22
Preprint
Table 19: Candidate configurations and estimation error under the cost constraint (corresponding to Figure 2) Sys.
Pooled configuration q/g/k
J
System-level candidate q/g/k
J
Stratified candidate q/g/k
J
S1 S2 S3 S4 S5
48/1/12 48/1/12 46/1/12 34/1/12 22/1/12
0.1987 0.3634 0.4996 0.6089 0.9935
48/3/5 48/1/16 48/1/11 34/1/12 22/1/12
0.1020 0.3014 0.5048 0.6089 0.9935
48/3/5 48/1/16 48/1/11 35/1/11 22/1/12
0.0899 0.3014 0.4708 0.5931 0.9480
Note: J=1024×MSE; the formal budget is 1024, with calibration cost reported separately. All configurations use the largest feasible q for the specified g and k. The stratified configuration samples at least one state from each category, estimates risk with equal weights across the four categories, and computes J using the corresponding stratified variance. Fixed-depth and system-level joint search use the same variance estimates, cost model, formal budget, and selection objective; fixed-depth search additionally constrains k to the specified value. The table retains absolute J and q/g/k configurations; Figure 2 plots each J in this table divided by the Table 18 value for the same system at fixed k=16. Table 20: Estimation efficiency in the cross-task-family transfer environment (corresponding to Figure 3) Method
Mean C0 ×MSE
Fixed k=4 0.992 Fixed k=8 0.787 Fixed k=9 0.771 Fixed k=10 0.742 Fixed k=12 0.757 Fixed k=16 0.795 Direct transfer of pooled configuration from 0.756 primary environment Leave-one-system-out pooled 0.790 Full-calibration system-level strategy 0.701 Small calibration + four-way stratification 0.667
Mean k
Relative to best in table
4.0 8.0 9.0 10.0 12.0 16.0 12.0
+48.8% +18.0% +15.6% +11.4% +13.5% +19.3% +13.4%
11.6 11.4 9.0
+18.5% +5.1% 0.0%
Note: The formal budget is C0 =1024, with calibration cost reported separately. “Relative to best in table” uses the lowest mean J in this table as the reference. Figure 3 plots the mean J values from this table.
G
A NONYMOUS E XECUTABLE M ATERIALS AND C ONSISTENCY C HECKS
This appendix lists only the relative paths, analysis entry points, and validation contracts for the anonymous executable materials accompanying the paper. It contains no external repository name, username, or URL. G.1
M AIN A RTIFACT M AP
23
Preprint
Table 21: Executable components in the anonymous supplementary materials Path / module
Role
data/templates/depthbenchc ad_tasks.json data/programs/ data/records/*_audits.jsonl data/generations/*.jsonl data/records/*expert_annota tions.jsonl configs/paper_protocol.json
120 templates, valid ranges, split metadata, and 1,920 frozen states CadQuery reference programs for the templates A/B state-level audit outcomes generation-level summaries and 16-state linkage 800 dual-expert validation records
seed, budget, cost, timeout, split, depth, replay, expert protocol src/depthbenchcad/audit.py / state construction; isolated execution and probes; executor.py / judge.py six-stage judgment src/depthbenchcad/variance three-level variance/FPC; feasible integer allocation .py / allocation.py src/depthbenchcad/replay.py nested replay; pooled/LOSO/system/stratified/paired / strategies.py src/depthbenchcad/paper_an statistical-analysis layer for the main text and appendices alysis.py scripts/reproduce_tables.py recompute paper analyses from record-level inputs scripts/validate_release.py release linkage and checks of paper values/mechanism / validate_paper_results.py invariants tests/ task/state, judge, statistics, strategy, and release tests
G.2
M INIMAL R EPRODUCTION E NTRY P OINT
python -m pip install -r requirements.txt python scripts/validate_release.py pytest -q python scripts/reproduce_tables.py \ --records-a data/records/depthbenchcad_A_audits.jsonl \ --records-b data/records/depthbenchcad_B_audits.jsonl \ --expert data/records/depthbenchcad_A_expert_annotations.jsonl \ --out-dir results/reproduced python scripts/validate_paper_results.py --results results/reproduced
G.3
AUTOMATED C ONSISTENCY C HECKS
Check layer
Table 22: Automated checks for the release and statistical analysis Frozen invariant
corpus state population challenge quality record linkage generation summary statistics expert sample paper analysis
120 templates; A/B=72/48; 14 families; split=24/48 and 12/36 16 states/template, 4 per category; unique, valid, non-no-op; distance≥0.10 boundary activates a true constraint; semantic closure≥2; family variants are not uniform-scale copies A/B=28,800/15,360 state records; each generation links to exactly 16 outcomes failure_count, failed_state_ids, and edit-type breakdown match state records exactly variance components are nonnegative; finite-pool variance does not exceed the corresponding superpopulation design variance; integer allocation is cost-feasible 800 unique system-template-generation-state keys key table values, benefit directions, regret/coverage, and frozen mechanism invariants pass an independent validation script
24
Preprint
These checks do not replace the statistical argument in the paper. Their purpose is to ensure that task definitions, the state population, record hierarchy, and analysis inputs contain no silent mismatches, so the reported results can be recomputed from the frozen record pool under the same protocol.
H
E XTENDED R ELATED W ORK
H.1
E XECUTABLE CAD G ENERATION AND E DIT E VALUATION
CAD generation research has expanded from final geometry to structured design histories and executable programs. DeepCAD represents CAD modeling as an operation sequence (Wu et al., 2021). Fusion 360 Gallery provides real design histories and programmatic construction data (Willis et al., 2021). Text2CAD maps natural language to parametric CAD sequences (Khan et al., 2024). CADRecode further uses language models to recover executable CAD code from point clouds (Rukhovich et al., 2025). These representations retain parameters, operation order, and construction logic, enabling generated outputs to remain executable and editable. Work toward editable programs spans interactive modeling, program generation, and reverse engineering. Sketch2CAD converts contextual sketches into sequential CAD modeling operations (Li et al., 2020), while Free2CAD parses freehand drawings into CAD commands (Li et al., 2022). ShapeAssembly uses executable programs to describe editable 3D structures (Jones et al., 2020). SECAD-Net recovers sketch-extrude operations from geometry (Li et al., 2023). ComplexGen performs CAD reconstruction over B-Rep chain complexes (Guo et al., 2022a), while neural halfspace representations learn implicit conversions of manifold B-Rep solids (Guo et al., 2022b). BrepGen uses diffusion models to generate B-Reps with structured latent geometry (Xu et al., 2024). Related work also obtains compact, interpretable shape representations through primitives, convex components, or constructive solid geometry. CSGNet learns to parse constructive solid geometry programs (Sharma et al., 2018). BSP-Net generates compact meshes through binary space partitioning (Chen et al., 2020), while CvxNet learns convex decomposition (Deng et al., 2020). CAPRI-Net studies adaptive primitive assembly for compact CAD shapes (Yu et al., 2022). Volumetric primitives provide compositional shape abstractions (Tulsiani et al., 2017), and superquadric representations extend shape parsing beyond cuboids (Paschalidou et al., 2019). Neural Parts uses invertible neural networks to learn more expressive 3D part abstractions (Paschalidou et al., 2021), while hierarchical cuboid representations support adaptive 3D shape abstraction (Sun et al., 2019). CAD datasets and representation learning have likewise shifted from generic 3D geometry toward native engineering structure. ABC provides a large CAD dataset with explicit parametric surfaces and curves (Koch et al., 2019). UV-Net learns directly from B-Rep geometry and topology (Jayaraman et al., 2021), and later work explores self-supervised representation learning for CAD (Jones et al., 2023). AutoMate targets automatic mating relations in CAD assemblies (Jones et al., 2021), while JoinABLe learns bottom-up assembly of parametric CAD joints (Willis et al., 2022). Scan2CAD studies alignment between RGB-D scans and CAD models (Avetisyan et al., 2019), and CAD-Deform adapts CAD models to 3D scans through deformable fitting (Ishimtsev et al., 2020). Together, these studies show that CAD evaluation must consider geometry, topology, structural relations, and executability. In reverse engineering and local geometric understanding, primitive fitting and structural segmentation provide another geometric foundation before execution. SPFN performs supervised fitting of parametric geometric primitives to point clouds (Li et al., 2019). ParSeNet fits parametric surfaces to 3D point clouds (Sharma et al., 2020). PrimitiveNet studies primitive instance segmentation with local primitive embeddings (Huang et al., 2021), while HPNet uses hybrid representations for primitive segmentation (Yan et al., 2021). Point2Cyl reconstructs editable extrusion cylinders from point clouds (Uy et al., 2022). PartNet provides a fine-grained hierarchical benchmark for part-level 3D understanding (Mo et al., 2019). Before learning-based methods, Efficient RANSAC was widely used for geometric shape detection in point clouds (Schnabel et al., 2007). Behavioral evaluation also depends on the validity of the automated judge. For labels that require semantic or expert judgment, automated metrics alone are insufficient to establish judge reliability. Cohen's kappa is a classical measure of agreement for nominal labels (Cohen, 1960). A survey in computational linguistics further emphasizes explicit reporting of annotation protocols, inter-coder 25
Preprint
agreement, and the limits of their interpretation (Artstein & Poesio, 2008). We therefore combine automated execution logs with CAD expert review and report precision, recall, expert agreement, and representative false positives and false negatives by edit type. H.2
B UDGET-C ONSTRAINED E VALUATION AND M ULTISTAGE S AMPLING
Randomness, repeated runs, and statistical power have become central concerns in reproducible model evaluation. Deep reinforcement learning experiments show that a single run can obscure substantial stochastic and implementation variation (Henderson et al., 2018). In sequence labeling, reporting score distributions can alter conclusions about model differences (Reimers & Gurevych, 2017). Limited statistical power also makes both positive and negative conclusions harder to interpret under finite experimental budgets (Card et al., 2020). Statistical tests for model comparison must match the sampling and replication design (Dietterich, 1998). Reusing the same crossvalidation process for model selection and error estimation can introduce systematic optimism in the final error estimate (Varma & Simon, 2006). The CAD behavioral evaluation studied here has three levels—task templates, within-template generations, and finite edit states—and three cost types: template preparation, program generation, and edit execution. In stratified sampling, variance and cost at different levels jointly determine sample allocation (Neyman, 1934). Under sampling without replacement, finite populations and inclusion probabilities must enter the estimator explicitly (Horvitz & Thompson, 1952). Repeated observations within the same template and program also exhibit within-cluster correlation and cannot be treated as independent (Liang & Zeger, 1986). We instantiate these statistical ideas as a three-level counterfactual audit of executable CAD programs and further account for execution cost, pilot cost, expert validation, and task-family transfer.
I
E XTENDED D ISCUSSION AND L IMITATIONS
I.1
M AIN F INDINGS AND P RACTICAL I MPLICATIONS
The experiments show that reliable evaluation first requires correctly identifying independent evidence: ignoring template-level correlation produces overconfident model conclusions. Deeper fixed auditing is a strong baseline in the primary environment, while cross-system and cross-environment differences further show that no single audit depth dominates universally. The value of audit depth must be interpreted jointly with the source of variance and execution cost. Cross-task-family results also caution against extrapolating conclusions from benchmark-specific distributions; the literature on dataset bias and shortcut learning shows that an advantage on one benchmark need not transfer to a new data composition (Torralba & Efros, 2011; Geirhos et al., 2020). The main value of the three-level model lies in explaining and diagnosing evidence allocation and predicting the direction of audit benefit, without guaranteeing that joint allocation will outperform a strong fixed-depth baseline in average error. When template variance is large, more templates should be covered; when within-template generation variability is large, more independent generations should be sampled; when edit-state variation is large and program startup cost is high, deeper within-program auditing is more favorable. The model can therefore identify when a system or environment departs from a shared configuration and indicate the appropriate direction of budget adjustment. Accounting for calibration cost further narrows the practical scope of system-level allocation. Full calibration costs more than one formal evaluation, and the break-even condition for small calibration still requires further validation. For a new system without target-system calibration data, the pooled DepthBenchCAD configuration can serve as a reference point that requires no additional calibration, but the current results do not support treating it as a universally superior default relative to fixed depth. System-level calibration is most likely to be worthwhile when execution costs or variance structure differ substantially from previously observed systems and enough subsequent evaluations are expected to amortize the calibration cost. Model-pair-specific calibration is likewise appropriate for focused review of closely matched systems and should not be treated as an implicit cost of a standard leaderboard. 26
Preprint
Leave-one-task-family-out and cross-task-family transfer results show that the preferred depth changes with task composition and its associated variance and execution costs. The shared configuration obtained in the primary environment can serve as a reference in similar settings, but its efficiency should be reported alongside fixed-depth baselines rather than treated as an optimal constant across systems or environments. Appendix G lists the reference configurations, analysis entry points, and consistency checks in the anonymous executable materials. The paired experiments support the same conclusion. Relative to a strong fixed baseline, model-pairspecific allocation yields only limited average improvement and loses its advantage at high budgets. Pairing itself can reduce comparison variance; whether additional calibration is worthwhile depends on model separation, cost shift, and reuse count. I.2
L IMITATIONS
The 16 edit states used here represent only a predefined finite counterfactual population and cannot cover a continuous parameter space or open-ended design interaction. The validity of the estimand depends on the quality of the state generator, valid ranges, and task constraints. Future work could expand audit coverage through constraint-driven state generation, importance sampling, and adaptive stress testing. The three-level variance formula assumes a prespecified sampling design and within-level exchangeability. When task difficulty, failure probability, and execution cost are correlated, or when nonrandom timeouts and adaptive early stopping occur, design weights or unit-level cost optimization are needed. Variance components from small calibration sets may also fluctuate substantially; future work could use REML or Bayesian hierarchical models to propagate uncertainty in parameter estimates. The automated judge reliably detects execution, topology, and explicit geometric errors, but its coverage of design semantics remains limited. Expert review can estimate false positives and false negatives but cannot fully eliminate differences between state definitions and professional judgment. The current validation also covers a limited number of task families, candidate systems, and execution environments; transfer across benchmarks and CAD kernels requires additional real experiments. Clear reporting of evaluation boundaries, failure modes, and conditions of use is therefore important for subsequent reuse (Mitchell et al., 2019). We focus on average failure risk and pairwise risk differences between models. Engineering safety certification also requires weighting critical states, worst-case analysis, and continuous-parameter verification. Unified leaderboards further involve multi-model ranking, uncertainty propagation, and multiple comparisons. These questions lie outside the scope of the present estimand.
J
S UPPLEMENTARY M OTIVATION AND P ROTOCOL E XPLANATION
The following paragraphs retain extended motivation, data checks, and interpretation from the fulllength manuscript. Core definitions and experimental settings remain in the main text. Samples in CAD behavioral evaluation are hierarchical: multiple generations from the same template share task characteristics, and multiple edits within one program share the same implementation. Treating these nested observations as independent evidence can underestimate evaluation uncertainty. Research on stochastic algorithms has shown that a single run can obscure substantial run-to-run variance (Henderson et al., 2018), while reporting score distributions can reduce overinterpretation of a single score (Reimers & Gurevych, 2017). At the same time, under a fixed budget, adding edit checks to each program reduces the number of templates or independent generations that can be covered, creating a tradeoff between thorough program inspection and accurate model evaluation. We study an apparently paradoxical question: why can more thorough auditing of each CAD program make model evaluation less accurate? Under the same budget, checking every edit state for a small number of programs may characterize those programs precisely while still failing to represent the model's behavior on other tasks and generations. Allocating part of the budget to new templates or independent generations can instead yield a more accurate model-level conclusion. The key is that the three sample types answer different questions: templates capture task variation, independent 27
Preprint
generations capture model stochasticity, and edit states capture behavioral variation within a program. Statistical tests for model comparison are highly sensitive to sampling and replication design (Dietterich, 1998). Reusing the same data for configuration selection and final evaluation can also introduce optimistic bias (Varma & Simon, 2006). We therefore build a three-level evaluation model and ask when the information gained by increasing one type of evidence is sufficient to offset the loss of the other two. We explicitly distinguish task templates (T), model generations (I | T ), and edit states (E), derive a three-level estimation variance with finite-population corrections, and incorporate the costs of template preparation, program generation, and edit execution into a unified budget. The final protocol jointly selects the number of templates (q), generations per template (g), and audits per program (k) through integer search. We further test, on held-out templates and new task families, whether three-level variance and cost estimates can predict the practical benefit of deeper auditing. Pilot-study overhead is included to determine when such allocation is worthwhile in real evaluations. Across two CAD environments, we compare program-level and model-level errors at different audit depths and test whether calibration data predict the direction of the benefit from deeper auditing. Program-level error can decrease while model-level error increases, and the direction of the effect across systems is jointly determined by the source of variance and execution cost. We also test how hierarchical treatment affects confidence intervals and report calibration overhead separately. Our analysis builds on classical multistage sampling, with contributions focused on three-level modeling, finite counterfactual auditing, end-to-end cost accounting, and systematic empirical validation for generative CAD. Optimal allocation in stratified sampling traces back to Neyman (1934). Horvitz & Thompson (1952) formalized estimation without replacement and inclusion probabilities in finite populations. Liang & Zeger (1986) provided a classical framework for correlated repeated observations within clusters. Environment A contains 72 templates, with five generations per template and 16 edit states per program, yielding 28,800 evaluation records across five systems. Environment B contains 48 templates, with four generations per template and 16 states per program, yielding 15,360 records. The calibration/test splits contain 24/48 templates in A and 12/36 templates in B. The two task sets are checked for near-duplicates before splitting. Parameter ranges and geometric constraints in environment B are constructed from its own task specifications, while the judgment criteria remain consistent with environment A. Cross-environment validation compares the allocation rule transferred directly from environment A with a rule recalibrated locally in environment B; the exact mapping is given in Appendix F.4. Equation (4) shows that additional checks of the same program can reduce only edit-state sampling error. Template heterogeneity and generation stochasticity remain even when all states are inspected. Under a fixed budget, deeper auditing also reduces the number of templates or independent generations that can be purchased, so total error may first decrease and then increase with audit depth. Whether this increase occurs, and where it begins, depends on the three variance components and execution costs rather than on the number of edits alone. The integer search compares these allocations and uses calibration-set parameters to predict the benefit of deeper auditing on held-out tasks. The continuous approximation and stationary-point derivation are given in Appendix A. When template or generation differences dominate, deeper auditing reduces measurement error for an individual program's risk but increases total estimation error for model-level risk. Redirecting the same budget to new templates or independent generations lowers total error. For systems with greater edit-state variation and higher generation costs, the relationship reverses and deeper auditing yields better estimates. The benefit directions predicted from calibration agree with the held-out observations, showing that differences in preferred depth across systems can be explained jointly by evidence source and execution cost. J.1
E XTENDED C ALIBRATION AND S PLIT P ROTOCOL
Variance components and cost parameters are estimated from an independent calibration set. Generation count and audit depth pooled across systems serve as a reference configuration that requires no 28
Preprint
target-system calibration; the feasible number of templates is then determined from measured targetsystem costs. We do not assume that this configuration outperforms a strong fixed-depth baseline. System-level calibration is used to analyze system-specific variance and cost structure and settings in which calibration results can be reused. The continuous approximation, break-even condition, and paired-calibration definition appear in Appendices A.3 and F.4. All splits are defined at the template level and stratified within task family. In the primary environment, each task family contains nine templates, of which three enter calibration and six enter testing, yielding 24/48 templates. In the transfer environment, each task family contains eight templates, with two for calibration and six for testing, yielding 12/36 templates. All generated programs and edit records from a template remain in the same split. The full calibration set executes all 16 states for every program to estimate three-level variance and execution costs. We also define a small calibration that samples a few templates per task family to study shrinkage estimation and reuse scenarios. Sample sizes, costs, and stability settings are given in Appendices D.3 and F.1. The test set is never used for variance estimation, shrinkage-weight selection, audit-depth selection, or cost-model fitting. The full test pool is used only to compute empirical reference risk and to run resampling replays. This strict separation prevents optimistic error estimates induced by model or configuration selection (Varma & Simon, 2006). J.2
E XTENDED C ONCLUSION
We study what evidence is sufficient to support reliable conclusions about generative CAD models under a fixed execution budget. For three sources of uncertainty—task templates, independent generations from the same template, and within-program counterfactual edits—we define a threelevel estimator of average failure risk and its finite-population variance, and place template scheduling, program generation, state execution, and calibration costs within a unified integer-allocation framework. The experimental protocol further includes cross-system pooling, leave-one-system-out calibration, small pilots, edit-type stratification, and model-pair-specific allocation. The experiments show that no audit depth is universally best across systems and environments. Deeper fixed auditing is a strong baseline in the primary environment, while the optimal fixed depth changes in the cross-task-family environment. Three-level variance and execution costs estimated during calibration explain these differences and reliably predict the direction of the benefit from additional auditing, although joint allocation strategies do not consistently outperform strong fixeddepth baselines. The value of system-level and stratified configurations therefore lies mainly in diagnosing system-specific deviations and guiding budget adjustment; their net efficiency also depends on calibration cost, task composition, and reuse count. Inspecting one program more thoroughly and judging a model more accurately are objectives at different levels. Under a fixed budget, they can conflict: additional edit checks reduce within-program uncertainty while consuming samples that could reveal task heterogeneity and generation stochasticity. Three-level variance and execution cost jointly explain this conflict and support prediction of audit benefit. CAD behavioral evaluation can therefore use the dominant source of uncertainty to decide when to keep auditing existing programs and when to spend the next unit of budget on new tasks or generations.
29