CIT-CAD: Constraint Intent Tree-based CAD Code Generation and Verification
arXiv:2609.07434v1 [cs.AI] 7 Sep 2026
Yali Du, Hui Sun, San-Zhuo Xi, Ming Li National Key Laboratory for Novel Software Technology, School of Artificial Intelligence, Nanjing University, China {duyl, sunh, lim}@lamda.nju.edu.cn Abstract—Natural-language Computer-Aided Design (CAD) code generation aims to turn design intent into executable and editable parametric programs. Large language models (LLMs) make this goal increasingly practical, but useful systems must preserve the construction process behind the rendered geometry. Existing benchmarks and methods mostly focus on how closely the generated CAD model matches the reference geometry, often using metrics such as Intersection over Union (IoU). Such metrics can miss errors in part decomposition, construction hierarchy, Boolean operations, sketch structure, and geometric relations. This gap calls for a representation that makes design intent explicit and lets a system check generated code against that intent. We propose CIT-CAD, a framework that infers a Constraint Intent Tree (CIT) from the input description to represent the intended entities, hierarchy, operations, and relations. The tree has two roles: it guides CAD code generation and defines expected constraints for verification. The framework extracts actual constraints from the generated program, compares them with the expected constraints, and uses mismatches to localize and repair design violations. Experiments show that the framework improves CAD generation performance, with larger gains on more complex multi-entity designs. By turning design intent into an explicit and checkable object, this work is the first attempt to move text-to-CAD generation beyond renderedgeometry matching toward construction-aware synthesis, verification, and repair. The complete artifacts are released at https://anonymous.4open.science/r/CIT-CAD-2FC2. Index Terms—CAD code generation, large language models, program verification, constraint reasoning, program repair
I. I NTRODUCTION Natural-language Computer-Aided Design (CAD) code generation studies how to produce executable CAD programs from design descriptions. Recent large language models (LLMs) have made this task increasingly feasible by generating CAD code from informal instructions [1]–[5]. Unlike a final geometric output, a CAD program is an editable construction procedure that records the modeling steps and parameters used to construct the object. These steps include 2D sketches, extrusions that turn sketches into 3D features, Boolean operations that add or cut parts, and geometric relations among parts. Parametric programming using CadQuery1 makes CAD code generation more amenable to LLMs by representing CAD models as executable construction procedures with explicit entities, parameters, and Boolean operations. Recent work has explored CAD program generation from design descriptions and visual inputs using LLMs or vision1 https://github.com/cadquery/cadquery
Fig. 1: Motivating example for construction-aware CAD evaluation. Geometric-level similarity can hide program-level errors such as missing entities, incorrect Boolean roles, or violated relations, while CIT-CAD exposes these errors as explicit constraint violations.
language models [3]–[7]. These studies demonstrate the potential of automated CAD code generation for turning informal specifications into executable programs. Their evaluations commonly report execution success and final-geometry similarity, typically using Intersection over Union (IoU) over the occupied regions of the 3D models produced by generated and reference CAD code [1], [3], [4], [8]. These geometric measures are necessary because a generated CAD program must produce the requested geometry. Final-geometry similarity is therefore necessary, but it is a coarse signal for CAD code generation. In realistic generation, a model may partially match the reference while violating construction constraints that determine how the model should be edited. An IoU score cannot localize whether an error comes from wrong decomposition, a missing feature, an incorrect Boolean role, or a violated relation between parts. Thus, natural-language CAD code generation should evaluate not only final geometry, but also the intended construction constraints behind it. This motivates a construction-aware view of CAD code generation. Instead of treating generation as text-to-code followed mainly by final-geometry scoring, a system should make the intended construction process explicit. Natural-language descriptions rarely enumerate every requirement, but they often imply expected entities, operations, and relations. Once made explicit, these requirements can provide a common target
for generation, verification, and repair. Under this view, the intended construction constraints are represented by a Constraint Intent Tree (CIT). A CIT records the expected CAD entities, their construction hierarchy, the operations that combine them, and the relations that should hold among them. It is intentionally lighter than a complete CAD program: rather than inferring every missing dimension or coordinate, it records the construction constraints that are important for program correctness. This makes the intended design explicit before code generation and comparable after code generation. Building on the CIT, we propose CIT-CAD, a framework that uses the inferred construction intent as a common reference for natural-language CAD code generation, verification, and repair. Given a design description, CIT-CAD first infers a CIT and then generates executable CAD code conditioned on both the original description and this CIT. For verification, CIT-CAD derives expected constraints from the CIT and extracts actual constraints from the generated CAD program through deterministic program and geometry analysis. It then compares the two to localize violated constraints, which are converted into repair feedback for the LLM. A revised program is accepted only when it preserves previously satisfied constraints and reduces the remaining violations. In this way, CITCAD keeps generation, verification, and repair aligned around construction constraints rather than only final geometry. We evaluate CIT-CAD on multi-entity samples derived from Text2CAD [1], comparing it with direct LLM-based CAD code generation. The evaluation combines valid-program rate, final-geometry similarity, and constraint satisfaction. CITCAD improves overall performance, especially on structurally complex designs, suggesting that explicit construction intent improves robustness beyond geometry matching alone. This paper makes the following contributions: • Constraint Intent Tree for CAD code generation. We introduce a tree structure that captures expected CAD entities, construction hierarchy, operations, and relations implied by natural-language descriptions. • CIT-based CAD generation and verification. We propose CIT-CAD, which infers a CIT, generates CAD code conditioned on it, and verifies generated programs against CIT-derived constraints. • Localized constraint repair. We convert entity-level and relation-level violations into repair feedback and accept a repair only if it preserves satisfied constraints and reduces remaining violations. • Empirical evidence beyond geometry similarity. We evaluate CIT-CAD on multi-entity Text2CAD samples and show consistent improvements in executability and constraint satisfaction across multiple LLM backbones. II. R ELATED W ORK CIT-CAD constructs CAD models via three core steps, generating CAD scripts from natural language or visual prompts, clarifying generation objectives via intermediate representations, and validating synthesized CAD programs.
A. Parametric CAD Code Generation and Benchmarks Parametric CAD programs encode both geometric shapes and editable construction history, forming the foundation of learning-based CAD synthesis. Existing datasets and paradigms support structured program generation, including ABC [9], Fusion 360 Gallery [10], DeepCAD [11], and SketchGraphs [12], as well as general programmatic shape generation methods such as CSGNet [13], ShapeAssembly [14], and CAD sequence modeling techniques [15]–[17]. Compared with vision- or point-cloud-based CAD reconstruction, text-to-CAD generation is more practically demanding and technically challenging. Natural language serves as the most intuitive and universal interaction interface for novice and expert designers, making text-conditioned CAD generation the most accessible and widely required design paradigm in real-world workflows. Recent text-to-CAD benchmarks and optimization methods [1], [3], [4], [18], [19] have validated the feasibility of LLM-based CAD programming. In contrast, multimodal CAD methods relying on drawing [20], point clouds [21], or 3D inputs [2], [5], [22]–[25] require specialized input data and are limited to specific reconstruction scenarios, resulting in less general applicability. Nevertheless, current text-to-CAD evaluations remain limited. Most approaches rely on executability checks and geometric metrics such as IoU and Chamfer distance [6], [8], which only measure final shape similarity. Generated CAD programs can produce visually identical geometries while deviating significantly in sketch structures, Boolean roles, and hierarchical modeling logic. B. Intermediate Intent Representations for Code Generation Code generation research often introduces intermediate representations to make user intent easier for LLMs to follow. Code representation models connect natural language and programs [26], while outline-based, coarse-to-fine, and structured reasoning methods use program sketches or intermediate steps to make generation more controllable [27]–[29]. Prompting and example-selection work provides a complementary form of guidance, showing that code structure can matter when constructing LLM inputs [30], [31]. Structure-aware code representations have also been used for bug localization, codechange modeling, and commit-message generation, where semantic-flow, control-flow, and context-aware representations help preserve program behavior beyond token sequences [32]– [35]. Recent software-engineering surveys further show that LLM-based development faces persistent challenges in controllability, evaluation, efficiency, and trustworthiness [36]– [38]. Formal studies of code-generation behavior examine license compliance, generalization beyond familiar problems, harmful off-the-shelf use, code simplification, and code translation [39]–[43]. Recent ICSE studies further evaluate LLMs in pragmatic and class-level code-generation settings, showing that realistic program structure and context remain difficult for general-purpose code models [44], [45]. Several studies further show that preserving structural, semantic, and executionstate information is important when translating code across languages or long contexts [46]–[48]. These studies suggest
that explicit intermediate objects or structurally informed prompts can improve controllability before code generation. For CAD code generation, the relevant intermediate object extends beyond a syntactic plan or a sequence of coding steps. It should represent construction intent, such as expected entities, hierarchy, Boolean roles, sketch properties, and interentity relations. C. Verification and Repair of Generated Programs Generated code often benefits from feedback after the first attempt. General code-generation work evaluates and improves model outputs through execution, tests, reranking, or robustness analysis [49]–[51]. Interactive and test-driven workflows further show that feedback can guide users and models toward better programs [52]. LLM-based testing research has studied unit-test generation, mutation testing, fuzzing, test-oriented model fine-tuning, and consistency-oriented evaluation of generated code [53]–[61]. At test time, self-refinement methods use feedback and iterative revision to improve LLM outputs without additional training [62]. Program repair systems extend this idea with autonomous repair agents, fact-selection mechanisms, templates, fine-tuning, and hybrid programanalysis feedback that determine what information should be supplied to the model [63]–[73]. These feedback signals are useful and typically operate at the level of tests, compiler behavior, visual review, or global geometry. They can reveal that a result is non-executable or geometrically inaccurate, while the repair target often remains global rather than tied to a specific expected entity, operation, or relation. III. M ETHOD This section presents CIT-CAD, a Constraint Intent Treebased framework for natural-language CAD code generation and verification. The key idea is to introduce an intermediate representation, called a Constraint Intent Tree (CIT), between a natural-language design description and executable CAD code. A CIT captures not only the hierarchical construction structure of a CAD model, but also the implicit geometric and semantic constraints that should hold among its parts. CITCAD uses the CIT for two purposes: guiding LLM-based CAD program generation and deriving deterministic constraints for post-generation verification. A. Overview As shown in Figure 2, given a natural-language CAD description D, CIT-CAD performs four stages. First, it infers a CIT T from D using an LLM with a schema-constrained prompt. Second, it generates a CAD program P conditioned on both D and T . Third, it deterministically derives an expected constraint set CT from T , and extracts an executed constraint set CP from P through static program analysis and runtime geometric reasoning. Finally, it compares CT and CP to validate the generated program, localizes violations to specific CIT nodes, relations, or subtrees, and uses the localized feedback to drive an iterative repair loop.
The framework is designed around the observation that geometry-only metrics, such as IoU, are insufficient for evaluating generated CAD programs. A program can produce plausible final geometry while using an incorrect construction hierarchy, missing a subtractive feature, violating coplanarity or shared-axis constraints, or losing editability. CIT-CAD therefore evaluates generated CAD code by checking whether it satisfies the structural and semantic constraints implied by the design intent. B. Constraint Intent Tree Definition 1 (Constraint Intent Tree). A Constraint Intent Tree is a tuple T = (V, E, A, R), where V is a set of nodes representing CAD entities or non-physical grouping nodes; E ⊆ V × V is a set of parent-child edges representing hierarchical construction structure; A is the set of node-level attributes and constraints; R is the set of inter-node relations encoding geometric and semantic constraints among nodes. Each entity node in a CIT corresponds to an expected parametric construction entity in the generated CAD program. In our implementation, a node contains an identifier, a human-readable name, an expected code variable, node-level constraints, relation-level constraints, and optionally child nodes. Construction semantics are represented by verifiable constraints, especially the node-level Boolean role field and relation-level spatial constraints. Non-physical grouping nodes are used only to organize related entities and do not correspond to generated solids. Node-level constraints describe properties that can be verified on individual generated entities. Let F denote the set of node-level constraint fields, and let Y denote the union of their possible values, including Boolean, categorical, and integervalued domains. We represent the node-level annotation set as A = {(v, f, y) | v ∈ V, f ∈ F, y ∈ Y},
(1)
where v is a CIT node, f is a node-level field such as Sketch type, Is connected, etc., and y is the expected value. Relation-level constraints describe expected relationships between entities or subtrees. Let G denote the set of relation types. These include Contact type, Intersection , Has coplanar faces, and Has shared axis. We represent the relation annotation set as R = {(vi , vj , g, y) | vi , vj ∈ V, g ∈ G, y ∈ Y},
(2)
where vi and vj are participating CIT nodes, g is a relation type, and y is the expected relation value, such as face for a face-contact relation or True for a shared-axis relation. The CIT is intentionally not a full CAD program. It omits low-level numeric details when they are underspecified in natural language, but retains the construction intent that should be preserved by the generated code. This makes the representation suitable both for generation and for validation.
Constraint Intent Tree of sample 00014943
Natural Language Design Description Start by creating a new coordinate system and drawing a twodimensional sketch to form a rectangular block. The block has a length of 0.7500000000000001 units, a width of 0.25000000000000006 units, and a height of 0.06250000000000001 units. Next, create another coordinate system and draw two circles on the top face of the block to define the positions of the cylindrical holes. The holes are located on opposite sides of the block, each with a diameter that fits within the block's dimensions. After defining the holes, extrude them through the block to remove material, creating the final shape. The final shape is a rectangular plate with two circular holes on opposite sides, maintaining a flat surface and sharp edges.
“Sketch_type”: “other”
Entity Constrains
Output Entity
CIT Construction
“Sketch_type”: “other”
“Is_open_profile”: False
The fundamental geometric primitive type of the sketch contour.
“Boolean_role”: True
Is_connected
Entity
“Is_connected”: True
“Sketch_type”: “circle”
“Is_open_profile”: False
“Is_open_profile”: False
“Boolean_role”: True
“Boolean_role”: False
Is_open_profile Whether the sketch curve is an unclosed open profile without connected endpoints.
Boolean_role:
“Has_shared_axis“: true
Sketch_Hole_1
“Sketch_type”: “circle”
“Contact_type": "face"
“Is_connected”: True
“Intersection“: true
“Is_open_profile”: False
“Contact_type": "face"
“Is_connected”: True
“Intersection“: true
Whether the sketch serves as a material-additive base solid instead of a materialsubtracting cutting feature during boolean modeling.
Expected Constrains
Relation Constrains Intersection
“Boolean_role”: False
Cut
“Sketch_type”: “rect”
Whether the sketch consists of a single uninterrupted connected curve segment.
Sketch_ Hole_2
“Is_connected”: True
Sketch_Rect_0
CadQuery
Sketch_type
“Is_connected”: True
Entities overlap geometrically.
Contact_type Touching relation such as face contact.
“Is_open_profile”: False
Has_shared_axis
“Boolean_role”: True
Circular/cutout features share an axis.
import cadquery as cq # Generating a workplane for sketch 0 wp_sketch0 = cq.Workplane(cq.Plane(cq.Vector(-0.75, 0.671875, 0.0), cq.Vector(1.0, 0.0, 0.0), cq.Vector(0.0, 0.0, 1.0))) loop0=wp_sketch0.moveTo(0.6484375, 0.0).lineTo(0.6484375, 0.21842105263157896).lineTo(0.0, 0.21842105263157896).lineTo(0.0, 0.0).close() solid0=wp_sketch0.add(loop0).extrude(0.0546875) solid=solid0 # Generating a workplane for sketch 1 wp_sketch1 = cq.Workplane(cq.Plane(cq.Vector(0.6796875, -0.5703125, 0.0546875), cq.Vector(1.0, 0.0, 0.0), cq.Vector(0.0, 0.0, 1.0))) loop1=wp_sketch1.moveTo(0.039473684210526314, 0.0).circle(0.039473684210526314) solid1=wp_sketch1.add(loop1).extrude(-0.0546875) solid=solid.cut(solid1) # Generating a workplane for sketch 2 wp_sketch2 = cq.Workplane(cq.Plane(cq.Vector(-0.25, 0.5703125, 0.0546875), cq.Vector(1.0, 0.0, 0.0), cq.Vector(0.0, 0.0, 1.0))) loop2=wp_sketch2.moveTo(0.039473684210526314, 0.0).circle(0.039473684210526314) solid2=wp_sketch2.add(loop2).extrude(-0.0546875) solid=solid.cut(solid2) final_body = solid.val() final_body.exportStep("model.step") final_body.exportStl("model.stl")
Verified CAD Program
Executed Constrains
Generated CAD Code LLMs
Localized Repair Feedback
Verified CAD Model
Verification
Fig. 2: Overview of the framework CIT-CAD. The framework infers a Constraint Intent Tree from the natural-language description, generates CAD programs conditioned on the tree, extracts expected and executed constraints, validates constraint satisfaction, and uses localized violations to guide monotonic repair. TABLE I: Constraint Category in CIT-CAD. Category
Constraint
Meaning
Values
Entity Entity Entity Entity
Sketch type Is connected Is open profile Boolean role
Geometric primitive category of the 2D sketch contour Indicator whether the sketch forms a single continuous connected curve Indicator whether the sketch corresponds to an unclosed polyline profile Functional boolean operation type for feature construction
rect , circle , etc. True/False True/False True: additive; False: subtractive
Relation Relation Relation Relation Relation Relation Relation Relation
Intersection Contact type Has coplanar faces Has shared axis Tangency Coaxiality Alignment Symmetry
Indicator whether two constructed solids have volumetric overlap Specific geometric contact form between adjacent solid components Total count of coplanar face pairs across paired entities Indicator whether arc or cylinder features share a central axis Indicator whether adjacent surfaces maintain smooth tangent connection Indicator whether two structural parts share the same central axis Spatial directional and coordinate alignment constraint between parts Structural mirror or rotational symmetry relationship of components
True/False point, line , face, none Integer count True/False True/False True/False x, y, z, parallel , perpendicular mirror, rotational
C. CIT Constraint Annotation CIT-CAD annotates each inferred CIT with a compact constraint vocabulary that can be checked against generated CAD programs. The vocabulary is divided into entity-level constraints and relation-level constraints. Entity-level constraints describe local properties of individual construction entities, while relation-level constraints describe spatial or semantic relations among entities. This annotation is part of the method pipeline: the expected constraints used for verification are derived from the CIT inferred from the input specification, not directly from the reference CAD program. Tab. I lists the constraint vocabulary used in the current implementation. Entity-level constraints include Sketch type, Is connected, Is open profile , and Boolean role. Relationlevel constraints include Intersection , Contact type, Has coplanar faces, and Has shared axis. These constraints are selected because they are common in multi-entity CAD programs and can be deterministically extracted from generated code through static analysis and runtime geometric reasoning. For each entity node, Sketch type identifies the primitive
or composite sketch class, Is connected checks whether the sketch consists of a single connected profile, Is open profile records whether the profile is represented as an open drawing chain before closure, and Boolean role indicates whether the entity contributes additively to the final model or serves as a subtractive cutting feature. For pairs or groups of entities, relation-level constraints check whether entities overlap, touch and how they touch, share coplanar faces, or share circular/cylindrical axes. These constraints capture construction semantics that are difficult to evaluate using final-shape overlap alone. D. CIT Inference from Natural Language Given a natural-language description D, CIT-CAD first prompts an LLM to infer a CIT in JSON format. The prompt specifies the node schema, node-level constraints, relationlevel constraints, and naming conventions. It instructs the model to use stable snake-case identifiers, create one node for every physical solid or cutting tool, and use null values for uncertain constraints rather than hallucinating precise geometry.
Formally, the CIT inference stage estimates: T = fθtree (D),
(3)
where fθtree is an LLM used as a structured information extractor. The output is parsed as JSON and normalized to ensure that it contains a root node, expected variables, node constraints, relation constraints, and child nodes. Although the CIT is inferred by an LLM, downstream verification does not trust the LLM output blindly. Instead, the CIT is treated as the explicit design intent extracted from the natural-language specification. CIT-CAD then uses deterministic procedures to derive constraints from this tree and to check generated code against them. E. Constraint-Guided CAD Code Generation After inferring the CIT, CIT-CAD generates CAD programs using the natural-language description and the CIT jointly: P0 = fθcode (D, T ).
(4)
The code-generation prompt requires the generated program to import cadquery as cq, create one variable for every entity node in the CIT, use the node’s expected variable name, construct each entity through a Workplane sketch followed by . extrude (...) , compose additive and subtractive entities according to their Boolean role, and assign the final object to a variable named solid . Conditioning code generation on the CIT has two benefits. First, it gives the LLM an explicit construction scaffold, reducing ambiguity in entity naming, part decomposition, and boolean composition. Second, it makes the generated code more analyzable: because expected entities are named in the CIT, the validator can compare extracted code constraints against the intended nodes and relations. F. Deterministic Constraint Extraction CIT-CAD derives two comparable short-constraint sets: the expected constraint set CT from the CIT T , and the executed constraint set CP from the generated program P . The expected set is obtained by normalizing the CIT annotations:
also records a node map that links each constraint back to its CIT node path, enabling later localization. Executed constraints from generated code. To extract CP , CIT-CAD analyzes the generated CAD program using deterministic program and geometry analysis. Static AST analysis identifies variables assigned from . extrude (...) , recovers sketch operations, classifies Sketch type, and computes Boolean role from union, cut, and intersect expressions. Runtime geometric reasoning executes the CAD program and inspects generated geometry to determine relation-level constraints such as Contact type, Intersection , Has coplanar faces, and Has shared axis. The extracted result is normalized into the same short-constraint format as CT . Formally, for a constraint source X ∈ {T, P }, where X = T denotes the expected source induced by the CIT and X = P denotes the executed source extracted from the generated program, let NX be the set of entity variables observed in that source. CIT-CAD represents entity constraints from source X as a partial map EX : NX × F ⇀ Y,
(6)
where EX (n, f ) = y means that entity variable n ∈ NX has field f ∈ F with value y ∈ Y. Relation constraints from source X are represented as canonical triples 2 QX = {(g, sort(u), y) | g ∈ G, u ∈ NX , y ∈ Y}.
(7)
Here, u is the tuple of participating entity variables, and sort(u) canonicalizes their order before comparison. The normalized short-constraint set for source X is therefore CX = (EX , QX ).
(8)
This gives CT = (ET , QT ) for the CIT-derived expected constraints and CP = (EP , QP ) for the generated-program constraints. This design separates uncertain LLM-based interpretation from deterministic validation. The LLM may infer or generate an imperfect result, but the verification step uses a fixed analyzer and geometry engine to decide whether the generated program satisfies the CIT-derived requirements. G. Structural and Semantic Verification
CT = Normalize(A, R) = (ET , QT ).
(5)
Here, A and R are the raw node-level and relation-level intent annotations stored on the CIT, ET is the expected entityconstraint map, and QT is the expected relation-constraint set. The conversion maps each CIT node v to its expected code variable n, removes unspecified or null-valued annotations, and canonicalizes relation participants so that pairwise relation comparison is order-invariant. Expected constraints from the CIT. The CIT-to-constraint conversion is deterministic. For each node-level annotation (v, f, y) ∈ A, if node v maps to expected variable n, CITCAD creates the entity constraint ET (n, f ) = y. For each relation annotation (vi , vj , g, y) ∈ R, if vi and vj map to expected variables ni and nj , CIT-CAD creates the canonical relation constraint (g, sort(ni , nj ), y) ∈ QT . The converter
Given expected constraints CT = (ET , QT ) and executed constraints CP = (EP , QP ), CIT-CAD verifies the generated program using exact equality over the normalized shortconstraint representation. For an expected entity variable n and field f , the entity-field satisfaction predicate is ( 1, n ∈ NP ∧ EP (n, f ) = ET (n, f ), satent (n, f ) = 0, otherwise. (9) If an expected entity variable does not appear in P , all constraints associated with that entity are marked as violated and the feedback records a missing-entity violation. For relation constraints, CIT-CAD first canonicalizes the participating entity names and then checks set membership: satrel (g, u, y) = I [(g, sort(u), y) ∈ QP ] ,
(10)
for every expected relation triple (g, sort(u), y) ∈ QT . The generated program satisfies the CIT if and only if all expected entity-field constraints and all expected relation constraints are satisfied: ^ Valid(P, T ) = satent (n, f ) (n,f )∈dom(ET )
^
∧
(11)
satrel (g, u, y).
(g,u,y)∈QT
The validator also constructs explicit satisfied and violated constraint sets. Let CT denote the finite set of individual expected constraints contained in CT : CT = {(n, f, ET (n, f )) | (n, f ) ∈ dom(ET )} ∪ QT .
(12)
CIT-CAD partitions CT into a satisfied set and a violated set: S(P, T ) = {c ∈ CT | sat(c, P ) = 1}, V (P, T ) = CT \S(P, T ). (13) Here, S(P, T ) is the set of expected constraints satisfied by program P , V (P, T ) is the set of expected constraints violated by P , and sat(c, P ) denotes Eq. (9) when c is an entityfield constraint and Eq. (10) when c is a relation constraint. This partition is exactly the information later consumed by the repair loop. The validator produces three outputs. First, it reports Valid(P, T ). Second, it records satisfied and violated constraint identifiers. Third, it localizes every violated constraint to the corresponding CIT node, relation, or subtree using the node map produced during CIT-to-constraint conversion. This localization is important because it turns validation failures into actionable repair feedback. Instead of reporting only that the final geometry is inaccurate, CIT-CAD can indicate that a specific node is missing, a specific sketch field is wrong, or a specific relation between two CIT nodes is not satisfied. H. Localized Reflexion Repair CIT-CAD uses localized validation feedback to iteratively repair the generated program. After each validation round, the feedback contains violated entities, violated relations, their corresponding CIT paths, and concrete repair guidance. The LLM repair prompt receives the natural-language description, the CIT, the current code, and the validation feedback, and returns a revised CAD program. However, CIT-CAD does not assume that the LLM repair model is reliable. A repair may fix one violated constraint while accidentally breaking a previously satisfied one. To prevent such regression, CIT-CAD wraps the LLM repair model in a monotonic constraint-preserving acceptance loop, shown in algorithm 1. Let K be the maximum number of repair attempts, and let Pt be the accepted program at repair iteration t. The deterministic validator partitions CT into satisfied and violated constraints for Pt : St = S(Pt , T ),
Vt = V (Pt , T ).
(14)
Algorithm 1: Monotonic CIT-guided repair Input: Description D, CIT T , initial program P0 , max repair attempts K Output: Verified or best accepted CAD program P 1 CT ← ExtractConstraints(T ) , P ← P0 , L ← ∅; 2 for t ← 0 to K − 1 do 3 (St , Vt ) ← Validate(P, CT ); 4 if Vt = ∅ then 5 return P ; 6 end 7 L ← L ∪ St ; 8 Ft ← Localize(Vt , T ); 9 P ′ ← LLMRepair(D, T, P, Ft , L); 10 (S ′ , V ′ ) ← Validate(P ′ , CT ); 11 if L ⊆ S ′ and |V ′ | < |Vt | then 12 P ← P ′; 13 end 14 if |V ′ | ≥ |Vt | then 15 return P ; 16 end 17 end 18 return P ;
Here, St and Vt are shorthand for the satisfied and violated sets at iteration t. The implementation maintains a locked set Lt of constraints that have been satisfied by any accepted program so far: t [ Lt = Si . (15) i=0
The validator also converts Vt into localized repair feedback Ft , which records the violated constraints and their corresponding CIT nodes, relations, or subtrees. After the LLM proposes a candidate repair Pt′ , CIT-CAD accepts the candidate only if both of the following conditions hold: Lt ⊆ S(Pt′ , T ).
(16)
|V (Pt′ , T )| < |Vt |.
(17)
If either condition fails, the candidate is rejected and the previous program Pt is retained. Thus, the LLM may propose arbitrary changes, but only validation-approved changes are committed. Proposition 1 (Regression-free accepted repairs). Assume the validator is deterministic and CIT-CAD accepts a candidate repair only when Eq. (16) holds. Then no accepted repair violates a constraint that was satisfied in the previous accepted program. Proof. At iteration t, every constraint satisfied by the current accepted program belongs to the locked set because St ⊆ Lt . A candidate Pt′ is accepted only if Lt ⊆ S(Pt′ , T ). Therefore, every constraint satisfied by Pt remains satisfied after accepting Pt′ . If any constraint in Lt becomes violated, the candidate fails the acceptance rule and is rejected, so the accepted
program remains Pt . Thus, accepted repairs are regressionfree with respect to the CIT-derived constraint set. The second acceptance condition, Eq. (17), ensures strict progress in the number of remaining violations for every accepted repair. Since the CIT-derived constraint set is finite, the number of accepted repairs is bounded by the number of initially violated constraints, although the system may still terminate earlier due to a fixed iteration budget or repeated rejected candidates. IV. DATASET AND M ETRICS We construct our evaluation set from Text2CAD [1] to study construction-aware CAD code generation. Each evaluated sample contains a natural-language description, a held-out reference CAD program, and generated CAD programs from the evaluated methods. This design allows us to evaluate generated programs from both geometry-level and constructionlevel perspectives without using the reference program as input to generation or CIT inference. A. Data Source and Sample Selection We use Text2CAD [1] as our benchmark dataset. Text2CAD contains about 178K text-to-CAD samples. Following its preprocessing protocol, we remove samples with missing naturallanguage descriptions, missing CAD programs, or duplicate descriptions, resulting in 151K valid text-to-CAD pairs. Since CIT-CAD targets construction-aware generation, we focus on multi-entity CAD programs. We parse each reference CAD program and count the number of explicit extruded construction entities. Samples with only one entity are excluded because they usually involve limited construction hierarchy, Boolean composition, or inter-entity relations. The final evaluation subset contains 26,783 multi-entity samples. Their entitycount distribution is 17,193 samples with two entities, 5,966 with three entities, 1,883 with four entities, 731 with five entities, 671 with six entities, and 339 with seven or more entities. This long-tail distribution allows us to evaluate both common simple multi-entity designs and more complex cases that require preserving richer construction structure. For each selected sample, the natural-language description is used as the input specification. The reference CAD program is used only for geometry-level evaluation and entity-count grouping, and is never provided to the code generator or the CIT inference stage. B. Evaluation Metrics We evaluate generated CAD programs using four complementary metrics: valid syntax rate, geometry-level success rate, Intersection over Union, and constraint satisfaction rate. a) Valid Syntax Rate: Valid syntax rate (VSR) measures the percentage of generated programs that can be executed successfully by the CadQuery runtime and converted into a valid solid for evaluation. Programs with syntax errors, unsupported CadQuery calls, runtime exceptions, or missing final solids are counted as invalid. VSR reflects whether a method can generate executable CAD code.
b) Geometry-level metrics: We use Intersection over Union (IoU) to measure final-shape similarity between the generated model and the reference model. Both models are normalized before comparison, and IoU is computed over their occupied 3D regions: IoU(G, R) =
Vol(G ∩ R) , Vol(G ∪ R)
(18)
where G and R denote the generated and reference solids, respectively. We report mean IoU and median IoU over samples whose generated programs execute successfully and whose IoU values fall in the valid range [0, 1]. We also report Success Rate ([email protected]), defined as the percentage of all evaluated samples whose generated geometry reaches an IoU greater than 0.95 with the reference geometry. c) Constraint Satisfaction Rate: We introduce Constraint Satisfaction Rate (CSR) to measure construction-level correctness based on CIT-derived structural and semantic constraints. For each sample i, let Ci be the deterministic constraint set derived from the CIT inferred from the naturallanguage description, and let Ĉi be the constraint set extracted from the generated program. CSR is defined as CSRi =
|{c ∈ Ci | c is satisfied by Ĉi }| . |Ci |
(19)
The final CSR is averaged over the evaluation set. Unlike IoU, CSR evaluates whether the generated program preserves the intended construction process, including entity decomposition, sketch properties, Boolean roles, and inter-entity relations. Therefore, CSR complements geometry-level metrics rather than replacing them. In our experiments, each natural-language description is used to infer a CIT and its corresponding construction constraints. These CIT-derived constraints are used for CSR computation and for checking whether the generated CAD program preserves the intended construction process. The reference CAD program is used only for geometry-level evaluation, such as building the reference solid for IoU computation. This separation ensures that constraint-level evaluation is based on the input design intent, while the reference program remains a held-out target for final-geometry comparison. V. E XPERIMENTS This section evaluates whether CIT-CAD improves naturallanguage CAD code generation across different backbone LLMs and different levels of construction complexity. We introduce the experimental settings and organize the evaluation around the following research questions. A. Experimental Setup Evaluated methods. For each backbone LLM, we compare two generation settings. The Vanilla directly prompts the LLM to generate CadQuery code from the natural-language description. CIT-CAD first infers a Constraint Intent Tree, generates CadQuery code conditioned on the tree, and then applies constraint-guided validation and repair. This comparison
(a) Constraint Satisfaction Rate vs. Mean IoU (b) Constraint Satisfaction Rate vs. Median IoU Pearson r=0.00, Spearman rho=-0.02 Pearson r=0.00, Spearman rho=-0.02
100%
80%
Mean IoU
isolates the effect of introducing explicit construction intent and deterministic feedback while keeping the underlying LLM family fixed. Backbone LLMs. We evaluate CIT-CAD with two opensource LLMs, Qwen3-A3B-Instruct [74] and DeepSeekCoder-V2-Lite-Instruct [75], and one closed-source LLM, GPT-5.4-mini [76]. The open-source models are served through vLLM, while GPT-5.4-mini is accessed through an OpenAI-compatible API. Each backbone is used for both the direct Vanilla and the CIT-CAD pipeline.
60%
40%
20%
0% 0%
20%
40%
60%
80%
Constraint Satisfaction Rate CIT-CAD (Qwen3) Sample
B. Quantitative Analysis ® RQ 1. How do LLM backbone and CAD entity complexity affect CIT-CAD performance? Tab. II reports the results grouped by reference entity count. Across the open-source LLMs, CIT-CAD consistently improves execution success over direct generation. The improvement is especially clear as entity count increases. The evaluated backbones show different behaviors. DeepSeek-Coder benefits strongly from the CIT-CAD wrapper in terms of VSR and CSR. Its CSR improves by 86.5% on two-entity samples and by 276.0% on samples with seven or more entities, indicating that the repaired programs satisfy substantially more extracted CIT constraints. Qwen3 also obtains consistent VSR gains and stronger Success Rate in most entity groups, suggesting that explicit construction intent and deterministic feedback are useful not only for weaker code models but also for stronger CAD code generators. GPT-5.4-mini shows a smaller but positive aggregate effect after excluding unavailable generated files: CIT-CAD improves overall VSR by 3.7%, Success Rate by 19.6%, mean IoU by 3.5%, median IoU by 23.5%, and CSR by 125.7%. Finding. CIT-CAD consistently improves the quality of CAD models across different backbones. The gain becomes larger as construction complexity increases for the open-source backbones, showing that explicit Constraint Intent Trees are particularly helpful for multi-entity CAD programs. The CSR results show that CIT-CAD substantially improves constraint-level structural validity.
® RQ 2. How about the correlation between geometriclevel and constraint-level correctness? RQ1 shows that IoU and CSR do not always move in the same direction. To quantify this gap, we pair each valid generated model’s normalized IoU with its corresponding CSR and compute sample-level Pearson and Spearman correlations. Fig. 3 visualizes the Qwen3 results at the sample level. The correlation between CSR and IoU is close to zero, with Pearson r = 0.022 and Spearman ρ = 0.023. This indicates that final-shape overlap does not reliably predict whether the generated CAD program preserves construction intent. We also observe high-IoU/low-CSR cases in which the final solid is geometrically similar to the reference but the program omits expected entities, assigns incorrect Boolean roles, or uses a different sketch decomposition. These cases motivate CSR as a complementary metric to IoU.
100%
0%
20%
40%
60%
80%
100%
Constraint Satisfaction Rate CIT-CAD (Qwen3) Bin Trend
Fig. 3: Correlation between geometric-level similarity and constraintlevel correctness on Qwen3. Each point is one generated sample. The weak Pearson and Spearman correlations show that mean/median IoU and CSR capture different aspects of CAD program quality.
Finding. Geometric-level similarity does not strongly reflect constraint-level correctness. IoU-only evaluation can therefore overestimate text-to-CAD generation quality, especially when a generated program produces a similar final shape through a structurally different construction process.
® RQ 3. How much does constraint-guided repair improve generated CAD programs? CIT-CAD does not stop after the first code generation attempt. It validates the generated program against the extracted Constraint Intent Tree, converts violated constraints into localized feedback, and asks the same backbone LLM to repair the program. A repaired candidate is accepted only when it preserves locked constraints and reduces the number of remaining violations. If a candidate fails to introduce newly satisfied constraints or reduces no remaining violation, the repair loop terminates early for that sample. For stopped samples, the last accepted CSR is carried forward in later iterations, so each iteration is evaluated over the same fixed sample cohort. Fig. 4 shows the CSR distribution of Qwen3 across repair iterations. The average CSR increases monotonically from 18.2% before repair to 21.6%, 25.0%, 28.4%, and finally 28.9% after iterative repair. This corresponds to a 10.7% absolute improvement, or a 58.8% relative improvement over the initial CSR. The distribution also shifts upward across iterations, indicating that localized constraint feedback improves not only a small subset of samples but the overall constraintsatisfaction profile. The improvement is most pronounced in the early repair rounds. The first three repair iterations each bring clear gains, while the improvement from Iteration 3 to Iteration 4 becomes smaller. This suggests that constraint-guided repair quickly fixes many actionable violations, such as missing or mismatched structural constraints, but later iterations face harder cases where remaining errors are less easily resolved by textual feedback alone. The diminishing gain also supports the use of early stopping: once a repair candidate no longer reduces the violation set, continuing to modify the program is unlikely to provide substantial benefit and may risk unnecessary changes.
TABLE II: RQ1 results by LLM backbone and reference entity count. Signed percentages after CIT-CAD results indicate relative change over the corresponding Vanilla setting, with green for improvement and red for degradation. Method
Metric
Number of reference entities
#Number of Samples DeepSeek-Coder-V2-Lite-Instruct Valid Syntax Rate Success Rate ([email protected]) Vanilla Mean IoU Median IoU Constraint Satisfaction Rate
CIT-CAD
Valid Syntax Rate Success Rate ([email protected]) Mean IoU Median IoU Constraint Satisfaction Rate
Qwen3-A3B-Instruct Valid Syntax Rate Success Rate ([email protected]) Vanilla Mean IoU Median IoU Constraint Satisfaction Rate
CIT-CAD
Valid Syntax Rate Success Rate ([email protected]) Mean IoU Median IoU Constraint Satisfaction Rate
GPT-5.4-mini Valid Syntax Rate Success Rate ([email protected]) Vanilla Mean IoU Median IoU Constraint Satisfaction Rate
CIT-CAD
Valid Syntax Rate Success Rate ([email protected]) Mean IoU Median IoU Constraint Satisfaction Rate
Entity=2 17193
Entity=3 5966
Entity=4 1883
Entity=5 731
Entity=6 671
Entity≥7 339
All 26783
77.0% 4.2% 32.6% 24.6% 14.8%
73.5% 1.5% 25.7% 17.6% 15.0%
68.1% 0.7% 22.6% 14.9% 15.5%
62.1% 0.4% 16.6% 10.1% 16.4%
49.6% 0.6% 8.5% 2.0% 13.8%
48.7% 0.3% 10.3% 1.3% 9.6%
74.1% 3.1% 29.0% 21.1% 14.8%
85.3%(+10.8%) 4.5%(+7.1%) 32.4%(-0.6%) 26.7%(+8.5%) 27.6%(+86.5%)
82.5%(+12.2%) 1.6%(+6.7%) 25.9%(+0.8%) 17.8%(+1.1%) 28.1%(+87.3%)
78.3%(+15.0%) 0.8%(+14.3%) 25.0%(+10.6%) 15.0%(+0.7%) 29.7%(+91.6%)
76.1%(+22.5%) 0.4% (0.0%) 19.9%(+19.9%) 13.8%(+36.6%) 30.9%(+88.4%)
68.7%(+38.5%) 0.8%(+33.3%) 8.8%(+3.5%) 1.6%(-20.0%) 35.4%(+156.5%)
67.5%(+38.6%) 0.3% (0.0%) 8.4%(-18.4%) 0.8%(-38.5%) 36.1%(+276.0%)
83.3%(+12.3%) 3.3%(+7.3%) 29.2%(+0.5%) 22.6%(+7.0%) 28.3%(+90.3%)
80.3% 7.2% 36.0% 27.0% 17.5%
77.4% 3.0% 28.8% 20.2% 18.0%
74.2% 1.5% 25.0% 17.7% 15.0%
65.3% 1.0% 19.5% 14.3% 14.3%
64.4% 0.9% 10.0% 1.9% 5.4%
46.3% 0.3% 9.2% 2.5% 6.6%
78.0% 5.5% 32.6% 23.3% 16.7%
87.1%(+8.5%) 7.6%(+5.7%) 36.0% (0.0%) 27.0%(+0.1%) 30.1%(+71.8%)
85.1%(+10.0%) 3.6%(+21.3%) 29.7%(+3.1%) 21.1%(+4.0%) 32.1%(+78.5%)
80.7%(+8.7%) 1.7%(+10.3%) 25.3%(+1.2%) 17.2%(-2.5%) 33.9%(+126.4%)
74.3%(+13.8%) 1.1%(+14.3%) 20.4%(+4.2%) 14.3%(+0.1%) 34.2%(+139.9%)
73.3%(+13.9%) 1.0%(+16.7%) 10.6%(+6.2%) 2.1%(+11.4%) 36.1%(+572.0%)
54.6%(+17.8%) 0.6%(+100.0%) 9.2% (0.0%) 2.0%(-20.0%) 43.3%(+560.7%)
85.1%(+9.1%) 5.9%(+7.9%) 32.8%(+0.6%) 23.5%(+0.9%) 31.8%(+90.1%)
88.5% 5.6% 33.9% 24.9% 18.0%
89.2% 3.6% 28.9% 20.1% 14.5%
86.7% 1.7% 27.3% 22.2% 16.8%
78.3% 0.8% 27.6% 17.3% 10.0%
69.8% 0.5% 9.3% 2.2% 11.5%
61.5% 0.3% 4.8% 0.8% 11.4%
87.2% 4.6% 31.2% 22.6% 16.7%
91.9%(+3.8%) 6.8%(+22.0%) 34.9%(+2.9%) 30.7%(+23.3%) 37.4%(+107.8%)
92.0%(+3.1%) 3.9%(+8.3%) 29.3%(+1.4%) 24.6%(+22.4%) 37.5%(+158.6%)
90.9%(+4.9%) 2.7%(+58.8%) 29.6%(+8.4%) 27.6%(+24.3%) 37.0%(+120.2%)
89.4%(+14.2%) 1.1%(+37.5%) 29.1%(+5.4%) 24.5%(+41.6%) 36.3%(+263.0%)
72.7%(+4.2%) 1.3%(+160.0%) 15.5%(+66.7%) 3.9%(+77.3%) 37.2%(+223.5%)
81.0%(+31.5%) 0.9%(+200.0%) 12.0%(+150.0%) 1.3%(+62.5%) 50.0%(+338.6%)
90.4%(+3.7%) 5.5%(+20.7%) 32.3%(+3.8%) 27.9%(+23.7%) 37.7%(+125.7%)
Average Constraint Satisfaction Rate
Constraint Satisfaction Rate
100% 80% 60% 40% 20%
18.2%
21.6%
25.0%
28.4%
28.9%
Iter 3
Iter 4
0% Iter 0
Iter 1
Iter 2
Repair iteration
Fig. 4: Constraint Satisfaction Rate Distribution Across Repair Iterations of Qwen3. The violin plots show the CSR distribution over all evaluated samples at each repair iteration, while the line reports the average CSR. Finding. Constraint-guided repair turns CIT-CAD from a one-shot generator into an iterative verifier-generator. On Qwen3, the repair loop improves the average CSR from 18.2% to 28.9%, demonstrating that deterministic validation and localized feedback can progressively increase constraint satisfaction. The gains are strongest in the first few iterations and gradually saturate.
® RQ 4. What is the distribution of violated construction constraints? We further analyze violated construction constraints to reveal program-level failure modes beyond final-shape IoU. We group violations by CIT field, such as sketch type, Boolean role, connectivity, contact, coplanarity, alignment, symmetry, and other geometric relations. Since one sample may violate multiple constraints, we report normalized counts as violations per 100 extracted constraints. Fig. 5 shows two main patterns. First, CIT-CAD reduces the total normalized violation count across all three backbones. The reduction is visible for DeepSeek-Coder-V2, Qwen3, and GPT-5.4-mini, indicating that explicit CIT scaffolding improves construction-intent preservation beyond a single model family. Second, the dominant remaining violations are sketchlevel constraints, especially Sketch type and Is connected. This suggests that CIT-CAD is effective at imposing a global entity scaffold, but fine-grained local sketch reconstruction remains the hardest part of the task.
Sketch type
Connectivity
DS-Coder-V2
Boolean role
Open profile
Intersection
35.3 Δ -7.0
+CIT-CAD
19.8
23.2
Δ -3.5
27.1
GPT-5.4-mini
19.7
34.6 Δ -8.6
+CIT-CAD
4.9
17.3
25
4.3
5.7 Δ +1.5
Tangency
4.2 Δ -0.9
Coaxiality
Symmetry
14.9
2.6
4.7 Δ -0.4 Δ -1.2
3.5
4.7
Correct
Δ +13.3 28.2
16.6
Δ +15.3 31.9
16.7
Δ +21.0 37.7
2.3
50
Alignment
3.5
3.7
Δ -0.4 4.3
Shared axis
5.1 3.9 Δ -0.2 Δ +1.0 Δ -1.2 2.5 4.9
23.6
Δ -6.3
26.0
0
Δ +0.0 5.7
5.7
34.8 Δ -7.7
+CIT-CAD
Coplanar faces
24.0
28.3
Qwen3
Contact type
Δ -4.2
75
100
Share of constraints (%)
Fig. 5: Normalized distribution of violated constraints. Each horizontal bar decomposes the number of violated constraints per 100 extracted constraints by constraint type. Lower total length indicates fewer construction-intent violations.
Sample: 00154369 Generated CAD Model (Vanilla)
(vs Ground Truth)
IoU: 98.58%
CSR: 8.33%
Repaired CAD Model (CIT-CAD)
(vs Ground Truth)
IoU: 99.99%
CSR: 100%
Sample: 00853530 Generated CAD Model (Vanilla)
(vs Ground Truth)
IoU: 90.38%
CSR: 20.00%
Repaired CAD Model (CIT-CAD)
(vs Ground Truth)
IoU: 100.00%
CSR: 100%
Sample: 00931238 Generated CAD Model (Vanilla)
(vs Ground Truth)
IoU: 90.41%
CSR: 10.00%
Repaired CAD Model (CIT-CAD)
(vs Ground Truth)
IoU: 99.99%
CSR: 100%
Sample: 00389769 Generated CAD Model (Vanilla)
(vs Ground Truth)
IoU: 93.97%
CSR: 4.76%
Repaired CAD Model (CIT-CAD)
(vs Ground Truth)
IoU: 99.99%
CSR: 100%
Fig. 6: Qualitative case study of CIT-CAD on representative multientity CAD samples.
Relation-level violations, such as contact type, coplanarity, shared axis, tangency, coaxiality, alignment, and symmetry, account for smaller portions of the normalized violation counts. This does not imply that such relations are less important. Many relation checks depend on first generating the correct entities and sketch profiles. Thus, the main bottleneck is improving sketch-type grounding and local profile construction while preserving the CIT-level entity and relation scaffold. Finding. CIT-CAD reduces normalized constraint violations across all evaluated backbones, but it still fails most often on fine-grained sketchlevel constraints. The remaining errors are dominated by sketch type and profile connectivity, while relation-level violations form a smaller but still meaningful portion of the error distribution. These results point to future improvements in fine-grained sketch extraction, parameter grounding, and relation-aware repair.
C. Qualitative Analysis (Case Study) ® RQ 5. What benefits does CIT-CAD bring on typical multi-entity CAD cases? Fig. 6 provides qualitative evidence for the quantitative findings. Although Vanilla generation can sometimes produce a visually plausible shape with high IoU, it may still fail to preserve the intended construction process. For example, in Sample 00931238, the Vanilla model generates the main cylindrical body and the central hole, but misses the side through-hole. This indicates that the model captures the coarse geometry while failing to execute an important subtractive construction step, leading to a low CSR of 10.00% despite an IoU of 90.41%. CIT-CAD makes such errors explicit through the Constraint Intent Tree. The omitted side hole is detected as a violated constraint related to a missing or incorrect subtractive entity. This localized feedback provides direct repair guidance: instead of regenerating the whole model blindly, CIT-CAD can focus on correcting the violated construction role and inserting the required cut operation. After repair, the model recovers the intended side-hole structure, improving IoU to 99.99% and CSR to 100%. Finding. The case study shows that CIT-CAD improves CAD generation by exposing construction-level errors that are not fully reflected by IoU. Vanilla may produce a globally similar shape while omitting entities or confusing additive/subtractive roles. By converting these errors into violated constraints, CIT-CAD provides localized and interpretable repair signals, leading to programs that better match the intended construction scaffold.
VI. C ONCLUSION This paper presented CIT-CAD, a Constraint Intent Treebased framework for natural-language CAD code generation and verification. The central idea is to make construction intent explicit before generating code: CIT-CAD infers a treestructured representation of expected CAD entities, operations, sketch attributes, Boolean roles, and inter-entity relations, then uses this representation to guide CadQuery generation and deterministic validation. By comparing CIT-derived expected constraints with constraints extracted from generated programs, CIT-CAD provides a program-level view of CAD correctness that complements final-shape IoU. Our experiments show that CIT-CAD improves the quality of generated
CAD programs from both geometric-level and constraint-level across different LLM backbones. Future work will extend the constraint detector vocabulary to cover more fine-grained construction relations. We will also improve parameter-aware CIT inference and develop repair models that can better target local sketch and relation errors.
R EFERENCES [1] M. S. Khan, S. Sinha, T. U. Sheikh, D. Stricker, S. A. Ali, and M. Z. Afzal, “Text2cad: Generating sequential cad designs from beginner-toexpert level text prompts,” Advances in Neural Information Processing Systems 37, pp. 7552–7579, 2024. [2] A. C. Doris, M. F. Alam, A. Heyrani Nobari, and F. Ahmed, “Cad-coder: An open-source vision-language model for computer-aided design code generation,” in International Design Engineering Technical Conferences and Computers and Information in Engineering Conference, 2025. [3] Y. Guan, X. Wang, X. Xing, J. Zhang, D. Xu, and Q. Yu, “Cad-coder: Text-to-cad generation with chain-of-thought and geometric reward,” arXiv preprint arXiv:2505.19713, 2025. [4] C. He, S. Zhang, L. Zhang, and J. Miao, “Cad-coder: Text-guided cad files code generation,” arXiv preprint arXiv:2505.08686, 2025. [5] K. Alrashedy, P. Tambwekar, Z. H. Zaidi, M. Langwasser, W. Xu, and M. C. Gombolay, “Generating CAD code with vision-language models for 3d designs,” in The Thirteenth International Conference on Learning Representations, 2025. [6] Z. Zhou, J. Han, L. Du, N. Fang, L. Qiu, and S. Zhang, “Cad-judge: Toward efficient morphological grading and verification for text-to-cad generation,” in ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1021–1025, IEEE, 2026. [7] M. F. Alam and F. Ahmed, “Gencad: Image-conditioned computer-aided design generation with transformer-based contrastive representation and diffusion priors,” arXiv preprint arXiv:2409.16294, 2024. [8] K. Niu, H. Yu, Z. Chen, M. Zhao, T. Fu, B. Li, and X. Xue, “From intent to execution: Multimodal chain-of-thought reinforcement learning for precise cad code generation,” arXiv preprint arXiv:2508.10118, 2025. [9] S. Koch, A. Matveev, Z. Jiang, F. Williams, A. Artemov, E. Burnaev, M. Alexa, D. Zorin, and D. Panozzo, “Abc: A big cad model dataset for geometric deep learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9601–9611, 2019. [10] K. D. Willis, Y. Pu, J. Luo, H. Chu, T. Du, J. G. Lambourne, A. SolarLezama, and W. Matusik, “Fusion 360 gallery: A dataset and environment for programmatic cad construction from human design sequences,” ACM Transactions on Graphics (TOG), vol. 40, no. 4, pp. 1–24, 2021. [11] R. Wu, C. Xiao, and C. Zheng, “Deepcad: A deep generative network for computer-aided design models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6772–6782, 2021. [12] A. Seff, Y. Ovadia, W. Zhou, and R. P. Adams, “Sketchgraphs: A large-scale dataset for modeling relational geometry in computer-aided design,” arXiv preprint arXiv:2007.08506, 2020. [13] G. Sharma, R. Goyal, D. Liu, E. Kalogerakis, and S. Maji, “CSGNet: Neural shape parser for constructive solid geometry,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. [14] R. K. Jones, T. Barton, X. Xu, K. Wang, E. Jiang, P. Guerrero, N. J. Mitra, and D. Ritchie, “ShapeAssembly: Learning to generate programs for 3d shape structure synthesis,” in ACM SIGGRAPH Asia 2020 Conference Papers, 2020. [15] Y. Tian, A. Luo, X. Sun, K. Ellis, W. T. Freeman, J. B. Tenenbaum, and J. Wu, “Learning to infer and execute 3d shape programs,” in International Conference on Learning Representations, 2019. [16] Y. Ganin, S. Bartunov, Y. Li, E. Keller, and S. Saliceti, “Computer-aided design as language,” arXiv preprint arXiv:2105.02769, 2021. [17] A. S. Karadeniz, D. Mallis, N. Mejri, K. Cherenkova, A. Kacem, and D. Aouada, “DAVINCI: A single-stage architecture for constrained CAD sketch inference,” in Proceedings of the British Machine Vision Conference, 2024. [18] P. Govindarajan, D. Baldelli, J. Pathak, Q. Fournier, and S. Chandar, “Cadmium: Fine-tuning code language models for text-driven sequential cad design,” arXiv preprint arXiv:2507.09792, 2025. [19] J. Liao, J. Xu, Y. Sun, M. Tang, S. He, J. Liao, S. Yu, Y. Li, and X. Guan, “Automated cad modeling sequence generation from text descriptions via transformer-based large language models,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 21720–21748, 2025. [20] W. Ma, S. Chen, Y. Lou, X. Li, and X. Zhou, “Draw step by step: Reconstructing cad construction sequences from point clouds via multimodal diffusion.,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 27154–27163, 2024.
[21] D. Rukhovich, E. Dupont, D. Mallis, K. Cherenkova, A. Kacem, and D. Aouada, “Cad-recode: Reverse engineering cad code from point clouds,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9801–9811, 2025. [22] J. Xu, C. Wang, Z. Zhao, W. Liu, Y. Ma, and S. Gao, “Cad-mllm: Unifying multimodality-conditioned cad generation with mllm,” arXiv preprint arXiv:2411.04954, 2024. [23] X. Wang, J. Zheng, Y. Hu, H. Zhu, Q. Yu, and Z. Zhou, “From 2d CAD drawings to 3d parametric models: A vision-language approach,” in Proceedings of the AAAI Conference on Artificial Intelligence, pp. 7961– 7969, 2025. [24] Z. Yuan, J. Shi, and Y. Huang, “Openecad: An efficient visual language model for editable 3d-cad design,” Computers & Graphics, vol. 124, p. 104048, 2024. [25] Z. Zhao, S. Wang, J. Gu, Y. Zhu, L. Mei, Z. Zhuang, Z. Cui, Q. Wang, and D. Shen, “Chatcad+: Toward a universal and reliable interactive cad using llms,” IEEE Transactions on Medical Imaging, vol. 43, no. 11, pp. 3755–3766, 2024. [26] Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of the Association for Computational Linguistics, pp. 1536–1547, 2020. [27] W. Zheng, S. Sharan, A. K. Jaiswal, K. Wang, Y. Xi, D. Xu, and Z. Wang, “Outline, then details: Syntactically guided coarse-to-fine code generation,” in International Conference on Machine Learning, pp. 42403–42419, 2023. [28] J. Li, G. Li, Y. Li, and Z. Jin, “Structured chain-of-thought prompting for code generation,” ACM Transactions on Software Engineering and Methodology, vol. 34, no. 2, pp. 1–23, 2025. [29] G. Yang, Y. Zhou, X. Chen, X. Zhang, T. Y. Zhuo, and T. Chen, “Chainof-thought in neural code generation: From and for lightweight language models,” IEEE Transactions on Software Engineering, 2024. [30] J. Li, C. Tao, J. Li, G. Li, Z. Jin, H. Zhang, Z. Fang, and F. Liu, “Large language model-aware in-context learning for code generation,” ACM Transactions on Software Engineering and Methodology, vol. 34, no. 7, pp. 1–33, 2025. [31] Y. Du, H. Sun, and M. Li, “Post-incorporating code structural knowledge into pretrained models via icl for code translation,” IEEE Transactions on Software Engineering, vol. 51, no. 11, pp. 3038–3055, 2025. [32] D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, A. Svyatkovskiy, S. Fu, M. Tufano, S. K. Deng, C. B. Clement, D. Drain, N. Sundaresan, J. Yin, D. Jiang, and M. Zhou, “Graphcodebert: Pre-training code representations with data flow,” in 9th International Conference on Learning Representations, 2021. [33] Y. Du and Z. Yu, “Pre-training code representation with semantic flow graph for effective bug localization,” in Proceedings of the 31st ACM joint European software engineering conference and symposium on the foundations of software engineering, pp. 579–591, 2023. [34] Y.-F. Ma, Y. Du, and M. Li, “Capturing the long-distance dependency in the control flow graph via structural-guided attention for bug localization.,” in IJCAI, pp. 2242–2250, 2023. [35] Y. Du, Y. Li, Y.-F. Ma, and M. Li, “Capturing the context-aware code change via dynamic control flow graph for commit message generation,” Machine Learning, vol. 114, no. 4, p. 94, 2025. [36] A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, and J. M. Zhang, “Large language models for software engineering: Survey and open problems,” in 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), pp. 31–53, 2023. [37] C. Gao, X. Hu, S. Gao, X. Xia, and Z. Jin, “The current challenges of software engineering in the era of large language models,” ACM Transactions on Software Engineering and Methodology, vol. 34, no. 5, pp. 127:1–127:30, 2025. [38] J. Shi, Z. Yang, and D. Lo, “Efficient and green large language models for software engineering: Literature review, vision, and the road ahead,” ACM Transactions on Software Engineering and Methodology, vol. 34, no. 5, pp. 137:1–137:22, 2025. [39] W. Xu, K. Gao, H. He, and M. Zhou, “LicoEval: Evaluating LLMs on license compliance in code generation,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 1665– 1677, 2025. [40] Y. Zhang, Y. Xie, S. Li, K. Liu, C. Wang, Z. Jia, X. Huang, J. Song, C. Luo, Z. Zheng, R. Xu, Y. Liu, S. Zheng, and X. Liao, “Unseen horizons: Unveiling the real capability of LLM code generation beyond
the familiar,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 604–615, 2025. [41] A. Al-Kaswan, S. Deatc, B. Koç, A. van Deursen, and M. Izadi, “Code red! on the harmfulness of applying off-the-shelf large language models to programming tasks,” Proceedings of the ACM on Software Engineering, vol. 2, no. FSE, pp. 2477–2499, 2025. [42] Y. Wang, X. Li, T. N. Nguyen, S. Wang, C. Ni, and L. Ding, “Natural is the best: Model-agnostic code simplification for pre-trained large language models,” Proceedings of the ACM on Software Engineering, vol. 1, no. FSE, pp. 586–608, 2024. [43] Z. Yang, F. Liu, Z. Yu, J. W. Keung, J. Li, S. Liu, Y. Hong, X. Ma, Z. Jin, and G. Li, “Exploring and unleashing the power of large language models in automated code translation,” Proceedings of the ACM on Software Engineering, vol. 1, no. FSE, pp. 1585–1608, 2024. [44] H. Yu, B. Shen, D. Ran, J. Zhang, Q. Zhang, Y. Ma, G. Liang, Y. Li, Q. Wang, and T. Xie, “CoderEval: A benchmark of pragmatic code generation with generative pre-trained models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), pp. 37:1–37:12, 2024. [45] X. Du, M. Liu, K. Wang, H. Wang, J. Liu, Y. Chen, J. Feng, C. Sha, X. Peng, and Y. Lou, “Evaluating large language models in class-level code generation,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), pp. 81:1–81:13, 2024. [46] Y. Du, Y.-F. Ma, Z. Xie, and M. Li, “Beyond lexical consistency: Preserving semantic consistency for program translation,” in 2023 IEEE International Conference on Data Mining (ICDM), pp. 91–100, IEEE, 2023. [47] Y. Du, H. Sun, and M. Li, “A joint learning model with variational interaction for multilingual program translation,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, pp. 1907–1918, 2024. [48] L. Xin-Ye, D. Ya-Li, and L. Ming, “Enhancing llms in long code translation through instrumentation and program state alignment,” arXiv preprint arXiv:2504.02017, 2025. [49] F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, A. Guha, M. Greenberg, and A. Jangda, “MultiPL-E: A scalable and polyglot approach to benchmarking neural code generation,” IEEE Transactions on Software Engineering, vol. 49, no. 7, pp. 3675–3691, 2023. [50] A. Mastropaolo, L. Pascarella, E. Guglielmi, M. Ciniselli, S. Scalabrino, R. Oliveto, and G. Bavota, “On the robustness of code generation techniques: An empirical study on GitHub Copilot,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pp. 2149– 2160, 2023. [51] T. Zhang, T. Yu, T. Hashimoto, M. Lewis, W.-t. Yih, D. Fried, and S. Wang, “Coder reviewer reranking for code generation,” in International Conference on Machine Learning, pp. 41832–41846, 2023. [52] S. Fakhoury, A. Naik, G. Sakkas, S. Chakraborty, and S. K. Lahiri, “Llmbased test-driven interactive code generation: User study and empirical evaluation,” IEEE Transactions on Software Engineering, 2024. [53] J. Wang, Y. Huang, C. Chen, Z. Liu, S. Wang, and Q. Wang, “Software testing with large language models: Survey, landscape, and vision,” IEEE Transactions on Software Engineering, vol. 50, no. 4, pp. 911–936, 2024. [54] Z. Yuan, M. Liu, S. Ding, K. Wang, Y. Chen, X. Peng, and Y. Lou, “Evaluating and improving ChatGPT for unit test generation,” Proceedings of the ACM on Software Engineering, vol. 1, no. FSE, pp. 1703–1726, 2024. [55] Y. Tang, Z. Liu, Z. Zhou, and X. Luo, “ChatGPT vs SBST: A comparative assessment of unit test suite generation,” IEEE Transactions on Software Engineering, vol. 50, no. 6, pp. 1340–1359, 2024. [56] F. Tip, J. Bell, and M. Schäfer, “LLMorpheus: Mutation testing using large language models,” IEEE Transactions on Software Engineering, vol. 51, no. 6, pp. 1645–1665, 2025. [57] Y. Deng, C. S. Xia, H. Peng, C. Yang, and L. Zhang, “Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,” in Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), pp. 423– 435, 2023. [58] Y. Shang, Q. Zhang, C. Fang, S. Gu, J. Zhou, and Z. Chen, “A large-scale empirical study on fine-tuning large language models for unit testing,” Proceedings of the ACM on Software Engineering, vol. 2, no. ISSTA, pp. 1678–1700, 2025. [59] C. Zhang, Y. Zheng, M. Bai, Y. Li, W. Ma, X. Xie, Y. Li, L. Sun, and Y. Liu, “How effective are they? exploring large language model
based fuzz driver generation,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), pp. 1223–1235, 2024. [60] H. Sun, Y.-J. Zhang, Z. Xie, R.-B. Liu, Y. Du, X.-Y. Li, and M. Li, “Aces: Who tests the tests? leave-one-out auc consistency for code generation,” arXiv preprint arXiv:2604.03922, 2026. [61] C. S. Xia, M. Paltenghi, J. L. Tian, M. Pradel, and L. Zhang, “Fuzz4All: Universal fuzzing with large language models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), pp. 126:1–126:13, 2024. [62] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-refine: Iterative refinement with self-feedback,” arXiv preprint arXiv:2303.17651, 2023. [63] Q. Guo, J. Cao, X. Xie, S. Liu, X. Li, B. Chen, and X. Peng, “Exploring the potential of ChatGPT in automated code refinement: An empirical study,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), pp. 34:1–34:13, 2024. [64] Y. Peng, S. Gao, C. Gao, Y. Huo, and M. R. Lyu, “Domain knowledge matters: Improving prompts with fix templates for repairing python type errors,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE), pp. 4:1–4:13, 2024. [65] I. Bouzenia, P. Devanbu, and M. Pradel, “RepairAgent: An autonomous, LLM-based agent for program repair,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 2188–2200, 2025. [66] N. Parasaram, H. Yan, B. Yang, Z. Flahy, A. Qudsi, D. Ziaber, E. T. Barr, and S. Mechtaev, “The fact selection problem in LLM-based program repair,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 2574–2586, 2025. [67] K. Huang, J. Zhang, X. Meng, and Y. Liu, “Template-guided program repair in the era of large language models,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 1895– 1907, 2025. [68] R. Ehsani, E. Parra, S. Haiduc, and P. Chatterjee, “Hierarchical knowledge injection for improving LLM-based program repair,” in 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 1440–1452, 2025. [69] G. Li, C. Zhi, J. Chen, J. Han, and S. Deng, “Exploring parameterefficient fine-tuning of large language model on automated program repair,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 719–731, 2024. [70] K. Huang, X. Meng, J. Zhang, Y. Liu, W. Wang, S. Li, and Y. Zhang, “An empirical study on fine-tuning large language models of code for automated program repair,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 1162–1174, 2023. [71] F. Li, J. Jiang, J. Sun, and H. Zhang, “Hybrid automated program repair by combining large language models and program analysis,” ACM Transactions on Software Engineering and Methodology, vol. 34, no. 7, pp. 202:1–202:28, 2025. [72] K. Huang, J. Zhang, X. Bao, X. Wang, and Y. Liu, “Comprehensive finetuning large language models of code for automated program repair,” IEEE Transactions on Software Engineering, vol. 51, no. 4, pp. 904–928, 2025. [73] A. Z. H. Yang, S. Kolak, V. J. Hellendoorn, R. Martins, and C. Le Goues, “Revisiting unnaturalness for automated program repair in the era of large language models,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 2561–2573, 2025. [74] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [75] D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. Li, et al., “Deepseek-coder: When the large language model meets programming–the rise of code intelligence,” arXiv preprint arXiv:2401.14196, 2024. [76] OpenAI, “Gpt-5-mini.” https://cdn.openai.com/gpt-5-system-card.pdf, 2025. Online. Accessed 31 December 2025.
Supplementary Material Constraint Intent Tree-based CAD code Generation and Validation
This appendix provides supplementary details for reproducing and interpreting CIT-CAD. It first reports the prompt templates used for CIT extraction, CIT-guided CAD code generation, and constraint-guided repair. It then formalizes the short-constraint validation procedure, describes the dataset preprocessing and parameter-augmentation pipeline, summarizes the experimental configuration, and discusses the main threats to validity. VII. P ROMPT T EMPLATES This appendix provides the inference-time prompt templates used by CIT-CAD. We report the prompt content at the level needed for reproduction while omitting private deployment details such as API keys, local endpoints, and machine-specific paths. CIT-CAD uses two prompt stages: an initial generation stage that first extracts the Constraint Intent Tree and then generates CadQuery code conditioned on the tree, and a repair stage that revises the generated code using localized constraintvalidation feedback. A. Constraint Intent Tree Extraction The initial generation stage contains two calls. The first call extracts a Constraint Intent Tree from the natural-language description. The second call generates executable CadQuery code from the description and the extracted tree. CIT Extraction Prompt System Prompt: You extract a CAD structure tree from a natural-language CAD description. Instruction: Output requirements: 1. Return only valid JSON. Do not include Markdown fences. 2. Include the sample id, the original description, and a root node. 3. Use stable snake-case ids and expected_variable names. 4. Create one child node for every physical solid or cutting tool. 5. For each entity, extract sketch_type, is_connected, is_open_profile, and boolean_role when supported by the description. 6. For relations, extract intersection, contact_type, has_coplanar_faces, and has_shared_axis when they are explicitly implied or strongly entailed. 7. Use null values for uncertain entity constraints rather than guessed geometry.
8. Use boolean_role=true for additive solids and boolean_role=false for subtractive cutting tools. 9. Use sketch_type values from rect, circle, with_arc, only_line, and other. User Input: Sample id: {sample id} Natural-language description: {description}
B. CIT-Guided CAD Code Generation CIT-Guided CAD Code Generation Prompt System Prompt: You generate executable CadQuery Python code from a natural-language CAD description and a Constraint Intent Tree. Instruction: Output requirements: 1. Return only Python code. Do not include Markdown fences. 2. Import cadquery as cq. 3. Create one variable for every non-group node in the Constraint Intent Tree. 4. Use each node’s expected_variable exactly as the variable name. 5. Each entity variable should be produced by a Workplane sketch followed by .extrude(...). 6. Compose entities using .union(...), .cut(...), or .intersect(...) according to the intended construction. 7. Assign the final object to a variable named solid. 8. Prefer simple explicit dimensions when the description is underspecified. 9. Preserve the entity constraints and relation constraints in the Constraint Intent Tree as much as possible. 10. Do not call .close() after .rect(...), .circle(...), .ellipse(...), or helpers that already create closed wires. 11. Use .close() only after explicit open chains such as .moveTo(...).lineTo(...).lineTo(...). User Input: Natural-language description: {description} Constraint Intent Tree: {tree json}
C. Constraint-Guided Repair The repair stage receives the natural-language description, the inferred CIT, the current generated code, and deterministic
constraint-validation feedback. The feedback identifies violated entity or relation constraints, their localized CIT paths, and constraints that have already been satisfied and should be preserved. Constraint-Guided Repair Prompt System Prompt: You repair CadQuery Python code using localized Constraint Intent Tree validation feedback. Instruction: Output requirements: 1. Return only the revised Python code. Do not include Markdown fences. 2. Keep the same intended object and preserve variable names from the Constraint Intent Tree. 3. Focus on the violated entity constraints, relation constraints, or subtrees listed in the feedback. 4. The final object must be assigned to a variable named solid. 5. Preserve constraints that were already satisfied in previous validation rounds. 6. Prefer localized edits to the violated entities, relations, or subtrees. 7. A candidate repair must preserve satisfied constraints and reduce the number of violated constraints. 8. Do not call .close() after .rect(...), .circle(...), .ellipse(...), or helpers that already create closed wires. 9. Use .close() only after explicit open chains such as .moveTo(...).lineTo(...).lineTo(...). 10. If feedback mentions ‘‘Cannot convert object type ... Wire ... to vector’’, remove the extra .close() before .extrude(...). User Input: Natural-language description: {description} Constraint Intent Tree: {tree json} Current code: {code} Constraint-validation feedback: {feedback json}
VIII. C ONSTRAINT E XTRACTION AND V ERIFICATION D ETAILS CIT-CAD converts both expected and generated CAD artifacts into the same short-constraint representation before verification. The expected representation is produced from the CIT. For each entity node, the converter reads the serializable implementation keys sketch type, is connected, is open profile , and boolean role when they are present and non-null. For each relation attached to a node, the converter keeps relation types in { intersection , contact type , has coplanar faces , has shared axis}, sorts the two participating entity names, and stores a triple containing entities and value. A node map
Algorithm 2: Short-constraint validation Input: Expected constraints CT , actual constraints CP Output: Validation result, satisfied IDs, violated IDs 1 S ← ∅, V ← ∅; 2 foreach expected entity n in CT do 3 if n is missing from CP then 4 V ← V ∪ {entity : n : exists}; 5 continue; 6 end 7 foreach expected field f of n do 8 if CP [n][f ] = CT [n][f ] then 9 S ← S ∪ {entity : n : f }; 10 end 11 else 12 V ← V ∪ {entity : n : f }; 13 end 14 end 15 end 16 foreach relation type g do 17 Ag ← canonicalized actual relation triples of type g; 18 foreach expected triple (g, u, y) do 19 if (g, sort(u), y) ∈ Ag then 20 S ← S ∪ {relation : g : u : y}; 21 end 22 else 23 V ← V ∪ {relation : g : u : y}; 24 end 25 end 26 end 27 return V = ∅, S, V ;
is also stored so each expected constraint can later be traced back to a CIT node path. The generated program is analyzed by a deterministic CadQuery analyzer. Static AST analysis identifies variables assigned from . extrude (...) expressions, recovers sketch construction calls, classifies local sketch fields, and infers Boolean roles from composition expressions such as union, cut, and intersect . Runtime geometric analysis executes the generated program and inspects the resulting solids to extract relation constraints. The analyzer then strips metadata and emits the same short representation used for expected constraints. Algorithm 2 mirrors the implemented validation procedure. Entity constraints are checked by exact field equality. Relation constraints are checked by membership in a canonicalized relation set. This intentionally avoids fuzzy matching: if the CIT expects a face-contact relation and the generated code produces no contact or a different contact type, the relation is counted as violated even if the final geometry has non-zero IoU with the reference model.
IX. DATASET P REPROCESSING The evaluation set is derived from Text2CAD [1]. We first join the natural-language CSV and CadQuery CSV by normalized sample ID, remove samples with missing descriptions or missing CAD programs, and remove duplicate naturallanguage descriptions. This yields 151K valid text-to-CAD pairs, as reported in the main paper. CIT-CAD is designed for construction-aware multi-entity generation. We therefore filter the valid pairs using an implementation-level entity parser. The parser counts explicit extruded construction entities in the reference CAD program and keeps only samples with at least two entities. To ensure that the natural-language description exposes the intended construction complexity, the preprocessing script also estimates the entity count mentioned in the description using numeric and ordinal expressions and keeps samples whose natural-language entity count matches the code entity count. Samples whose reference programs cannot be parsed by the entity parser are excluded from the final selected set. After preprocessing, the final evaluation subset contains 26,783 samples. The natural-language inputs used by the final experiments include parameter augmentation. The augmentation script prompts an LLM to extract compact parameter summaries from the held-out reference CadQuery code, including visible dimensions, radii, heights, translations, rotations, entity counts, and Boolean composition. The summary is appended to the natural-language description before generation. This step addresses a limitation of the original descriptions, which often describe the object qualitatively but omit numeric parameters needed for executable CadQuery code. The reference code is not provided to the code-generation model at evaluation time; only the augmented natural-language description is used as input. X. E XPERIMENTAL C ONFIGURATION All evaluated methods consume the same augmented natural-language descriptions. For each backbone LLM, the Vanilla baseline directly prompts the model to generate a standalone CAD program and does not use CITs, constraints, or repair. CIT-CAD runs the complete pipeline: CIT extraction, CIT-conditioned code generation, deterministic validation, localized repair, and monotonic candidate acceptance. The open-source backbones are served through vLLMcompatible OpenAI-style chat endpoints. The Qwen setting uses Qwen3-30B-A3B-Instruct, and the DeepSeek setting uses DeepSeek-Coder-V2-Lite-Instruct in the reported experiments. GPT-5.4-mini is accessed through an OpenAI-compatible closed-source API. All LLM calls use chat-completion format with temperature 1.0 when the endpoint supports temperature; if an endpoint rejects the temperature argument, the implementation automatically retries without it. The maximum retry count is three. Dataset-level generation and repair are parallelized across samples, with 16 workers used by default and 32 workers used in some large closed-source runs.
The maximum number of repair iterations is controlled by −−max−reflexion− iters . The default implementation uses two repair iterations, while some extended GPT runs use five iterations. Existing generated samples are skipped unless explicit regeneration is requested. This skip policy is used to support long-running experiments and resume interrupted generations without overwriting completed samples. For geometry-level evaluation, generated and reference solids are normalized to the same voxel grid before IoU computation, and invalid or missing generated programs are counted as failures for validity and success-rate metrics. XI. T HREATS TO VALIDITY We discuss the main factors that may affect the interpretation and generalizability of our results. These threats mainly arise from the quality of the inferred CITs, the coverage of the deterministic constraint detectors, the construction-aware dataset filtering strategy, and the execution-based nature of CAD code evaluation. a) CIT extraction quality.: CIT-CAD treats the inferred CIT as the explicit design intent extracted from the naturallanguage input. If the extraction model misses an entity or infers an incorrect relation, the verifier will faithfully check the wrong intent. We reduce this risk by using a fixed output format, requiring null values for uncertain constraints, and evaluating across multiple LLM backbones, but the quality of CIT extraction remains an important source of variance. b) Constraint detector coverage.: CSR measures satisfaction of the constraints currently supported by the deterministic analyzer. The implementation covers entity-level sketch fields, Boolean role, intersection, contact type, coplanar faces, and shared axes. It does not yet cover all possible CAD design intents, such as precise symmetry, tolerances, manufacturing constraints, or arbitrary parametric dependencies. CSR should therefore be interpreted as construction-intent correctness under the implemented vocabulary, not as a complete proof of CAD equivalence. c) Dataset and parameter augmentation.: The final dataset is filtered to multi-entity samples whose naturallanguage entity count matches the code-derived entity count. This selection is appropriate for evaluating construction-aware generation, but it may underrepresent single-solid designs or descriptions with implicit entities. Parameter augmentation improves executability by making numeric details available in the natural-language input, but it also means the evaluated setting is closer to parameter-aware specification following than to generation from purely qualitative prompts. d) Execution-based evaluation.: Valid syntax rate, IoU, and CSR all depend on successful CadQuery execution and deterministic geometry analysis. Runtime failures are treated as invalid outputs and can affect both geometry-level and constraint-level metrics. This is intentional for CAD code generation, where executable code is the target artifact, but it may penalize models for small API-level errors even when the intended shape is partially recoverable.