Conceptio › Archive › arXiv CS
arXiv CSopen access

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

arXiv:2609.05364v1 [cs.PL] 4 Sep 2026

Samuel Kushnir1 Kimia Noorbakhsh2 Kavya Sreedhar4 Liqun Cheng4 Ming Liu4 Parthasarathy Ranganathan4 Mohammad Alizadeh2 Fred Kjolstad3 Suvinay Subramanian1 1 Google DeepMind 2 MIT 3 Stanford 4 Google

Abstract Machine-learning performance modeling is a uniquely hostile terrain for longlived software: the assumptions baked into today’s abstractions are invalidated by tomorrow’s models and systems, forcing perpetual refactoring of performancemodeling frameworks. Meanwhile, AI coding agents have become fast and capable enough that regenerating an entire library is cheaper than paying down the tech debt of incrementally patching it. We describe SMART, a rigorous symbolic performancemodeling library for ML systems whose main branch contains almost no code: the repository is a DAG of self-contained natural- language design docs, coding subagents regenerate the implementation from only the docs on new version updates, and every human change is a natural-language edit to a doc—self-documenting by construction. Two ingredients make regeneration reliable: (i) a design-doc style built around step-by-step worked examples that act as in-context demonstrations for the generating agents, and (ii) a minimal, recursively defined operator IR with symbolic (SymPy) cost expressions, a fast analytical roll-up mode for large sweeps, and a slow modulo-scheduling mode for fine-grained schedule studies. Regenerated implementations reproduce hand-audited reference models—including DeepSeekV3 serving on a TPU pod slice—to round-off precision, suggesting that design docs—not code—can be the durable artifact for ML-systems co-design tools.

1

Introduction

Machine-learning performance modeling is a punishing environment for software abstractions, because the assumptions of current architectures and systems are constantly evolving. An abstraction that seemed reasonable a year ago—for instance, that every transformer layer looks the same—is suddenly false in the face of mixture-of-experts routing, latent attention, and heterogeneous inference phases. Performance-modeling frameworks sit at the intersection of the two fastest-moving parts of the stack—rapidly evolving model architectures on top, and underlying hardware accelerators and interconnects on the bottom—so they absorb this churn from both directions; the result is a never-ending stream of refactorings and hacks to absorb each new architecture or system change into an aging framework. At the same time, AI coding tools have become remarkably capable and fast. Together, these two forces produce failure modes that compound technical debt: 1. Incremental generation debt (predates AI). Let St denote the specification of a library at time t and let G denote a code generator—a human engineer or an agent. A greenfield build computes Ct = G(St ). In practice, however, the generator at time t+1 is handed the previous implementation as an extra argument and computes Ct+1 = G(St+1 , Ct ). We can define the technical debt of this process as Dt+1 =

G(St+1 , Ct ) − G(St+1 ) ,

(1)

human edits: natural language only doc

doc

doc

doc

doc

agentic regeneration

generated library

one sub-agent per doc, topological order

symbolic cost model (build product)

validation

design-doc DAG (checked into master) repair

reference-model reconciliation, parameter guards, unit tests

Figure 1: The regeneration workflow: sub-agents regenerate the implementation from the design-doc DAG in dependency order, and the result must pass reconciliation against hand-built references before it replaces the previous build. Humans only ever edit the docs. the distance between the incrementally patched system and the one that would have been built from the current spec alone. Because starting over—deleting all of one’s code—is mentally challenging and time-consuming for humans, the technical debt is in practice much larger than zero: incremental patches inevitably introduce compromises that compound with every spec revision. 2. Context-window myopia AI coding agents are constrained by finite context windows, making it impossible to pass an entire mature codebase into a single prompt. Consequently, developers must feed the agent fragmented, localized code snippets when requesting revisions. Because the agent cannot see the whole picture—missing global invariants, crossmodule dependencies, and the broader architectural intent—it frequently generates code that is locally plausible but globally sub-optimal. Over time, these piecemeal, context-blind updates degrade the structural coherence of the framework. To solve these problems while exploiting the speed of AI coding agents, we make natural-language design docs the durable artifact and treat code as a regenerable build product (Section 2). Reliable regeneration imposes real demands on how docs are written (Section 2) and how the modeled system is factored into abstractions (Section 3). Because Ct = G(St ) is recomputed from scratch on a regular cadence, Equation (1) is driven to zero by construction. This complete regeneration workflow is illustrated in Figure 1.

2

Design docs as the source of truth

Our main branch contains almost no code. Instead, it consists of a folder structure of design document markdown files that compose into a directed acyclic graph (DAG). Some modules must be generated before others—for example, hardware topology and numerics precede the collective-cost models, which in turn precede the model catalog. The edges of this DAG are machine-discovered rather than hand-maintained. We utilize a distributed process where read-only agents analyze the design documents and infer the dependency edges between them. Once the graph is resolved, an orchestrator agent walks the DAG in topological order, assigning a dedicated coding sub-agent to implement each self-contained doc. Throughout this process, the orchestrator maintains a central log file recording areas where sub-agents struggled to interpret the prose, as well as bugs uncovered in the output of agents from earlier topological waves. Factoring the codebase into self-contained documents and deploying a distinct sub-agent for each yields three distinct benefits: 1. Bounded context windows. By restricting the scope of each generation step to a single document, we limit the context window and task length for each individual LLM. This avoids overwhelming the model, significantly increasing the probability of correct code generation for each isolated task. 2. Targeted human iteration. The orchestrator’s log provides fine-grained visibility into the “hardness” of each design doc. By observing where agents struggle the most and introduce the most bugs, human engineers know exactly which documents require prose refinement and clarification for the next version. 2

3. Dynamic model routing. The orchestrator can dynamically allocate varying levels of intelligence based on document complexity. While foundational documents (such as the core DSL design) might be routed to a larger, more capable LLM, many downstream design docs can be successfully translated into working code using smaller, more cost-effective models. Cost and iteration speed. In practice, a full clean-slate regeneration of the library takes between 1.5 and 3 hours. The API cost for a complete rebuild using Claude Code is around 100 USD and under a standard high-tier usage plan (e.g., Claude Max), this represents roughly 20% of a weekly usage budget, making continuous, full-library regeneration both practical and economically viable. An opinionated way of writing design docs Conventionally, in the age of AI, practitioners have focused on writing the right tests and high-level rules to ensure code correctness and adherence to project conventions. This is, in essence, a top-down approach—a constitution the generator must obey. We believe a complementary, bottom-up ingredient is crucial: worked examples. Writing out how a piece of pseudo-code executes on a given input, step by step—intermediate shapes, intermediate values, the exact closed-form cost expression that should result—is what maintains consistency across independent agentic code-generations. Like a human learner, in-context learning benefits from a tangible walk-through [Brown et al., 2020, Dong et al., 2022]: a concrete trace pins down semantics that prose alone leaves ambiguous, directly attacking failure mode 2 of Section 1. Our docs therefore favor executable-in-your-head vignettes (“on a 2×2×2 torus with wraparound, the per-node link count is 3, not 6; the all-gather of V bytes therefore costs . . . ”) and every number-bearing doc ends with a reconciliation anchor: a small preset whose expected outputs are stated exactly and enforced by generated tests.

3

A minimal symbolic IR for performance co-design

Reliable regeneration also constrains the artifact being specified: the abstractions must be few, orthogonal, and stable under architecture churn. We use a flexible and minimal IR with both a fast mode for large sweeps and a slow mode for careful scheduling studies. The Op abstraction. Op: inputs: List[Tensor] outputs: List[Tensor] cost: OpCost rrt: RRT

params:

A model is defined by a single recursively defined operation: # symbolic shapes

# SymPy exprs per key cost (compute, memory, comm) # resource reservation table: rows = resources, # cols = cycles, cell = units used, # e.g. ("MXU", cycle 3) -> 1 Union[InnerLoop(n_iter, body: Graph), # loop nest GraphParams(graph), # subgraph LeafParams(...)] # leaf op

An Op is therefore either an interior node—a loop with a trip count and a body graph, or a plain subgraph—or a leaf. The leaf nodes are where the software and system sides meet: a leaf is an op for which the system specifies an RRT and an OpCost. We target TPUs [Jouppi et al., 2023], so the leaves are TPU-shaped: an MXU matmul tile (mxu_op), a VMEM tile load (load_tile_to_vmem), or an ICI collective (allgather). The algorithm side composes leaves into loop nests; the system side prices them. Swapping either side—a new attention variant, a new interconnect generation—touches only its own docs. The builder DSL. Models are authored in a thin Python-embedded tracing DSL and never construct Op nodes by hand: a decorated block traces into a named subgraph, decorated loops become InnerLoop nodes (@smart_loop is a true reduction with a carried accumulator, @smart_map_loop a parallel map), and builder calls emit system-priced leaves. Every dimension is a SymPy symbol, so trip counts like Tq /qblk stay symbolic and one trace serves the whole design space. Listing 1 shows the (lightly condensed) flash-attention core: the (B, H, Tq , Tkv ) score matrix never leaves VMEM, and the asymmetric Tq /Tkv make the same nest serve prefill (Tq =Tkv =T ) and flash-decoding (Tq =1, Tkv =Tctx ). 3

Listing 1: the flash-attention core in the builder DSL. def flash_attention_core(b, q, k, v, *, B, H, T_q, T_kv, hd_qk, hd_v, q_blk, kv_blk): @smart_map_loop(b, n_iter=T_q/q_blk, name="q_blocks", gather_axis=2) def q_block(): # parallel MAP over query tiles q_i = b.load_tile(q, (B, H, q_blk, hd_qk)) acc0 = b.load_tile(v, (B, H, q_blk, hd_v)) @smart_loop(b, n_iter=T_kv/kv_blk, name="kv_blocks") def kv_block(o_acc): # REDUCTION over kv tiles k_t = b.load_tile(k, (B, H, kv_blk, hd_qk)) v_t = b.load_tile(v, (B, H, kv_blk, hd_v)) s = b.mxu_op("bhte,bhse->bhts", q_i, k_t, name="scores") pv = b.mxu_op("bhts,bhse->bhte", s, v_t, name="attn_v") return b.tensor_add(o_acc, pv, name="accumulate") return b.store_tile(kv_block(acc0)) # this tile’s (B,H,q_blk,hd_v) return q_block() # gathered (B, H, T_q, hd_v)

Distribution is expressed as sharding annotations, not hand-placed collectives: tensors name the mesh axes each dimension is sharded on, and a sharded-einsum wrapper infers the collectives from the operand/output shardings—a just-in-time AllGather of a sharded contracting weight, a ReduceScatter when an output is reduced over a sharded dimension. Only layout-moving collectives are explicit: in the DeepSeekMoE block, the dispatch all_to_all moves the expertparallel axis from the token-group dimension onto the expert dimension (the combine moves it back); every other collective in the block is inferred. Rolling up the Op hierarchy. Two modes turn a tree of per-op costs into wall-clock time. In fast mode, loops are rolled up coarsely: each leaf’s cost is scaled by the product of enclosing trip counts, and simple analytical schedulers model communication/computation overlap (weight pre-collection, all-to-all hiding) as composable transforms in the style of a roofline bound [Williams et al., 2009]; evaluation is closed-form, fast enough for sweeps over thousands of design points. In slow mode, each loop is modulo-scheduled [Rau, 1994] into its resource reservation table—software-pipelining the body against per-resource capacity—and the achieved initiation interval rolls up recursively up the tree, yielding dependency- and resource-aware schedules for the design points that sweeps flag as interesting. Symbolic propagation. All cost formulas are propagated upward symbolically and every roll-up produces a closed-form SymPy [Meurer et al., 2017] expression in the free variables of the design space (batch, sequence length, bandwidths, mesh axes, datatype widths). Numeric binding happens only at the edge—one substitution per design point—so a single symbolic build serves an entire sweep. A design doc can even state the exact expected expression for a collective’s cost, and the generated tests assert it.

4

Conclusion

Developers routinely use agents to modify code. We believe the SMART alternative approach of generating whole systems from design documentations has a lot of benefits and should increasingly be considered. SMART today comprises 50 design docs (∼9,000 lines of specification prose) spanning TPU topology, collective cost models, numerics, schedulers, and a catalog of frontier model families (dense, MoE, latent-attention, and robotics/VLA variants). We regenerate with every new version change: master is reduced to the docs plus a handful of leaf utilities, and the library is rebuilt by sub-agent orchestration. Treating design docs—rather than code—as the durable artifact pays the incremental-patching debt of Equation (1) down to zero at every regeneration and surfaces vague intent early, because a guess must be written into a doc to survive. The enablers—worked-example docs, a machine-discovered dependency DAG, and a minimal symbolic IR—should generalize wherever specs churn faster than software absorbs them.

References Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are 4

few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020. Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, et al. A survey on in-context learning. arXiv preprint arXiv:2301.00234, 2022. Norman P Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, et al. TPU v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. In Proceedings of the 50th Annual International Symposium on Computer Architecture, pages 1–14, 2023. Aaron Meurer, Christopher P Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K Moore, Sartaj Singh, et al. SymPy: symbolic computing in Python. PeerJ Computer Science, 3:e103, 2017. B Ramakrishna Rau. Iterative modulo scheduling: An algorithm for software pipelining loops. In Proceedings of the 27th Annual International Symposium on Microarchitecture, pages 63–74, 1994. Samuel Williams, Andrew Waterman, and David Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009.

5

Record · ID 660847 · SHA-256 98010c2ff6a519a0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.