ForgeTrain: Forging Production-Grade Training Frameworks via Harness-Driven AI Development Qingfeng He∗
Tsinghua University Beijing, China
Yaojian Chen
arXiv:2609.13645v1 [cs.SE] 12 Sep 2026
Tsinghua University Beijing, China
Leshan Li
Zhui Zhu∗
Tsinghua University Beijing, China
Haojun Sun
Shangzhan Li∗
Harbin Institute of Technology Harbin, China
Xu Chen
ModelBest Inc. Beijing, China
ModelBest Inc. Beijing, China
Tsinghua University Beijing, China
ModelBest Inc. Beijing, China
Yifei Shen
Changjingxing Zhao
Mengyuan Fan
Wenyu Guan
Yiyun Zheng
Peking University Beijing, China
ModelBest Inc. Beijing, China
Tsinghua University Beijing, China ModelBest Inc. Beijing, China
ModelBest Inc. Beijing, China
ModelBest Inc. Beijing, China
Zhen Li
Zhenghang Luo
Yuxuan Li†
Xu Han†
Zhiyuan Liu†
Yuxuan Zuo
Tsinghua University Beijing, China [email protected]
Tsinghua University Beijing, China [email protected]
1
Abstract
Training large models still relies on general-purpose frameworks such as Megatron-LM, whose generality tax constrains scenario-specific optimization and adds runtime overhead through accumulated abstraction. AI code generation reduces the cost of building a framework, and makes it affordable to forge one per scenario. We propose Forge Engineering: building a dedicated implementation from scratch for each scenario and iteratively optimizing it toward peak performance under correctness and usability constraints. Dedicated implementations inherit no abstraction boundaries, so they can integrate optimizations across the stack and reach a higher performance ceiling. We instantiate this paradigm for training frameworks as ForgeTrain, which holds a trusted framework as a golden reference and relaxes equivalence monotonically from Bit-for-Bit to Surpass. Experiments across multiple model–hardware configurations show that ForgeTrain consistently produces correct training engines and improves MFU over established training frameworks by 4.7–33.2%. To our knowledge this is the first production-grade training framework forged end-to-end by AI to match or surpass its human reference.
Introduction
ModelBest Inc. Beijing, China
Tsinghua University Beijing, China [email protected]
Training a frontier model has become one of the most capitalintensive activities in computing: a single frontier run already costs tens of millions of dollars and is headed toward $1B [3], on top of industry-wide AI-infrastructure capital on the order of $700B in 2026 [17]. The training framework fixes the throughput, the memory footprint, and the modelFLOPs utilization (MFU) of every run. The arithmetic is blunt: against a $700B capital base, raising training performance by a mere 10% is worth at least $70B a year, and every point of MFU a framework loses costs millions per run. Capturing this value means driving each training scenario to its peak performance. In practice, however, training a large model hardly avoids a general-purpose engine such as Megatron-LM [16]. Its megatron/ tree has accumulated abstraction layers, configuration switches, and special-case branches for every model, scale, and device it has ever served (Section 2.1). Yet any one concrete model-and-hardware pair exercises only a small slice of them. This generality tax has two components. (1) Suboptimality: a single framework must serve all scenarios at once, so the tensor sharding, memory layout, operator fusion, and communication schedule that would be optimal for the case at hand are constrained to stay
∗ Equal contribution.
† Corresponding authors.
1
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
viable across all the others; scenario-specific peak performance is out of reach. (2) Runtime overhead: the abstraction, indirection, and configuration dispatch accumulated to host every scenario erode MFU at every step. The heavier the inherited baggage, the greater the suboptimality and runtime overhead. The general framework was the rational response to an economic premise. Writing code was expensive and slow, so software engineering maximized reuse of what was already written, and extending a shared framework always beat rewriting one per scenario. Coding agents, however, have driven the cost of writing code sharply down [1, 18], and that changes the calculus. The better move now is to forge a dedicated framework for each scenario, fitted natively to its model and hardware, exempt from both taxes at once. Prior AI attempts. Agents have already been set to write entire systems from scratch. VibeTensor [18], the first fully AI-generated deep-learning runtime, was synthesized without a reference and runs 1.7–6.2× slower than PyTorch [13], because no source of truth forces its global behavior onto a production baseline. Claude’s C compiler [1] was developed under differential testing against GCC and grew to roughly 100K lines that compile the Linux kernel, yet it still delegates assembly and linking to GCC, and its authors state that it is not production-ready. These failures share a diagnosis. Building a complex system with agents poses three essential problems. Performance must be pushed to the human bar, correctness must hold throughout, and the development itself must stay efficient. Each attempt secures at most two of the three. VibeTensor anchors neither correctness nor performance, and the compiler holds correctness yet stops short of production performance. A production-grade system requires all three at once, and none has yet emerged. ForgeTrain. For training frameworks, we answer with ForgeTrain1 (Figure 3), an autonomous agent loop that forges a dedicated framework from an empty repository. The loop answers the three problems in turn. • How to reach peak performance? The forged framework must approach and ultimately surpass the handtuned framework it replaces; anything less defeats the purpose of forging. The obstacle, however, is planning rather than capability: agents already forge kernels that rival vendor libraries; what a bare loop cannot do is lay out the path from empty repository to peak as a sequence of steps it can walk. ForgeTrain marks that path with Milestones and adjudicates each with a Gate: a real training run comparing training quality and throughput at once. Passing a Gate locks the gain into the baseline, ratcheting measured MFU toward the target round by round.
• How to maintain correctness? A faster framework that trains wrong is worthless, and correctness here is the hardest thing to adjudicate: it is a global property, surfacing only across whole training trajectories. ForgeTrain therefore disciplines the forge with a golden reference, a trusted framework such as PyTorch [13], Megatron-LM [16], or MindSpeed [8] held as ground truth, so correctness never needs to be specified from scratch. The crux is the order of correctness requirements: it first reproduces the reference’s anchor, the artifacts of its training process, bit for bit, and relaxes the requirement only afterward. Bit-for-bit reproduction delivers a provably correct limit, not just a starting point. Forging under a bar that relaxes gradually from that limit turns debugging from conjecture over global behavior into pointwise comparison, which suppresses the risk of forging wrong. • How to forge efficiently? The case for forging rests on the falling cost of writing code, and a loop that spends its rounds on anything but the engine forfeits that advantage. The obstacle is that the optimizations reaching toward the peak are multi-step and pass through worse intermediate states, so a bare loop, rewarded by the Gate one step at a time, stops at the nearest local optimum, while every round lost to a broken build or a drifted dataset is a round not spent on the engine at all. ForgeTrain therefore hardens the forging environment, freezing dependencies, build, data, and evaluation, and supplies a knowledge prior, a corpus of optimization experience recording which directions historically pay off and which intermediate costs are worth tolerating. The environment keeps every round on the engine; the prior lets the loop commit to a multi-step optimization before the Gate can reward it. These three answers are ForgeTrain’s method in sum: the performance and correctness answers shape the protocol, the efficiency answer shapes the environment, and together they hold the loop to one trajectory: reproduce the reference exactly, then leave it behind. The underlying paradigm, Forge Engineering, is to build a dedicated implementation from scratch for each scenario and iteratively optimize it toward peak performance under correctness and usability constraints. ForgeTrain instantiates this paradigm for training frameworks. The Harness is itself reused across scenarios: each new model-and-hardware pair starts a fresh forge, drawing on the same construction knowledge and executable evaluations. As coding costs fall, this approach makes scenariospecific construction increasingly viable, just as falling design cost once pushed hardware from general-purpose CPUs toward domain-specific architectures [6]. Results. From an empty repository, ForgeTrain forges a dedicated ForgeEngine for seven model–hardware settings, the same Harness retargeted across two vendors: MegatronLM on NVIDIA H100 and MindSpeed on Ascend 910 NPUs. Every engine surpasses its reference, by 4.7–33.2% MFU at
1 Code and forged engines: https://anonymous.4open.science/r/forgetrain_
anon-FE80. 2
ForgeTrain Megatron-LM / MindSpeed
ForgeTrain, H100
46.2 (+9.1%)
CPM4 0.5B
46.1 (+4.9%)
CPM5 1B
50.8 (+4.7%)
CPM5 16B-A3B
27.6 (+17.7%)
CPM4 8B
51.2 (+5.2%)
CPM5 1B
46.4 (+15.5%)
CPM5 130M 0
however, exercises only a small slice of them. Two measurements make the tax visible. First, suboptimality: across the 128,275 lines of its megatron directory in core_v0.15.0, roughly one line in every 134 is a configuration branch on args, config, or self.config. The tensor sharding, memory layout, operator fusion, and communication schedule that would be optimal for the case at hand are constrained to stay viable for all the scenarios the framework serves. Second, runtime overhead: the abstraction built to host those scenarios is equally deep. Whereas a dedicated engine can invoke the kernel directly, a single QKV projection traverses seven Python module boundaries before reaching the GEMM, including GPTModel, the block, layer, and attention modules, a Transformer Engine wrapper layer, an autograd.Function, and the aten dispatcher (Appendix J). Communication–computation overlap illustrates the suboptimality this generality imposes. Megatron-LM’s tensorparallel overlap relies on the process-wide environment variable CUDA_DEVICE_MAX_CONNECTIONS being set to 1: with a single hardware queue, the all-gather issued ahead of the GEMM is guaranteed to be scheduled first, and the framework asserts this value whenever tensor or context parallelism is enabled. Expert-parallel overlap for MoE models needs a value larger than 1, so that dispatch communication and expert computation can actually run concurrently. When a MoE model uses both forms of parallelism, the framework can only warn that the user should “set CUDA_DEVICE_MAX_CONNECTIONS to 1 or 32, depending on which parallelization you want to prioritize.” Both overlaps are implemented and each is optimal in isolation, yet a single global knob that must remain valid for every scenario forces the combined case to forfeit one of them. A dedicated implementation for a fixed workload can instead express the required ordering through explicit stream and event dependencies and obtain both overlaps at once. Forge Engineering uses this freedom to build a dedicated implementation from scratch and iteratively optimize it toward peak performance under correctness and usability constraints. The deployment scenario specifies the workload, hardware, and operational requirements. Iteration proceeds through implementation, execution, evaluation, and revision; each accepted candidate must satisfy the applicable correctness and usability checks. A trusted reference anchors the behavior to preserve. The Harness supports this process with reusable construction knowledge and executable evaluations, allowing the same process to guide separate implementations across scenarios. ForgeTrain instantiates this paradigm for training infrastructure.
ForgeTrain, Ascend
Qwen3 0.6B
30.6 (+33.2%) 15
30 MFU (%)
45
60
Figure 1. Framework-level MFU of each ForgeEngine against its golden reference; relative gain annotated. matched training quality (Figure 1), and its forged FlashAttention and GEMM kernels run level with FlashAttention-3 and cuBLAS. Correctness is validated on the three engines carried into long-horizon training, including MiniCPM4 0.5B, MiniCPM5 1B, and MiniCPM5 130M, each holding loss parity with its reference through the run and downstream parity with the reference-trained baseline. Concretely, we make three contributions: • We propose Forge Engineering, a paradigm for building systems software from scratch for each scenario and iteratively optimizing the implementation toward peak performance under correctness and usability constraints. A reusable Harness supports this process through construction knowledge and executable evaluations. • We propose ForgeTrain, which instantiates Forge Engineering for training frameworks. A trusted framework serves as the golden reference: the loop first reproduces its training artifacts bit for bit, then relaxes the requirement stage by stage to training-quality parity, so the forged engine first equals the reference and then beats it. Milestones and the Gate drive the throughput gains; long-run validation checks correctness after each relaxation; a hardened environment and a knowledge prior keep the loop efficient. • We deliver ForgeEngine, the per-scenario engines ForgeTrain forges: they match their golden reference on training quality and surpass it by 4.7–33.2% in throughput, and we run MiniCPM4 0.5B’s decay phase with its engine to downstream-evaluation parity with a Megatron-trained baseline. To our knowledge this is the first productiongrade training framework forged end-to-end by AI to match or surpass its human reference.
2
Background
2.1
Generality Tax and Forge Engineering
2.2
Megatron-LM [16] makes the generality tax concrete. The framework’s source code has grown over time to accommodate new models, scales, and devices, and has accumulated layers of abstraction, configuration switches, and specialcase branches. Any one concrete model-and-hardware pair,
Coding Agents and Harness
A coding agent is a language model that acts through tool calls: it inspects a repository, edits files, and invokes build and test commands. These calls become a development process only under an agent harness, or scaffold, which supplies the 3
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu (a) General framework
one implementation serves 𝑁 scenarios — choices stay open at run time Suboptimality Runtime overhead training step GPTModel
one setting
TransformerLayer SelfAttention TE wrapper autograd.Function aten dispatch
128,275 lines of code a config branch every 134 lines
specialize
GEMM
× 7 module boundaries vertically integrate training step one execution path GEMM
✓ optimal for it
(b) Forge engineering
ForgeTrain: Method
3.1
Overview: The Agent Loop and the Golden Reference
ForgeTrain turns one training scenario into a framework built for that scenario alone (Figure 3a). Its inputs are a golden reference (PyTorch [13], Megatron-LM [16], or MindSpeed [8]) and a thin per-scenario configuration: a reference script that drives the golden reference and a single TOML file describing the target model and hardware. The Harness, the one component reused across every forge, guides and constrains a coding agent through the build, and what the agent yields is a dedicated ForgeEngine fitted to that scenario and to nothing else. No human stands inside the loop: once started, the agent generates, runs, and revises the framework on its own. Prior to the forging loop, the agent instantiates a scenariospecific Harness. Working from the reference script and the TOML, it runs the golden reference once and captures the artifacts of that run as the bit-for-bit anchors the engine will be held to, and it writes the concrete Milestone and Gate scripts for the target model, hardware, and parallel configuration. Once forging begins, these scripts are frozen on the Harness side, out of the develop agent’s reach (§3.2). Forging then unfolds in two distinct stages. The first demands an exact match at every anchor: the agent loop establishes a bit-for-bit implementation of the training behavior the golden reference specifies. The second relaxes that demand to training-quality parity, and the agent uses the resulting reference-anchored evidence to guide performance optimizations whose numerical paths may differ (Table 1). The objective is realized through repeated forging rounds (Figure 3b): the develop agent proposes a change, the Harness evaluates it against the active milestone, and the review agent verifies the implementation and the evaluation path. The measured outcome then determines the context of the next round. A training framework decomposes into two layers, and the loop proceeds through both. The framework layer holds the pipeline, communication, optimizer-state, and scheduling code that treats each operator as a unit; beneath it, the operator layer holds the kernels the pipeline invokes. Both layers are forged under the same discipline, and the agent uses profiling evidence to select operators for further optimization after the framework layer is established. The forged engines along these two layers are analyzed in §4.2.1–§4.2.2. The following subsections describe the performance protocol (§3.2), the correctness chain (§3.3), and the mechanisms for efficient forging (§3.4).
TransformerBlock
× optimal for none
3
3 passes → 1 kernel
one scenario — choices close at build time
Figure 2. The two forms of the generality tax (top) and the two freedoms that answer them one for one (bottom). ★ marks each scenario’s own optimal setting.
execution loop, tools, context management, and instructions that enable a model to act. Restrained by that loop, an agent revises its implementation using execution feedback, and can take on complex, often multi-file software-engineering tasks that take human professionals hours or days, as SWEBench Pro evaluates [4]; it can even forge low-level operators, with Sakana’s CUDA Engineer synthesizing thousands of test-verified CUDA kernels [9]. The general harness turns the model into a sustained developer rather than a single-turn responder: across context windows the agent must recover project state, identify unfinished work, preserve progress, and verify changes. Anthropic’s harness for long-running coding agents supports exactly this, where an initializer prepares the environment and feature list while later sessions make incremental changes, test them, and leave progress records and Git commits for subsequent sessions [19]. Harnesses are, however, not one-size-fits-all: task-specific harnesses adapt the loop and its evaluation to a domain. For training infrastructure, correctness is a global property of an entire optimization run rather than of any local unit test, and it must be judged alongside throughput. ForgeTrain therefore provides a task-specific Harness that organizes construction through Milestones and executable Gates, progressing from exact reproduction of reference anchors to performance optimization under training-quality constraints. Section 3 describes the protocol.
3.2
Peak Performance: Milestones and the Gate
For a coding agent, reaching peak performance is primarily a planning problem rather than an implementation problem. 4
ForgeTrain
Golden reference Golden reference
Reference script
Reference Scenario script TOML
1. Bit-for-Bit 1. Bit-for-Bit
Develop Develop Agent Agent
PyTorch / Megatron-LM PyTorch / Megatron-LM / MindSpeed / MindSpeed
forge
Scenario TOML
forge
1 0 Reference Reference
11
10
01
11
00
01
0
0
1 0 Candidate Candidate
11
10
01
11
00
01
0
0
bit-identical: bit-identical: max_abs_diff max_abs_diff =0 =0
Knowledge prior
Knowledge Milestones prior
(b)
Milestones Evaluation
forge
Evaluation
Pass / next round
(b) Agent loop Agent loop
Fail / rollback
HarnessHarness
Fail / rollback
instantiateinstantiate
Pass / next round
Implementation Implementation Training pipeline
Training Kernel pipeline fusion
Kernel Overlap fusion
Overlap
2. Surpass 2. Surpass Quality parity Quality + higher parity throughput + higher throughput
CUDA graphs
CUDAKernel graphsforge
Kernel ... more forge
... more
Higher MFU Higher MFU reference reference
evaluate evaluate
Gate
Gate
Rounds criteria
criteria
Rounds
Reference Reference Candidate Candidate
Quality parity Quality parity forge
ForgeEngine ForgeEngine
verdict
verdict
Loss
Review Agent Review Agent
Steps
per-scenario, per-scenario, regenerated regenerated
(a) Pipeline (a) Pipeline
Loss
(b) The loop (b) The loop
Steps
(c) Targets (c) Targets
Figure 3. ForgeTrain at a glance. (a) The Harness turns a golden reference and a thin per-scenario configuration into a ForgeEngine; (b) the forge–evaluate–verdict loop that builds it; (c) the targets the Gate enforces: bit-for-bit identity, then higher MFU with a loss curve that tracks the baseline. Table 1. The three phases of a forge. Phase
Equivalence demanded
Freedom granted
Success criterion
0: Instantiation 1: Bit-for-Bit 2: Surpass
— Bit-level identity Training-quality parity
— None beyond the reference Restructure, fuse, rewrite
Milestones, Gate scripts, and bit-for-bit anchors in place Exact match at every anchor Quality parity, higher throughput
Frontier agents can already implement many local optimizations, including specialized kernels [9, 10] and large system components [1], but local capability does not determine which dependent changes to attempt or how to organize them into a complete path from an empty repository to peak throughput. ForgeTrain addresses this gap with Milestones and Gates. Milestone decomposition. A Milestone is an intermediate target verified against the golden reference, and the decomposition is recursive: a long-range target splits into sub-goals, each sub-goal into shorter steps, until every step is short enough for the develop agent to complete and the reference to check. The decomposition is designed for the capabilities and dependencies of the target model, hardware, and parallel configuration, so different scenarios may expose different intermediate milestones. It is written at instantiation (§3.1): the MiniCPM4 8B scenario, for instance, carries a Milestone in which the agent implements tensor parallelism on its own, whereas the 0.5B scenario needs no tensor parallelism and carries a data-parallelism Milestone in its place. During reference reproduction, a Milestone can be an operator-level checkpoint; during performance optimization, it can be an intermediate performance objective. Milestones are therefore the finest granularity at which the loop advances, and the decomposition fixes both how the
work proceeds and when each piece of it may be admitted. The Forge Loop is precisely this progression: round by round, the agent advances through the sequence of Milestones. The Gate. A Gate is the Harness-controlled verdict on whether a candidate implementation satisfies a prescribed condition, measured in a real training run. Milestones set the targets; Gates decide when a target counts as reached, so every Milestone is closed by a Gate that checks its advancement condition, such as training quality within bounds and throughput at the mark for a performance Milestone. Over a multi-round performance optimization the Gate serves three functions. First, measured positioning: every verdict returns the candidate’s measured quality and throughput, so the develop agent judges its progress from data rather than from an unverified performance estimate. Second, actionable failure: a rejection states whether the problem is training quality out of bounds or speedup short of the mark, which gives the next round a concrete direction. Third, safe accumulation: optimizations that pass are written into the current baseline and become the starting point of later rounds, while failed changes are rolled back whole, so performance already won is preserved without carrying a failed exploration into the state that follows. The performance target is set above the current implementation and no single optimization reaches 5
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
it, so MFU rises through successive measurement and admission (Figure 3c). A capable agent may try to hack a Milestone’s Gate rather than earn it, so the Harness enforces a privilege split: the agent can only edit the candidate implementation, while Gate conditions, run shapes, and reference-side artifacts belong to the Harness and stay out of its reach. The review agent additionally audits for proxy execution, fabricated metrics, altered workloads, or weakened evaluation logic. Progress can therefore come only from candidate changes that pass the prescribed Gates. 3.3
identical inputs and seeds, every anchored tensor satisfies max_abs_diff = 0. One deviating bit at any anchor is failure. We demand this over the usual loss-curve agreement because it rules out agreement for the wrong reasons: canceling errors, an unstable kernel passing on a benign input, and a wrong reduction order hidden inside an acceptable loss can each fool a curve-level comparison, but none survives zero tolerance. Any deviation is reported as a localized defect to be fixed, never an optimization to be kept. Our singlenode, eight-GPU ForgeEngine build clears the full anchor set at roughly 80% of Megatron-LM’s throughput: slower, but provably computing the reference’s result. Long-run validation: holding parity after the relaxation. The Surpass stage then drops the equivalence one notch to training-quality parity, and the develop agent may restructure freely, fusing operators, reordering reductions, overlapping communication, and rewriting recomputation policies. These moves break the bit-level anchors, so the enforcement of correctness passes to the real training run behind every gate (§3.2), read through sensitive indicators: the loss trajectory must coincide with the baseline within runto-run variation, downstream evaluation must hold under an identical fine-tuning recipe, and gradient-norm statistics, which are more sensitive to numerical perturbation than the loss, expose a deviation before it reaches the loss, pinning the error to the round that introduced it. The runs must be long because a framework’s errors, such as a wrong reduction order or a subtle mixed-precision slip, are invisible in a single step and accumulate only over a trajectory. And the criterion is valid only in this order: first prove the engine computes the quantities the reference intends, then allow it to compute them by a different numerical route.
Correctness: The Equivalence Chain and Long-Run Validation
Every step of performance presupposes correctness, and at this scale a training framework’s correctness is a global property. Collective communication, optimizer-state management, mixed-precision policy, and scheduling go wrong in ways that surface only across an entire training trajectory, never in a local unit test. The requirement is moreover double: the framework must reproduce the golden reference exactly to be trustworthy, yet diverge from it to be faster. No fixed target is at once “equal to” and “better than” the reference. The equivalence chain resolves the tension by separating the two requirements in time, descending a lattice of equivalence relations from bit-level identity to trainingquality parity; relaxation is one-way and happens at stage granularity, which is one level coarser than the Milestones of §3.2. The order carries the argument: before any optimization begins, the engine agrees with the reference to the bit at every anchor captured during instantiation, so any error a later rewrite introduces localizes to a concrete deviation at a concrete tensor rather than a guess about global behavior. This progression has not, to our knowledge, been articulated as a methodology: prior reference-anchored work uses the reference only for differential testing, and reference-free synthesis has no reference to tighten against. Bit-for-Bit: zero-tolerance adjudication. The anchors are the artifacts of the golden reference’s own training run on the target scenario, captured at instantiation as machinecheckable assertions: the activations, the gradients, the optimizer states, the loss-scaling trajectory, and the collectivecommunication pattern, together with the invariants over their ordering (a given all-reduce must complete before the corresponding parameter update). They are taken at peroperator granularity, at exact numerical values, and over the edge cases production must survive but unit tests omit (gradient overflow and loss-scale backoff, checkpoint saveand-restore across a parallelism change, the first step after resumption); weaker anchors admit an implementation correct under normal inputs yet silently wrong in production. The anchor set is discovered by the agent’s own instrumentation rather than curated by hand. The forged framework must then clear those anchors with exact equivalence: given
3.4
Forging Efficiency: The Environment and the Knowledge Prior
The protocol above answers what each step should do and whether it was done right; forging efficiency decides how many rounds those steps take. A loop that spends rounds on anything but the artifact erodes the near-zero forging cost the enterprise rests on, and the waste comes from two places, friction in the environment and detours in the search. ForgeTrain answers each with one design. A hardened forging environment. The first waste has nothing to do with the artifact: a dependency that fails to install, a build that breaks intermittently, an evaluation entry point that does not reproduce. Every round the develop agent loses to such accidents is pure loss, and the accidents cascade: one mis-installed dependency contaminates every verdict after it. The Harness therefore hardens the environment up front: dependencies, build, data, and evaluation entry points are frozen at instantiation (§3.1), before the loop starts, and the fully codified evaluation of §3.2 pays a second dividend here, since a verdict that cannot be argued with can also be triggered cheaply at any moment. Every round the develop 6
ForgeTrain
agent spends therefore goes into the engine itself rather than into the surrounding infrastructure. The knowledge prior. The second waste comes from the search itself. The rewrites that approach the peak are multistep and their intermediate states are worse: a fusion pays off only after the memory layout is also changed; a reduction reorder helps only once communication is overlapped. An unaided search stops at the nearest local optimum, because the gain lies past a run of worse candidates. The Gate admits each of those stopped steps as it should, but it does not supply the next one. The limit here is not capability: a frontier coding agent can write each of these rewrites, but it will not commit to a sequence whose early steps measure worse. The knowledge prior addresses this by raising the search’s effective learning rate. It is a durable corpus of optimization experience the develop agent consults, recording which aggressive directions pay off in a given regime, which intermediate costs are worth tolerating, and which changes compound. Consulting it leads the develop agent to accept a run of worse candidates and reach the optimum beyond them. Every such attempt still passes the Gate, so the prior widens the search without weakening correctness: the prior determines what the loop attempts, the Gate what is admitted. The prior itself never writes code, and it lives in the Harness rather than in the forged framework, leaving the zero-human-written-code discipline intact. The three groups of mechanisms are independent of one another: each governs a different property of the forge, and none weakens the guarantees the other two provide. Together they are what makes the match-then-surpass progression against the golden reference executable by an imperfect agent.
4
Experiments
4.1
Setup
Model architectures are given in Appendix B for the H100 settings and Appendix H for the Ascend settings; the parallel degrees and batch sizes for each setting are in Table 2. Performance experiment. We forge training frameworks for seven model–hardware settings. On H100, we use the official Claude Opus 5 agent (1M-token context window) with ForgeTrain to forge Qwen3 0.6B, MiniCPM4 0.5B, MiniCPM5 1B, MiniCPM5 16B A3B, and MiniCPM4 8B. We compare their MFU with Megatron-LM 0.15 [16]. On Ascend, we forge MiniCPM5 1B and MiniCPM5 130M and compare their MFU with their golden references, MindSpeed on Megatron core 0.12.1 and PyTorch respectively. MiniCPM4 0.5B and MiniCPM5 1B on H100 and MiniCPM5 1B on Ascend each carry a one-layer MTP (Eagle) head; the other settings train without MTP. Table 2 collects the framework configurations and the corresponding MFU trajectories. Correctness experiment. On H100, we use Claude Opus 4.6 to forge a MiniCPM4 0.5B engine, with MegatronLM as the golden reference. We run the forged engine through the decay phase of pretraining, then fine-tune the resulting checkpoint with Megatron-LM and evaluate it on downstream benchmarks. On Ascend, we use Opus 4.6 with a 1M-token context window to forge the MiniCPM5 1B engine (with a one-layer MTP head) with MindSpeed as the golden reference, and GLM-5.2 with a 1M-token context window to forge the MiniCPM5 130M engine with PyTorch as the golden reference. We run decay training with the 1B engine and a complete stable–decay–SFT training pipeline with the 130M engine, and evaluate the resulting checkpoints of both on downstream benchmarks. Ablation experiment. We use Claude Opus 5 with a 1M-token context window for two ablations on H100. We remove the Harness entirely for MiniCPM5 16B A3B, and remove the quality constraints while retaining the Milestones for MiniCPM4 0.5B. We use the corresponding completeHarness runs in Section 4.2 as comparisons. Metric. We report throughput as model-FLOPs utilization (MFU), computed from an exact per-operator FLOP count rather than the usual closed-form approximation (the full expression is given in Appendix C). The same expression and the same peak-times-devices denominator are used for the forged engine and for the baseline.
We evaluate ForgeTrain through three experiments. The first measures performance by comparing the MFU reached during forging with established training frameworks. The second checks correctness by carrying the forged engines through end-to-end forging and production training. The third evaluates the usefulness of the Harness by conducting a forging task after removing selected Harness components. Sections 4.2, 4.3, and 4.4 describe these experiments in this order. Hardware and software. We conduct the experiments on two platforms: NVIDIA H100 GPUs and Huawei Ascend 910 NPUs. On H100, we use PyTorch 2.8.0 with CUDA 12.9, cuDNN 9.10.02 [2], and NCCL 2.27.3. The forged H100 engines use Transformer Engine 2.4 [12], and the hand-written operators use CUTLASS 4.5.0 [11]. On Ascend, we use PyTorch 2.7.1+cpu with CANN 8.5.0 and torch_npu 2.7.1. All experiments use a sequence length of 4096 and the mixedprecision settings specified for the corresponding platform.
4.2
Main MFU Results
We first report framework-level performance across seven model–hardware settings, then present the kernels selected and forged by the agent during the MiniCPM4 0.5B run. We conclude with an analysis of the framework and kernel implementations. 4.2.1 Framework-Level Forging. ForgeTrain surpasses the corresponding Megatron-LM or MindSpeed baseline across all seven model–hardware settings (Tables 2), with 7
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
relative MFU gains of 4.7%–33.2%. The dense H100 models gain approximately 5%–9%, while MiniCPM5 16B A3B and the Ascend models show larger improvements of about 16%–33%. Table 2 shows how each setting reaches that result. Forging cost has two readings: the round at which the engine first passes its baseline, roughly 12–20 on H100 and 3–5 on Ascend, and the round at which it peaks, the last checkpoint 𝑘 reported for each setting in Table 2. These results demonstrate that scenario-specific forging can improve performance across both hardware platforms.
4.2.3 Sources of the Speedup. The forged 0.5B engine’s advantage over Megatron-LM 0.15 comes from three sources. The first is a reproduced standard set: optimizations that exist in Megatron with the same boundaries, reproduced component for component, including bucketed reduce-scatter overlapped with backward, ZeRO-1 optimizer sharding, fused RMSNorm, RoPE, SwiGLU, and cross-entropy kernels, and a direct Transformer-Engine attention path. The harness rediscovered them in the first nine steps of the framework-level optimization trajectory (Table 8, Appendix A). The second is vertical integration: components that Megatron also has but keeps as separate stages, which the forged engine merges into a single execution path for the fixed scenario. Megatron’s optimizer tail runs gradient clipping, the Adam update, and the FP32-to-BF16 parameter copy as three passes over every parameter; the forged engine fuses them into one Triton kernel with a single pass. Likewise, rather than graphing individual transformer layers, the forged engine captures a whole microbatch’s forward, loss, and backward as one CUDA Graph replay. The third is deep customization: scenario-specific points with no counterpart in Megatron, where the forged engine rewrites the execution path around the fixed model, shapes, and parallel layout. Optimizer scalars and the clipping coefficient stay on the device, so the step never synchronizes with the host; the loss all-reduce is moved off the critical path; and the read-only and layout-specific paths are specialized to the fixed shapes. Each mechanism’s source locations, configuration gates, and the per-step MFU delta it contributes are detailed in Appendix F.
4.2.2 Operator-Level Forging. This section analyzes the operator forge carried out as part of the H100 MiniCPM4 0.5B correctness experiment. The framework and operator results reported here come from the same no-MTP MiniCPM4 0.5B forge. The agent first profiles the training step to identify the operators that account for the largest runtime cost. The profile (Table 3) leads the agent to select five GEMM variants and the attention core for further forging. The GEMMs cover the QKV projection, output projection, and the two FFN projections. For attention, the agent forges both the forward and backward kernels. The optimization families these forgings instantiate are cataloged in Appendix E. For each selected operator, the agent writes a specialized implementation for the shapes issued by the MiniCPM4 0.5B training step. Each implementation is checked against the corresponding reference computation before it is included in the forged engine. The resulting kernels are then compared with the corresponding vendor implementations at the target shapes (Table 4). GEMM results. The agent-forged CuTeDSL kernels remain close to cuBLAS on most target shapes and surpass it on several forward and backward passes. The strongest results occur for the attention-output and FFN projections. For example, the forged attention-output backward kernel reaches 73.0% MFU, compared with 71.5% for cuBLAS. Attention results. The forged FlashAttention forward trails cuDNN, FlashAttention-3, and FlashAttention-4 but leads Transformer Engine, while the forged backward leads every baseline, including FlashAttention-3 and FlashAttention-4. This backward result is especially significant because attention backward carries twice the FLOPs of forward; Appendix D traces the full forging trajectory of this backward kernel, from the naive atomic-accumulate baseline to the final warped-specialized, layout-refactored implementation. The attention contribution therefore comes mainly from the backward kernel rather than from a uniform improvement in both directions. End-to-end result. After the selected GEMM and FlashAttention kernels are integrated into the same correctness-forged framework, the complete MiniCPM4 0.5B engine reaches 44.13% MFU, compared with 41.66% before operator-level forging. The operator stage therefore adds 2.47 percentage points of end-to-end MFU.
4.3
End-to-End Correctness Validation
We evaluate ForgeTrain’s correctness end to end by using the forged engines for real production training and comparing their training trajectories, resulting checkpoints, and downstream model quality with those of the trusted reference implementations described in Section 4.1. The engines under test are the framework-forged engines running at full production throughput: the MiniCPM4 0.5B engine trains at 41.66% MFU on 16×H100 against 40.3% for Megatron-LM (Figure 4 traces how forging raised it from 30.45% to this level), and the MiniCPM5 1B and MiniCPM5 130M engines train on Ascend 910 at 37.0% MFU on 32 devices (DP= 32) against 33.6% for MindSpeed and at 24.9% MFU on 2 devices (DP= 2) against 23.0% for MindSpeed. In wall-clock terms, the MiniCPM4 0.5B engine first passed Megatron-LM after roughly ten hours of forging and reached this level after two to three further days. 4.3.1 Training-Trajectory Agreement. We compare loss trajectories under matched checkpoints, data, and schedules. Figure 5 summarizes the three runs. For MiniCPM4 0.5B, we 8
ForgeTrain
Table 2. Framework configurations and MFU trajectories during forging. Par. gives the data (DP), tensor (TP), and expert (EP) parallel degrees; GBS and MBS are global and micro-batch size; per-model details are in Appendix B. 𝑘 indexes a setting’s accepted forging rounds (Appendix A); a slash marks a run that ended before that checkpoint. MiniCPM4 0.5B and MiniCPM5 1B on H100 and MiniCPM5 1B on Ascend carry a one-layer MTP (Eagle) head; the remaining settings train without MTP. Setting
Par.
GBS MBS Baseline (%)
H100 (Megatron-LM)
MFU (%) at checkpoint 𝑘
Final (%)
Gain
46.2 46.1 50.8 27.6 51.2
+9.1% +4.9% +4.7% +17.7% +5.2%
46.4 30.6
+15.5% +33.2%
𝑘=4 𝑘=8 𝑘=12 𝑘=16 𝑘=20 𝑘=24 𝑘=28 𝑘=32 𝑘=36
Qwen3-0.6B DP2 MiniCPM4-0.5B DP2 MiniCPM5-1B DP2 MiniCPM5-16B-A3B EP8 MiniCPM4-8B DP4 TP2
80 80 80 32 32
10 10 4 4 2
42.4 43.9 48.5 23.5 48.7
36.2 36.3 25.0 36.9 43.6 44.9 18.4 22.2 23.1 44.5
Ascend 910 (MindSpeed) MiniCPM5-1B MiniCPM5-130M
DP2 DP2
80 80
2 5
40.2 23.0
44.3 44.2 / 23.0 46.1
45.4 45.2 / 23.7 47.2
46.3 46.1 / 24.7 50.9
45.6 / / 24.1 /
𝑘=1 𝑘=3 𝑘=5
𝑘=7
𝑘=9 𝑘=11 𝑘=13 𝑘=15 𝑘=17
30.1 44.8 46.3 10.0 22.6 27.1
46.4 29.4
/ 29.9
/ 30.4
46.8 / / 25.9 / / 30.6
/ / / 26.7 / / /
/ / / 27.6 / / /
Table 3. Operator time budget of a steady-state training step (bfloat16; forward+backward+optimizer, averaged over 5 profiled steps). Shares are of GPU compute time, communication excluded; the GEMM row covers forward, dgrad, and wgrad. Operator class
Share of step
GEMM (qkv/o/fc1/fc2/lm-head) Attention core (cuDNN FlashAttention fwd+bwd) Elementwise (RMSNorm, RoPE, activation, Adam) Other (Triton concat, sort)
55.9% 16.6% 24.4% 3.1%
Figure 4. Framework-level forging trajectory of the MiniCPM4 0.5B engine on 16×H100. Each point is an accepted optimization step; MFU rises from 30.45% to 41.66% and crosses the Megatron-LM baseline of 40.3% (dashed) after the kernel-fusion steps give way to overlap, bucketing, and CUDA-graph optimizations. This is the engine used for the MiniCPM4 0.5B correctness runs below.
Table 4. Forged-kernel throughput (MFU, %) at the 0.5B engine’s own shapes (H100, bfloat16; weight gradients accumulated in FP32). FA-3 and FA-4 are FlashAttention-3 [15] and FlashAttention-4 [20]; TE is Transformer Engine. “n/a” marks passes the fused qkv/output backward kernel does not expose separately. Bold marks the best value in each comparison (ties within 0.1 pp both bold).
we resume from the same 480k-step checkpoint and run the same 48,000-step decay schedule on Ascend; the final losses are 1.3227 and 1.3226. For MiniCPM5 130M, the forged engine runs the complete stable–decay–SFT pipeline on Ascend, and its loss trajectory is compared with a PyTorch reference over the same schedule. The final SFT losses are 0.845 and 0.861.
(a) Attention Direction
Forged
cuDNN
FA-3
FA-4
TE
Forward Backward
42.6 45.7
47.8 35.0
44.3 43.4
43.7 43.3
39.0 36.2
(b) GEMM (forged / cuBLAS) Operator
Forward
Backward
dgrad
wgrad
qkv attn-out fc1 fc2 output
69.7/69.9 70.8/69.8 77.1/74.2 80.2/80.5 67.5/67.9
68.4/73.7 73.0/71.5 81.4/81.5 79.9/78.9 73.6/70.4
72.9/72.8 69.9/69.9 81.9/82.1 76.9/76.9 n/a / 75.9
62.4/71.6 72.7/69.3 83.3/81.1 83.3/81.7 n/a / 66.1
4.3.2 Downstream Model Quality. For MiniCPM4 0.5B, we fine-tune the decay-phase checkpoints from both runs with the same Megatron-LM SFT procedure and evaluate the resulting models on six benchmarks. This keeps SFT fixed and isolates the effect of the pretraining engine. The model trained with the forged engine matches or exceeds the Megatron baseline on five of the six benchmarks. For MiniCPM5 1B on Ascend, we evaluate the checkpoints produced by the two decay runs in Section 4.3.1 on eight
resume the decay phase from a shared stable-phase checkpoint and compare with Megatron-LM. For MiniCPM5 1B, 9
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
Table 5. Downstream evaluation of MiniCPM4 0.5B on H100 after the same decay schedule with Megatron-LM and with the forged engine, followed by identical Megatron-LM SFT. Benchmark
Megatron
ForgeTrain
Δ
CMMLU MBPP (sanitized) GSM8K MATH (PRM800K-500) IFEval C-Eval (CoT)
65.23 57.98 49.96 31.00 50.65 64.16
65.69 57.98 50.95 32.20 50.28 65.94
Average
53.16
53.84
+0.46 0.00 +0.99 +1.20 −0.37 +1.78
+0.68
Table 6. Downstream evaluation of MiniCPM5 1B on Ascend 910 after the same decay schedule with MindSpeed and with the forged engine. Benchmark
MindSpeed
ForgeTrain
Δ
C-Eval MMLU-Redux HumanEval MBPP (sanitized) GSM8K MATH-500 BBH IFEval
37.10 42.90 36.59 44.75 58.91 30.20 26.72 34.01
35.53 42.41 37.20 48.25 54.59 31.40 22.01 34.38
Average
38.90
38.22
−1.57 −0.49 +0.61 +3.50 −4.32 +1.20 −4.71 +0.37
−0.68
Table 7. Downstream evaluation of MiniCPM5 130M on Ascend 910 after the complete stable–decay–SFT pipeline with the PyTorch (torch_npu) reference and with the forged engine. Figure 5. Loss trajectories for the three long-horizon comparisons: MiniCPM4 0.5B decay on H100, MiniCPM5 1B decay on Ascend 910, and the complete MiniCPM5 130M stable–decay–SFT run on Ascend 910. benchmarks, covering Chinese and English knowledge (CEval, MMLU-Redux), code (HumanEval, MBPP), mathematical reasoning (GSM8K, MATH-500), general reasoning (BBH), and instruction following (IFEval). Table 6 reports the results. The model trained with the forged engine exceeds the MindSpeed baseline on four benchmarks and trails it on the other four. For MiniCPM5 130M on Ascend, both the forged engine and the PyTorch reference run the complete stable–decay– SFT pipeline in Section 4.3.1, so this comparison also covers the SFT stage. We evaluate the final checkpoints on five English benchmarks spanning general reasoning (BBH), mathematical reasoning (GSM8K, MATH), and code (HumanEval,
Benchmark
PyTorch
ForgeTrain
Δ
BBH GSM8K MATH HumanEval MBPP
32.93 17.29 9.80 20.12 34.63
32.08 18.42 10.20 21.34 37.35
Average
22.95
23.88
−0.85 +1.13 +0.40 +1.22 +2.72
+0.93
MBPP). Table 7 reports the results. The model trained with the forged engine exceeds the reference on four of the five benchmarks. Across the three comparisons, the average difference stays within one point and changes sign, so the forged engines are interchangeable with their references at the level of downstream quality on both hardware ecosystems. 10
ForgeTrain
Related Work
5.1
Large-Scale Training Frameworks
ForgeTrain forges against a mature body of humanengineered training systems. Megatron-LM [16] popularized tensor and pipeline model parallelism for transformers; DeepSpeed’s ZeRO partitions optimizer states, gradients, and parameters across data-parallel ranks to fit trillionparameter models [14]; GPipe introduced micro-batch pipeline parallelism with synchronous updates [7]; and Alpa automatically searches a hierarchical space of inter- and intra-operator parallel plans [21]. These are the canonical incumbents of the one-codebase paradigm: a single general framework whose abstraction layers and configuration knobs are built to cover every model, scale, and hardware topology. Even Alpa’s automation searches over a fixed framework’s plan space rather than regenerating the framework itself. ForgeTrain inverts this: instead of one framework amortized across all scenarios, it forges a per-scenario ForgeEngine from scratch.
Figure 6. MiniCPM4 0.5B forging without quality constraints. The dashed line is the complete-Harness engine of Section 4.2.1, at 46.1% on the same 2×H100 setting.
4.4
5
Ablation
We examine two reduced settings with Claude Opus 5 (1Mtoken context window) on H100. The first removes the Harness for the MiniCPM5 16B A3B task in Section 4.2.1. The second retains the Milestones but removes the quality constraints for MiniCPM4 0.5B. Each is compared with the complete Harness on its corresponding model; the two ablations address different failure modes. Without the Harness. During the MiniCPM5 16B A3B implementation, the agent reduces optimizer precision from FP32 to BF16 to accelerate training. This changes the numerical computation of the optimizer, so throughput alone cannot establish that the resulting engine preserves training quality. The run therefore provides no verified performance gain at training-quality parity. Without quality constraints. On MiniCPM4 0.5B, MFU remains near 37% within 33 forging rounds, far below the 46.1% that the complete Harness reaches for the same setting (Table 2). Figure 6 shows the available trajectory through round 32. The agent’s reasoning repeatedly treats Bit-forBit agreement as a restriction on optimization. In round 18, it rules out changing deterministic attention backward because of its effect on numerical agreement. In round 32, it again treats attention and GEMM implementations as fixed and considers changes to the reduction order too risky. The search consequently concentrates on smaller changes that preserve the existing arithmetic. Together, these observations illustrate why the Harness needs an explicit quality criterion. Without one, the agent may either change optimizer precision without establishing training-quality parity, or continue to require exact numerical agreement even during performance optimization. ForgeTrain specifies the transition from Bit-for-Bit agreement to training-quality parity, giving the agent a criterion for evaluating changes to the reference computation.
5.2
AI-Generated System Software
The closest neighbors to ForgeTrain are whole-system software artifacts written end-to-end by AI. VibeTensor [18] is presented as the first fully AI-generated deep learning system, a PyTorch-style eager runtime assembled by agents. It is, however, released for agentic-systems research only: it runs 1.7–6.2× slower than PyTorch, and its authors describe a “Frankenstein composition effect” in which locally correct components compose into a globally suboptimal whole. The shortfall is not implementation skill but the absence of a truth source that forces global behavior to match a production baseline. Claude’s C Compiler [1] is methodologically the closest relative: roughly one hundred thousand lines of Rust written by sixteen cooperating agents, compiling the Linux 6.9 kernel, with GCC used as a differential-testing oracle. Promoting a reference implementation to a golden reference is precisely the discipline ForgeTrain formalizes in its Bit-forBit stage. Yet the compiler falls back to GCC for assembly and linking, is explicitly disclaimed as not production-ready, and targets the compiler domain rather than training infrastructure. ForgeTrain differs on all three axes these systems leave open: it is whole-system, production-grade, and matches and surpasses its reference on the reference’s own benchmarks. 5.3
AI-Generated Kernels and Operators
A productive line of work forges performance-critical code at node granularity. AlphaEvolve [5] couples an evolutionary search loop with automated evaluation and has discovered scheduling heuristics and matrix-multiplication kernels deployed in production; KernelEvolve [10] evolves productiongrade kernels under throughput targets; and Sakana’s CUDA Engineer [9] synthesizes thousands of test-verified CUDA 11
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
kernels. The combination of iterative or evolutionary search, test-driven verification, and profiling feedback used in these systems is precisely the Harness-driven loop, and these systems show it is already production-ready at the level of a single kernel or operator. The distinction from ForgeTrain is granularity. Each of these efforts forges one self-contained node with a clear local objective, where correctness and speed can be measured in isolation. A training framework is an integrated system whose behavior is correct only globally, across an entire optimization trajectory. Forging at that granularity is what the match-then-surpass discipline of §3 makes possible, and it is the gap these node-level efforts leave open.
6
Can AI Agents Solve Long-Horizon Software Engineering Tasks? arXiv preprint arXiv:2509.16941 (2025). https://arxiv.org/abs/2509. 16941 [5] Google DeepMind. 2025. AlphaEvolve: A coding agent for scientific and algorithmic discovery. Google DeepMind Blog. [6] John L. Hennessy and David A. Patterson. 2019. A New Golden Age for Computer Architecture. Commun. ACM 62, 2 (2019), 76–85. [7] Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism. In Advances in Neural Information Processing Systems (NeurIPS). [8] Huawei Ascend. [n. d.]. MindSpeed: An Acceleration Library for Large-Model Training on Ascend NPUs. https://gitee.com/ascend/ MindSpeed. Accessed 2026-09-09. [9] Robert Tjarko Lange, Aaditya Prasad, Qi Sun, Maxence Faldor, Yujin Tang, and David Ha. 2025. The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization & Composition. Sakana AI Technical Report. https://sakana.ai/ai-cuda-engineer/ [10] Gang Liao, Hongsen Qin, Ying Wang, Alicia Golden, Michael Kuchnik, Yavuz Yetim, Jia Jiunn Ang, Chunli Fu, Yihan He, Samuel Hsia, Zewei Jiang, Dianshi Li, Uladzimir Pashkevich, Varna Puvvada, Feng Shi, Matt Steiner, Ruichao Xiao, Liyuan Li, Nathan Yan, Xiayu Yu, Zhou Fang, Roman Levenstein, Kunming Ho, Haishan Zhu, Alec Hammond, Richard Li, Ajit Mathews, Kaustubh Gondkar, Abdul Zainul-Abedin, Ketan Singh, Hongtao Yu, Wenyuan Chi, Barney Huang, Sean Zhang, Noah Weller, Zach Marine, Wyatt Cook, Carole-Jean Wu, and Gaoxiang Liu. 2026. KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta. In Proceedings of the 53rd ACM/IEEE International Symposium on Computer Architecture (ISCA). https://arxiv.org/abs/2512.23236 [11] NVIDIA. 2025. CUTLASS: CUDA Templates for Linear Algebra Subroutines. GitHub repository. [12] NVIDIA. 2025. Transformer Engine. GitHub repository. [13] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems (NeurIPS). [14] Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC). [15] Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. 2024. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision. arXiv preprint arXiv:2407.08608 (2024). [16] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. In arXiv preprint arXiv:1909.08053. [17] TechInsights. 2026. Analysis: Hyperscaler Earnings Takeaways Q1 2026 – The $700 Billion AI Arms Race. TechInsights Industry Analysis. https://www.techinsights.com/blog/analysis-hyperscalerearnings-takeaways-q1-2026-700-billion-ai-arms-race [18] Bing Xu, Terry Chen, Fengzhe Zhou, Tianqi Chen, Yangqing Jia, Vinod Grover, Haicheng Wu, Wei Liu, Craig Wittenbrink, Wen mei Hwu, Roger Bringmann, Ming-Yu Liu, Luis Ceze, Michael Lightstone, and Humphrey Shi. 2026. VibeTensor: System Software for Deep Learning, Fully Generated by AI Agents. arXiv:2601.16238 [cs.SE] https://arxiv. org/abs/2601.16238
Discussion and Conclusion
Across seven model–hardware settings on two hardware ecosystems, every forged ForgeEngine surpasses its golden reference by 4.7–33.2% MFU at matched training quality. MiniCPM4 0.5B weights trained on a ForgeEngine match the Megatron-LM-trained baseline on downstream evaluation, and on Ascend a forged engine carried a 130M model through a complete stable–decay–SFT production run. ForgeTrain thus shows that a production-grade training framework can be forged end-to-end by an autonomous agent from an empty repository. To our knowledge, it is the first such framework forged by AI to match and surpass its human reference. ForgeTrain is thoroughly validated at smaller model scales, but forging for larger models remains to be demonstrated experimentally, and the long-term maintenance and evolution of a forged engine are untested. ForgeTrain instantiates Forge Engineering: building dedicated systems software from scratch for each scenario and iteratively optimizing it toward peak performance under correctness and usability constraints. A reusable Harness carries construction knowledge and executable evaluations across scenarios, while each implementation is tailored to its workload, hardware, and operational requirements. As coding agents improve, Forge Engineering could make dedicated systems a practical alternative to general-purpose implementations across a wider range of domains.
References [1] Anthropic. 2026. Claude’s C Compiler. Anthropic Engineering Blog. https://www.anthropic.com/engineering/building-c-compiler [2] Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cuDNN: Efficient Primitives for Deep Learning. arXiv preprint arXiv:1410.0759 (2014). [3] Ben Cottier, Robi Rahman, Loredana Fattorini, Nestor Maslej, Tamay Besiroglu, and David Owen. 2024. The Rising Costs of Training Frontier AI Models. arXiv preprint arXiv:2405.21015 (2024). [4] Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Nitin Pasari, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy, Sean Hendryx, Zifan Wang, Vijay Bharadwaj, Jeff Holm, Raja Aluri, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler. 2025. SWE-Bench Pro: 12
ForgeTrain
(ℎ𝑞 =𝑛ℎ𝑑, ℎ𝑘𝑣 =𝑛𝑘𝑣 𝑑), 𝑉 the vocabulary size, 𝐿 the layer count, and 𝑠 the sequence length. The forward per-token FLOPs of each operator class are
[19] Justin Young. 2025. Effective Harnesses for Long-Running Agents. Anthropic Engineering Blog. https://www.anthropic.com/engineering/ effective-harnesses-for-long-running-agents [20] Ted Zadouri, Markus Hoehnerbach, Jay Shah, Timmy Liu, Vijay Thakkar, and Tri Dao. 2026. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling. arXiv preprint arXiv:2603.05451 (2026). [21] Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning. In USENIX Symposium on Operating Systems Design and Implementation (OSDI).
A
Attn. projections (Q,K,V,O) : 2ℎ (ℎ𝑞 + 2ℎ𝑘𝑣 ) + 2ℎ𝑞 ℎ,
FFN (gate, up, down; SwiGLU) : 2 (2ℎ𝑓 ) + 2𝑓 ℎ,
Attention core (𝑄𝐾 ⊤, 𝐴𝑉 ; causal) : 2𝑛ℎ 𝑠𝑑, Output / cross-entropy head : 2ℎ𝑉 .
The attention core carries a factor 12 for the causal mask (folded into the expression above), and grows with 𝑠 where the others do not. The backward pass costs twice the forward for every GEMM (one matrix for the data gradient, one for the weight gradient) and twice the forward for the attention core, so each per-layer term is multiplied by three (one forward, two backward) and summed over 𝐿 layers; the output head is counted once, also at the 1:2 forward-to-backward ratio. Table 11 evaluates this for the two evaluation scenarios. Two properties of this count matter for the operator budget of §4.2.2. First, the GEMMs and the attention core together are the entire compute FLOP budget—normalization, RoPE, activations, and the optimizer contribute negligible FLOPs and are memory-bound, which is why they are absent from this table yet still consume measurable time (the elementwise bucket of Table 3). Second, the attention core is the one term that scales with 𝑠: at the 4096 sequence length it is 18.8% of the 0.5B FLOP budget, and two-thirds of that is backward, which is why the attention backward is the single operator whose forging trajectory we trace in detail (Appendix D).
Framework-Level Optimization Trajectory
Table 8 lists, in order of acceptance, the framework-level optimizations the agent loop applied to the MiniCPM4 0.5B ForgeEngine during Stage C, together with the measured MFU after each step (16×H100, DP only). This is the per-step detail behind the trajectory of Table 8: the first eight steps form the kernel-fusion regime, and the remainder the overlap/bucket/CUDA-graph regime. Two candidates (†) were measured but rejected by the Stage-C gate and reverted, so they do not carry into the final engine.
B
Model Architectures
The H100 evaluation scenarios use three released model families without architectural modification: MiniCPM4 0.5B and 8B 2 , Qwen3 0.6B 3 , and MiniCPM5 1B and 16B-A3B. All are decoder-only transformers with grouped-query attention, SwiGLU feed-forward blocks, and pre-normalization RMSNorm; MiniCPM4 uses LongRoPE positional encoding and the MiniCPM 𝜇P-style scaling (scale_emb, dim_model_base, scale_depth), Qwen3 uses standard parameterization with RoPE (𝜃 =106 ) and per-head QK-Norm, and MiniCPM5 is a dense or sparse (MoE) model with RoPE and 𝜇P scaling. Tables 9 and 10 give the configurations that fix the FLOP count and the shapes the forged operators target; the experiments train at a sequence length of 4096 (§4.1), within the context each architecture supports.
C
D
The attention backward (§4.2.2) is the clearest illustration of how the loop forges a kernel, and its trajectory (Figure 8) separates two kinds of decision. We optimize it on an H100 SXM at the engine’s attention shape (𝐵=10, 𝐻 =16, 𝑁 =4096, 𝐷=64, FP16, GQA 8:1, causal), organized—following the FA-3/FA-4 design—as a three-kernel pipeline (preprocess → main → postprocess). From a naive baseline that accumulates d𝑄 through atomicAdd (5.19 ms, 16.5% MFU), the loop first lays a structural foundation: a 128-row 𝑀-tile (versus 64), warp specialization into one producer and two consumer warp-groups (384 threads, a 24/240-register split), double-buffering on every pipeline stage, and an asynchronous five-GEMM issue schedule. These choices are mutually dependent and do not pay off in isolation—while they are being put in place latency regresses to 6.3–6.9 ms, because the d𝑄-accumulator data-flow is still the bottleneck. The decisive step is a data-layout refactor that replaces the WGMMAfragment-aware d𝑄 accumulator with a flat global buffer and a flat thread-value register-to-shared copy, and sources the d𝐾/d𝑉 GEMM operands directly from registers; this cuts shared-memory bank conflicts from 73 M to 8.2 M and brings
Model-FLOPs Utilization
We report MFU against an exact per-operator FLOP count rather than the usual closed-form 6𝑁 𝐷 approximation, which omits attention and miscounts the GQA projections and the vocabulary head. We count a multiply–accumulate as two FLOPs and report per token; a quantity is the same for every token in the batch, so the per-step numerator is this total times the number of tokens processed. Let ℎ be the hidden size, 𝑓 the FFN intermediate size, 𝑛ℎ and 𝑛𝑘𝑣 the query and key/value head counts, 𝑑 the head dimension 2 https://www.modelscope.cn/models/OpenBMB/MiniCPM4-0.5B
Anatomy of a Forged Kernel
and
https://www.modelscope.cn/models/OpenBMB/MiniCPM4-8B. 3 https://huggingface.co/Qwen/Qwen3-0.6B. 13
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
Table 8. Framework-level optimizations applied to MiniCPM4 0.5B during Stage C, in acceptance order. MFU is measured on 16×H100 (data-parallel). † marks candidates that were reverted after the gate rejected them.
(a) Qwen3 0.6B
#
Phase
Optimization
1 2 3 4 5 6 7 8
Kernel fusion Kernel fusion Kernel fusion Kernel fusion Kernel fusion Kernel fusion Kernel fusion Kernel fusion
Stage-B baseline (no fusion) Fused cross-entropy + SwiGLU Cross-entropy NaN + embedding-backward fix Direct Transformer-Engine attention Fused residual-add + RMSNorm Fused gradient-clip + Adam + param-sync as_strided cross-entropy + FFN fusion Cross-entropy prefill + fused RMSNorm backward
30.45 31.64 34.73 34.76 35.24 35.24 37.23 37.81
9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30
Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph Overlap/graph
Gradient bucketing (100 MB) Per-bucket all-gather + optimizer overlap All-gather–optimizer overlap Wgrad-overlap by default NCCL tuning + async loss all-reduce Production-default verification Batch-flush write-grad restructure Batch-flush gate ZeRO-1 sharded optimizer Sharded optimizer by default Routing / baseline fixes Fused norm-backward + multi-tensor Adam Fused-operator micro-optimizations RoPE strided-K (drop K copy) Document-aware fused RoPE Wgrad bucket 200 MB RoPE backward in-place† Forward-only CUDA Graph Full-step CUDA Graph (phase A) Full-step CUDA Graph (phase B/C) Wgrad bucket 400 MB (single bucket) Cross-step pipelining†
38.55 38.55 38.55 38.55 38.16 38.57 38.57 38.57 39.70 39.40 39.40 39.52 39.15 39.30 41.00 41.22 41.34 40.99 41.37 41.37 41.66 41.66
(b) MiniCPM4 0.5B
MFU (%)
(c) MiniCPM5 1B
(d) MiniCPM5 16B A3B
(e) MiniCPM4 8B
Figure 7. Per-setting MFU trajectories of the five H100 settings from Table 2, each normalized to its own Megatron-LM baseline (dashed). The x-axis is the setting’s 𝑘-th recorded forging round. the kernel to 2.22 ms (38.6% MFU). Three independent tactical optimizations then close the gap: 128-bit vectorized preprocess/postprocess copies (2.03 ms), skipping the causal mask on all but the diagonal block—which halves arithmeticpipe instructions (−54%)—(1.97 ms), and removing a stray cudaStreamSynchronize that had disabled programmatic dependent launch, restoring cross-kernel overlap (1.90 ms, 45.7% MFU). The result is 5.6% faster than FlashAttention-4 (2.007 ms, 43.3% MFU) and 31% faster than cuDNN (2.483 ms,
35.0% MFU); note that the margin over FA-4 runs the other way in the per-kernel NCU timings, which are slower than FA-4’s (2072 versus 1990 µs). The entire end-to-end margin therefore comes from cross-kernel launch overlap: programmatic dependent launch hides 0.17 ms across our four kernels, against no measurable overlap for FA-4’s five. The methodological point generalizes beyond this one kernel: structural decisions must be fixed jointly and up front, even through a 14
ForgeTrain
Table 9. MiniCPM4 architecture configuration for the two evaluation scenarios.𝑎 Hyperparameter
MiniCPM4 0.5B
MiniCPM4 8B
Transformer layers Hidden size FFN intermediate size Attention heads Key/value heads (GQA) Head dimension Vocabulary size Activation Normalization Position encoding Max context length Tied input/output embeddings 𝜇P (scale_emb, dim_model_base, scale_depth) MTP / Eagle head Training precision
24 1024 4096 16 2 64 73448 SwiGLU (SiLU) RMSNorm (𝜖=10−5 ) LongRoPE 32768 Yes 12, 256, 1.4 On (1 layer) bfloat16
32 4096 16384 32 2 128 73448 SwiGLU (SiLU) RMSNorm (𝜖=10−6 ) LongRoPE 32768 No 12, 256, 1.4 Off bfloat16
𝑎 The 0.5B column carries a one-layer MTP (Eagle) head. The framework-level setting of Section 4.2.1 forges it, whereas the operator-level forge of
Section 4.2.2 and the correctness experiment of Section 4.3 run the same model with this head disabled (no-MTP). The 8B column trains without MTP.
Table 10. H100 architecture configurations for the Qwen3 and MiniCPM5 scenarios. MiniCPM5 16B-A3B is a sparse MoE model (160 routed experts, top-16, plus a 512-wide shared expert); its dense FFN width applies to layer 0 only, and layers 1–27 use moe_ffn_hidden_size instead. Hyperparameter Transformer layers Hidden size FFN intermediate size MoE FFN size (routed experts) Routed experts (top-𝑘) Attention heads Key/value heads (GQA) Head dimension Vocabulary size Activation Normalization Position encoding Tied input/output embeddings 𝜇P (scale_emb, dim_model_base, scale_depth) MTP / Eagle head Training precision
Qwen3 0.6B
MiniCPM5 1B
MiniCPM5 16B-A3B
28 1024 3072 — — 16 8 128 151936 SwiGLU (SiLU) RMSNorm (𝜖=10−6 ) RoPE (𝜃 =106 , QK-Norm) Yes — Off bfloat16
24 1536 4608 — — 16 2 128 130560 SwiGLU (SiLU) RMSNorm (𝜖=10−6 ) RoPE No 12, 256, 1.4 On (1 layer) bfloat16
28 2048 8192 512 160 (16) 32 2 128 130560 SwiGLU (SiLU) RMSNorm (𝜖=10−6 ) RoPE No 12, 256, 1.4 Off bfloat16
transient regression, whereas tactical optimizations compose independently on top of a correct structure.
E
projection are captured into a single CUDA graph replayed once per step, with the backward pass, communication, and the optimizer left outside. Megatron-LM offers a wider graph spectrum—per-layer graphs, a full-iteration graph that includes gradient reduction, and a separately graphed optimizer step; the agent’s narrower boundary is an adjudicated outcome: the wider candidates (forward-plusbackward graphs, collectives and optimizer in-graph) were each tried in earnest and rejected by the gate for negative measured returns or an unsupporting runtime stack.
Optimization Catalog
This appendix details the optimizations summarized in §4.2.1 and §4.2.2. Framework level, second class (goals the reference shares, boundaries redrawn on measured returns). • Whole-forward graph capture. The roughly three hundred kernel launches from the embedding to the output 15
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
Table 11. Per-token FLOP budget (forward+backward) by operator class for the two MiniCPM4 scenarios at sequence length 4096, with each class’s share of the per-step total. The total is the numerator of the MFU computed in §4.1. 0.5B Operator class
8B
GFLOP/tok
Share
GFLOP/tok
Share
FFN GEMM (gate/up/down) Attention core (FlashAttention) Output / CE GEMM Attn. projection GEMM (Q/K/V/O)
1.812 0.604 0.451 0.340
56.5% 18.8% 14.1% 10.6%
38.655 3.221 1.805 6.845
76.5% 6.4% 3.6% 13.5%
Total
3.207
100%
50.526
100%
• Fused optimizer step. Gradient clipping, the AdamW update, and the master-to-bf16 write-back are fused into a single kernel (fused_clip_adam_sync), where MegatronLM keeps the three as separate steps.
with a small set of flags and performs no such automatic selection. • High-priority auxiliary streams. The streams carrying wgrad and gradient reduction are set to high CUDA priority, letting them land ahead of the main stream’s large GEMMs at SM scheduling boundaries and shortening the exposed communication and wgrad tail at the end of each step; Megatron-LM runs all its streams at default priority. Kernel level, the four recurring families. • Warp specialization and software pipelining. Threads within a block are split into producer roles that move data and consumer roles that compute, with every pipeline stage multiply-buffered and the constituent matrix multiplies issued asynchronously, so memory movement and computation overlap across stages rather than serializing. A general-purpose library fixes one thread-to-work mapping for a whole shape class; forging it for the engine’s exact dimensions recovers the occupancy that mapping leaves on the table.
• Optimizer–gather pipeline. On the 8B engine, the parameter all-gather is interleaved with the optimizer step bucket by bucket: the moment a bucket’s Adam update completes, the all-gather of its BF16 parameters is issued on the communication stream, ordered by forward consumption so it overlaps the next step’s forward, which waits only at bucket boundaries on the matching events. Megatron-LM’s overlap_param_gather likewise overlaps the all-gather with the next forward, but dispatches it lazily, bucket by bucket, from forward pre-hooks; the agent moved the dispatch into the optimizer step itself, interleaved with per-bucket Adam as a single pipeline. Framework level, third class (no Megatron-LM counterpart). • Asynchronous loss all-reduce. The scalar loss allreduce is issued on the communication stream and overlapped with the optimizer step, hiding it entirely in measurement; Megatron-LM leaves this reduction synchronously on the critical path.
• Memory-layout specialization. On-chip data is laid out for the operator’s exact shape—flattening accumulator buffers to remove shared-memory bank conflicts, vectorizing the load/store paths, and sourcing matrix-multiply operands directly from registers rather than staging them through shared memory. These are the layout choices a vendor kernel cannot hard-code without knowing the shape, and they convert otherwise-stalled memory pipes into useful throughput.
• Device-side gradient norm and clipping. The gradient norm, clip coefficient, and scaling are computed entirely on device—l2norm, all-reduce, square root, clamp, and scale in one uninterrupted pipeline—where Megatron-LM reads the norm back to the host (.item()) between norm and clip, placing a CPU–GPU synchronization on the hot path.
• Backward-pass fusion. An operator’s backward is computed by two matrix multiplies—one for the data gradient, one for the weight gradient—and, under sharding, a gradient reduction; the agent folds all three into a single kernel, sparing the launch overhead and intermediate write-back that a vendor library’s separate dgrad, wgrad, and reduce kernels pay.
• Zero-copy gradient views. Having established that its fused Adam kernel only reads gradients, the agent had the all-gather return views rather than copies, eliminating roughly one hundred and fifty clones and about 2 GB of peak memory per step. • Per-shape operator selection. A dispatcher benchmarks candidate implementations across CuTeDSL, cuBLAS, Triton, and Transformer Engine and selects, per operator shape, the kernel of highest model-FLOPs utilization (MFU); Megatron-LM relies on a fixed backend
• Cross-kernel scheduling. Adjacent kernels are overlapped across their launch boundary through programmatic dependent launch, and computation that the problem structure renders redundant is pruned outright (for 16
ForgeTrain
Figure 8. Optimization trajectory of the FlashAttention backward kernel (𝐵=10, 𝐻 =16, 𝑁 =4096, 𝐷=64, FP16, GQA 8:1, causal; H100 SXM). Top: backward latency (lower is better) across milestones. A structural build-up phase (warp-specialized threekernel pipeline, 128-row 𝑀-tile, double-buffered stages) is mutually dependent and regresses latency to 6.3–6.9 ms before a d𝑄 data-layout refactor (flat global buffer, register-to-shared copy, register-sourced GEMM operands) drops the kernel to 2.22 ms; three independent tactical steps—128-bit vectorized preprocess/postprocess copies, diagonal-only causal masking, and removing a stray synchronization that had disabled programmatic dependent launch—then reach 1.90 ms (45.7% MFU). Middle: the four tactical steps magnified from the dashed region of the top panel, on a 1.8–3.1 ms scale, against the cuDNN and FlashAttention-4 end-to-end reference lines; the gains are individually small (2.22 → 2.03 → 1.97 → 1.90 ms) and are what the top panel cannot resolve. Bottom left: per-kernel NCU time breakdown versus FlashAttention-4 (forged sum 2072 µs vs. 1990); our per-kernel times are the slower ones, and the two post-d𝐾 kernels of FA-4 appear fused at 0.024 ms in ours. Bottom right: end-to-end backward latency—against FlashAttention-4 (2.007 ms, 43.3% MFU) the forged kernel is 5.6% faster and against cuDNN (2.483 ms, 35.0% MFU) 31% faster, despite the slower per-kernel sum: the entire margin comes from programmatic dependent launch hiding 0.17 ms across our four kernels, where FA-4’s five show no measurable overlap.
F
a causal operator, evaluating the mask only on the diagonal block). Both recover time that lives between or inside kernels, where a single-kernel benchmark cannot reach.
Why the Engine Is Faster, Mechanism by Mechanism
§4.2.3 traces the forged 0.5B engine’s margin over MegatronLM 0.15 to three sources. This appendix expands each source with its source location and key code. Citations are file:line into the exported ForgeEngine package; snippets are trimmed of unrelated lines only. 17
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
Reproduced standard set. The optimizations MegatronLM already has, reproduced component for component with the same boundaries (nccl.py:152,654 — bucketed reducescatter overlapped with backward; optimizer.py:668 — ZeRO-1 sharding; triton_kernels.py:362,699 — fused RMSNorm, :208,268 — SwiGLU, :1005,1166 — RoPE, :112 — fused cross-entropy; and the direct Transformer-Engine attention path). These were rediscovered in the first nine steps of Table 8. Vertical integration. Components Megatron keeps as separate stages, merged into one execution path. The optimizer tail is the clearest case: Megatron runs gradient clipping, the Adam update, and the master-to-bf16 writeback as three passes over every parameter; the engine fuses all three into one Triton kernel (optimizer.py:357-385, triton_kernels.py:415-446).
def compute_clip_coeff_device(grad_norm_sq, clip_coeff_buf, max_norm, eps): grad_norm = grad_norm_sq.view(()).clamp(min=0.0).sqrt() coeff = (max_norm / (grad_norm + eps)).clamp(max=1.0) clip_coeff_buf.view(()).copy_(coeff) return clip_coeff_buf
# triton_kernels.py:544 -- kernel reads per-step scalars from buffers. lr = tl.load(lr_ptr).to(tl.float32) clip = tl.load(clip_coeff_ptr).to(tl.float32) bc1 = tl.load(bc1_ptr).to(tl.float32) bc2 = tl.load(bc2_ptr).to(tl.float32)
Because the values live in buffers rather than frozen launch arguments, the host refreshes them once per step (optimizer.py:492-524) and the replay stays correct; keeping the scalars on the device is what lets the step avoid synchronizing with the host. The remaining deepcustomization points — the loss all-reduce moved off the critical path, and the read-only, layout-specific paths specialized to the fixed shapes — complete the set.
# optimizer.py:357 -- one fused_clip_adam_sync does clip + Adam + cast. def fused_clip_adam_sync(state, fp32_grads, params, lr, clip_coeff): state.step_count += 1 for name in state.param_names: g = fp32_grads[name].float() fused_adam_sync(g, master, exp_avg, exp_avg_sq, bf16_param, lr=param_lr, step=state.step_count, clip_coeff=clip_coeff, wd=wd)
G
Ascend Framework-Level Optimization Trajectory
Table 12 lists, in acceptance order, the gated optimization rounds behind Table 2 for the MiniCPM5 130M engine, together with each round’s own reported avg MFU(standard) over its gate window. This is the per-round detail behind that trajectory, analogous to Table 8 on H100. Unlike the H100 trajectory, which is governed by a single MFU target throughout, the Ascend log applies two gates in sequence — perf-bitwise (≥ 13%) then long-horizon (≥ 25%) — so the phase column marks which gate each round is scored against.
# triton_kernels.py:415 -- one pass per element: clip, Adam, cast. g = tl.load(grad_ptr+offs) * clip_coeff m = beta1*m + (1-beta1)*g v = beta2*v + (1-beta2)*g*g m_hat = m / bias_correction1 v_hat = v / bias_correction2 update = m_hat / (tl.sqrt(v_hat) + eps) + wd * p p_new = p - lr * update tl.store(master_ptr+offs, p_new) tl.store(bf16_ptr+offs, p_new.to(tl.bfloat16))
H
Ascend Model Architectures
Table 14 lists the geometry for the two Ascend scenarios referenced throughout this section, analogous to Table 9 on H100. The 130M geometry (18 layers, minicpm5_130m) is the one the recorded forging rounds and the bit-for-bit anchors of Table 13 were measured on, chosen for cheaper iteration; the 1B geometry (24 layers) is the engine pretrained end-to-end in §4.3.1. Two points are worth flagging rather than smoothing over. First, the 1B scenario’s FLOP accounting (Appendix I) uses the full MiniCPM5 tokenizer vocabulary (130560) because that is the field config.py’s _global_flops_per_token reads, whereas the 130M scenario pads to 73448 to match the pre-tokenized data shards it trains against — a real difference in what the two model families’ embedding/LM-head matrices are sized to, not a reporting inconsistency. Second, the 1B config carries a one-layer Eagle (MTP) head that the 130M scenario does not; its forward-plus-backward FLOPs
The same source also merges the per-layer graphs Megatron builds into one captured CUDA graph over a whole microbatch’s forward, loss, and backward (step_graph.py:235-260). Deep customization. Points with no Megatron counterpart, where the engine rewrites the execution path around the fixed model, shapes, and parallel layout. The optimizer scalars and clipping coefficient are the cornerstone: instead of computing the clip coefficient on the host from a gradientnorm .item() round trip and passing the resulting Python floats as kernel arguments (which would then be frozen at capture), they stay on the device, in 1-element FP32 buffers the kernel reads via tl.load (optimizer.py:420-457, optimizer.py:460-524, triton_kernels.py:525-566). # optimizer.py:420 -- clip coefficient computed entirely on device. 18
ForgeTrain
Table 12. Gated optimization rounds applied to the MiniCPM5 130M engine (18 layers), in acceptance order. MFU is that round’s own avg MFU(standard) over its gate window (Ascend 910, DP-only). #
Gate
Optimization
1 2 3
perf-bitwise (≥ 13%) perf-bitwise (≥ 13%) perf-bitwise (≥ 13%)
Baseline (bitwise-preserving) Mask/RoPE cache + foreach grad ops Async grad-hash overlap
9.67 9.73 15.27
4 5 6 7 8 9 10 11
long-horizon (≥ 25%) long-horizon (≥ 25%) long-horizon (≥ 25%) long-horizon (≥ 25%) long-horizon (≥ 25%) long-horizon (≥ 25%) long-horizon (≥ 25%) long-horizon (≥ 25%)
det-off + chunked CE (memory) Dataloader prefetch CE upcast fused into log-softmax Reuse forward log-softmax in backward Phase-breakdown + host-sync lift foreach-batched AdamW Fused NPU cross-entropy (FP32) Defer grad-norm .item() off hot path
19.21 20.10 20.05 21.87 22.04 22.08 24.87 24.88
Table 13. Bit-for-bit anchor counts captured from the golden reference. All listed tensors must match exactly; MiniCPM5 130M uses blake2b equality. Scenario MiniCPM4 0.5B MiniCPM4 8B MiniCPM5 130M (Ascend)
Forward
Backward
Total
183 164 128
339 195 237
522 359 365
lower still (15.2%), because the one-layer Eagle (MTP) head – which repeats the backbone layer’s attention, projection, and FFN plus a second output head – absorbs 18.3% of the budget and dilutes the backbone shares. This is the reason the fused-CE win of §4.2.1 (which sits inside the outputhead and attention-adjacent work) moves MFU by a larger fraction on the 130M geometry than an equivalent fusion would on the H100 0.5B geometry. Confirming that this numerator matches §C’s methodology narrows — but does not close — the denominator question: what remains unverified is whether MindSpeed’s own –log-throughput uses the same per-operator count on its side of Table 2, or a coarser approximation.
are the last row of Appendix I’s table, bringing the 1B total to 7.940 GFLOP/tok, while the 130M total stays at the backbone-only 1.070.
I
MFU (%)
Ascend Model-FLOPs Utilization
J
The Ascend engines report MFU against the same exact peroperator FLOP count as §C, not the closed-form approximation: both engines compute the sum of attention projections, FFN (SwiGLU), attention core, and output head per layer, forward, then multiply the total by 3 for forward-plus-backward — the identical accounting §C uses, down to the same 1:2 forward-backward ratio applied to the output head. The 1B engine adds the forward FLOPs of its one-layer Eagle (MTP) head to this sum when MTP_ENABLED is set, and we include that head as a separate row in Table 15; the 130M engine has no such head, so its total is backbone-only. We reproduce the count here (Table 15) rather than take the engines’ self-reported denominators on faith: applying the formula to the 0.5B/24-layer H100 geometry first reproduces Table 11’s 3.207 GFLOP/tok exactly, which is what justifies applying the same formula to the Ascend geometries of Table 14. The attention core is a noticeably larger share on the Ascend 130M scenario (26.5%) than on the H100 0.5B scenario (18.8%, Table 11): the 130M model runs the same 4096 sequence length as H100 but with fewer, narrower heads (10 × 64 versus 16 × 64), so the FFN does not dominate the budget as heavily. On the 1B scenario the attention share is
Ascend Optimization Catalog
This appendix supplements §4.2.1 with additional optimizations recorded in the exported source but not detailed in the main text, in the same two-tier shape as Appendix E. Citations are file:line references into the exported training_engine_tensor packages. Framework level (goals MindSpeed shares, structure redrawn on Ascend). • No autograd engine. Both engines run their entire step under torch.no_grad() with a hand-written static backward that replays the reference’s own primitive backward ops in a fixed, pre-determined order (backward.py:1-41). MindSpeed’s Megatron core builds and walks an autograd graph every step; the forged engines skip that construction and traversal entirely, at the cost of hand-maintaining the backward op sequence themselves. • Host-sync removal on the optimizer path. The 130M engine lifts the optimizer step counter into a host Python int (train_loop.py:896-904) and defers the grad-norm .item() call past the step-end device sync (optimizer.py:207-240), removing two CPU–GPU synchronization points from the hot path — the same class 19
Qingfeng He, Zhui Zhu, Shangzhan Li, Yaojian Chen, Haojun Sun, Xu Chen, Leshan Li, Yifei Shen, Changjingxing Zhao, Mengyuan Fan, Wenyu Guan, Yiyun Zheng, Yuxuan Zuo, Zhen Li, Zhenghang Luo, Yuxuan Li, Xu Han, and Zhiyuan Liu
Table 14. MiniCPM5 architecture configuration for the two Ascend scenarios. Hyperparameter Transformer layers Hidden size FFN intermediate size Attention heads Key/value heads (GQA) Head dimension Vocabulary size (as used in FLOPs) Activation Normalization Position encoding Sequence length Tied input/output embeddings 𝜇P (emb/depth/base-hidden) MTP / Eagle head Training precision
130M
1B (production)
18 640 1920 10 2 64 73448 SwiGLU RMSNorm (𝜖=10−6 ) RoPE (base 10000) 4096 Yes 12.0, 1.4, 256 Off bfloat16
24 1536 4608 16 2 128 130560 SwiGLU RMSNorm RoPE 4096 — 12.0, 1.4, 256 On (1 layer) bfloat16
Table 15. Per-token FLOP budget (forward+backward) by operator class for the two Ascend scenarios at sequence length 4096, computed with the §C formula. The 1B total includes its one-layer Eagle (MTP) head, shown as its own row; the 130M scenario trains without MTP, so that row is not applicable there. 130M Operator class
1B (production)
GFLOP/tok
Share
GFLOP/tok
Share
FFN GEMM (gate/up/down) Attention core Output / CE head Attn. projection GEMM (Q/K/V/O) Eagle / MTP head
0.398 0.283 0.282 0.106 —
37.2% 26.5% 26.4% 9.9% —
3.058 1.208 1.203 1.019 1.452
38.5% 15.2% 15.2% 12.8% 18.3%
Total
1.070
100%
7.940
100%
of fix as H100’s “device-side gradient norm and clipping” catalog item (Appendix E), independently rediscovered on Ascend.
principle as H100’s “zero-copy gradient views” catalog item (Appendix E), applied on the parameter side instead of the gradient side.
• Chunked / streaming cross-entropy. Beyond the fused FP32 CE operator described in §4.2.1, the 1B engine carries a streaming CE path that folds the LM-head GEMM, the CE, and both gradient passes into one chunked loop that never materializes the full logits tensor at all — not even a row-chunk at a time (backward.py:276-368). It is implemented and correctness-tested but not enabled in production (train_loop.py:1279, _streaming_ce=False); the in-place chunked path actually running in production (§4.2.1) is the more conservative of the two. Framework level (no MindSpeed counterpart). • Weights as all-gather-buffer views. The 1B engine’s bf16 parameter buffers are not copied out of the all-gather receive buffer after each collective; the weights are views into it, replaced in place on the next sync (optimizer.py:1521-1552), sparing roughly 1.8 GiB of duplicate bf16 weight storage. This is the same zero-copy
• Batched Newton–Schulz for Muon. The 1B optimizer’s Muon branch stacks same-shaped 2D parameters and runs Newton–Schulz iteration as batched bmm calls instead of one launch per parameter, cutting roughly 2500 kernel launches to 60 (optimizer.py:93-118) — an optimizerlevel fusion with no Megatron-LM-side analog in §4.2.2, since the H100 engine does not run Muon. Kernel level (1B custom AscendC GEMM family; structural evidence only — see the header note on why no benchmark table is given here). • Per-shape specialized tiling. Fifteen AscendC GEMM kernels (five projections × three directions: forward, dgrad, dwgrad) hard-code 𝑀/𝐾/𝑁 as compile-time constants and tile dimensions sized to fill the 910’s L1/L0C capacity exactly (custom_gemm.py:59-91; e.g. fc2_fwd_entry.cpp:16-33 fixes AIC_CNT=20, baseM=128, shareL1Size= 516096), the same per-shape 20
ForgeTrain
specialization principle as H100’s kernel-level catalog (Appendix E), retargeted to a different vendor ISA. • Transpose elimination via isTrans. The weight and gradient-output tensors are consumed in their natural physical layout rather than materialized-transposed first (fc2_fwd_entry.cpp:81,114), sparing a full-tensor HBM round-trip the vendor GEMM path would otherwise pay on the weight gradient. • Direct dispatch, bypassing the operator registry. The custom kernels are invoked via ctypes against a compiled .so, passing raw data_ptr()s and the current stream directly (custom_gemm.py:247-263,359-363), skipping the aten dispatcher, TensorIterator, and the torch_npu aclnn wrapper layer entirely.
(a) MiniCPM5 1B
(b) MiniCPM5 130M
Figure 9. Per-setting MFU trajectories of the two Ascend 910 settings from Table 2, each normalized to its own MindSpeed baseline (dashed). The x-axis is the setting’s 𝑘-th recorded forging round.
21