arXiv:2609.10632v1 [cs.SE] 9 Sep 2026
Numbat: Building and Verifying a Self-Contained Machine-Learning Stack Thang Tran∗
Lan Dang
CloudKites AI Lab New South Wales, Australia [email protected]
Monash Business School, Monash University Victoria, Australia [email protected]
September 2026
Abstract Machine-learning systems are built almost exclusively on a few large Python-orchestrated frameworks, and they inherit those stacks’ engineering costs: environments of hundreds of version-coupled packages, separate export toolchains for deployment, and the split between the language research is written in and the language products ship in. We report on the construction and verification of numbat, a machine-learning stack written in one general-purpose language (Zig) with no third-party runtime dependencies. The stack spans tensor computation, automatic differentiation, neural-network modules, mixed precision, multi-GPU training, data loading and monitoring; an SDK exposes it behind a stable, additively versioned C ABI of over 1,400 entry points, with bindings for six languages; and its clinical domain planes encode regulatory requirements as executable acceptance gates rather than documentation. Verifying such a stack is the harder half of building it: a defective training run rarely fails, it converges quietly to a slightly worse model. We treat a widely used reference implementation as an executable specification and verify against it at five levels, from operator gradient checks to an automated trajectory gate against a same-machine reference run — the arrangement our companion study formalizes as a trajectory-level differential oracle. The protocol surfaced ten silent recipe divergences, which we catalog with mechanisms and symptoms. As the acceptance test, we train a 25.9M-parameter detector of the YOLOv8m class from random initialization on COCO 2017 for the full 500-epoch schedule: the exported weights score 0.4956 mAP50–95 under the official protocol, scored by the reference stack’s own validator (published endpoint 0.502), with single-GPU step time at parity on identical hardware. Weights, per-epoch metrics and the full run manifest are released.
1
Introduction
Contemporary deep-learning practice is concentrated on a small number of large Python-orchestrated frameworks [1, 5, 27]. These systems are mature and productive. They are also expensive in ways their own authors acknowledge [2, 27]. Per-operator dispatch through the interpreter adds host-side overhead, which PyTorch 2 mitigates with an additional bytecode-capture and graph-compilation layer [2] at the price of further system complexity. A GPU-enabled installation spans gigabytes across hundreds of packages, and the version coupling among framework, accelerator libraries, driver and Python dependencies makes environments costly to reproduce and audit. Deployment to edge, embedded, mobile or air-gapped targets goes through separate export toolchains rather than ∗
Corresponding author: [email protected]
1
through the stack that trained the model, yielding the familiar two-language pattern — research in Python, production re-implemented in a systems language — which duplicates engineering and invites divergence. Organizations that must audit their supply chain, or bound memory behavior on constrained devices, inherit all of it. The demand for self-contained runtimes is directly evidenced by the adoption of inference engines such as llama.cpp [8] and ONNX Runtime [25]; such systems deliberately stop short of training. Independent training stacks outside the mainstream ecosystem exist [10, 14, 33], but public evidence that one can reproduce the training outcome of a mature, competitive recipe — rather than merely execute gradient descent — remains scarce. numbat removes these costs at their root: one dependency-free stack in a single general-purpose language, covering training through deployment, with applications shipping as single static binaries. Building it, however, is only half the work, and the smaller half. A modern training recipe is specified de facto by its implementation rather than by any paper: initializer distributions, mixed-precision autocast placement, optimizer parameter grouping, exponential-moving-average (EMA) semantics, gradient-clipping policy, loss-scaler dynamics, augmentation details and random-number-generator (RNG) structure are behavioral properties of code, individually undocumented and individually capable of degrading the final metric without any overt failure. This is a recognized reproducibility problem even within one framework [4, 13, 28]; across independently implemented frameworks it is the central difficulty, and it is a software-engineering difficulty — a verification problem with no oracle — before it is a numerical one. In a companion study we formalized the two-stack arrangement as a trajectory-level differential oracle and applied it to language-model fine-tuning, where the comparison exposed seventeen faults that single-implementation development had missed [35]. This paper is about the stack itself: how it is put together, how its interfaces are kept stable, how domain requirements are made executable, and what it took to reproduce a competitive from-scratch training outcome end to end. The paper makes five contributions. 1. Design. We describe numbat, a self-contained ML framework written in Zig covering tensors, autograd, modules, mixed precision, distributed data-parallel training, data pipelines and monitoring, with CPU and CUDA backends in production use and vendor-neutral GPU paths including a single-source kernel plane compiled by the framework’s own compiler (§3). One design principle carries most of the weight: reference-default semantics — numbat adopts the dominant ecosystem’s defaults (initializers, autocast policy, optimizer and loss-scaler defaults, RNG streams) as its own, so recipes transfer across stacks unchanged. 2. Interface discipline. The numbat SDK packages the full training and inference capability behind one C ABI that is versioned and strictly additive, with thin bindings for six languages whose numerical agreement is enforced by shared tests and, at trajectory level, evidenced in the companion study (§4). 3. Requirements as executable gates. For the clinical domains the stack targets, design requirements are encoded as acceptance gates — executable checks with verdicts and exit codes — backed by typed diagnostic registries whose coverage is measured rather than asserted (§5). We describe the method and what it caught. 4. Reproduction protocol and divergence catalog. We define a five-level cross-framework training-reproduction protocol — operator, module, step, trajectory, outcome — with an automated trajectory gate against a same-machine reference trajectory (§6), and we catalog the ten silent recipe divergences it surfaced, with mechanisms and symptoms (§7). The catalog is of independent value to anyone porting a training recipe between frameworks. 2
5. End-to-end evidence. We train YOLO-NB-M, a YOLOv8m-class detector, from scratch on COCO 2017 for 500 epochs on three power-capped consumer GPUs, tracking the reference trajectory throughout, and report accuracy under the official COCO protocol as scored by the reference stack’s own validator on exported weights, plus same-machine throughput, memory and energy accounting, including negative results (§8). Weights, metrics and the run manifest are released (§9).
2
Related work
Frameworks and runtimes. PyTorch [27], TensorFlow [1], and JAX [5] dominate training workloads and anchor a large ecosystem of Python packages. Outside this ecosystem, MLX [10] targets Apple silicon, Burn [33] provides a Rust-native training stack, and tinygrad [14] explores a minimal operator set; llama.cpp/ggml [8] and ONNX Runtime [25] demonstrate the deployment value of self-contained inference. numbat differs in combining (i) a single-language, dependency-free implementation of the full training stack, (ii) explicit behavioral compatibility with PyTorch defaults, and (iii) a packaged SDK exposing that stack from six languages behind one stable ABI (§4). The YOLO family. Single-stage real-time detection descends from YOLO [29, 30] through successive architecture generations [3, 16, 18, 37]. The eighth-generation “v8” class combines a crossstage-partial (CSP) backbone [36], a path-aggregation feature-pyramid (PAN-FPN) neck [22, 23], an anchor-free decoupled head with distribution focal loss (DFL) box regression [20], task-aligned label assignment [7], and complete-intersection-over-union (CIoU) box loss [39]. We train a detector of this architecture class (25.9M parameters at medium scale) as our case study. Training detectors from scratch. DSOD [32] and subsequent analysis [12] established that detectors can be trained from random initialization to pretrained-equivalent accuracy given sufficient schedule length; modern YOLO recipes are routinely trained from scratch. Our contribution is orthogonal: not whether from-scratch training works, but whether an independent framework can reproduce a reference recipe’s outcome exactly, and what silently breaks along the way. Reproducibility and cross-stack validation. Sensitivity of outcomes to seemingly minor implementation details is well documented [4, 13, 28]. We extend this line of work to the crossframework setting, where the reference implementation must be treated as the specification, and provide an operational protocol for verifying convergence-level — not merely operator-level — parity. The companion study [35] formalizes the underlying idea: two independently implemented stacks running one specification act as differential oracles for each other at the level of the learning trajectory, addressing the oracle problem that makes training pipelines hard to test at all. There the workload was language-model fine-tuning and the second stack was numbat driven through its own SDK; here the same discipline is applied to a from-scratch detector run, and to the construction of the stack itself.
3
The numbat framework
numbat is written in Zig.1 The language choice is doing real work, not signalling: memory is managed explicitly through caller-supplied allocators, there is no hidden runtime and nothing that 1
https://ziglang.org
3
Applications
model zoo: 100+ models (LLM · VLM · vision · OCR · speech · medical) · trainers io: nbq (native container) · safetensors · GGUF · NPY · PyTorch pickle · NIfTI · codecs
Modules & training services
nn layers · detection/LLM losses · SGD/Adam(W) · EMA (incl. buffers) · fused grad-clip AMP (f16 compute, f32 BN+loss) · DDP (shared-mem all-reduce · NCCL) · prefetch loaders atomic resumable checkpoints · metrics streaming · trajectory gates
Tensor & autograd generic Tensor(T), 12 dtypes, 64-B-aligned storage · views/broadcast · tape autograd incremental activation freeing · Philox RNG (reference bit-parity) · gradient checks
Backends & kernels
CPU: thread pool · SIMD · Winograd conv · optional BLAS CUDA: driver-API kernels · selective cuDNN/cuBLAS · NVRTC-JIT fusion (3 tiers) ROCm · Metal · Vulkan · WebGPU (WGSL)
Hardware
x86-64 SIMD · NVIDIA GPU · AMD GPU · Apple GPU · WebGPU targets
numbat SDK libnumbat — one library for inference AND training stable C ABI (1,400+ functions) opaque handles · status codes versioned, additive Bindings (open, thin): Zig · C · Python · Rust Go · TypeScript — identical numerics in all six clinical domain planes (§5) 7 dtypes · autograd · SDPA nn · 14 optims · LLM generation image/audio io · safetensors prebuilt library + thin bindings — ready to use from all six languages
Figure 1: The numbat stack and the numbat SDK. The framework is one vertically integrated codebase in a single general-purpose language (Zig); the SDK (§4) packages its full capability as prebuilt libraries behind a stable, versioned C application binary interface (ABI) with thin open bindings, so applications in six languages build on the framework directly.
can pause a training step to collect garbage, and the compile-time metaprogramming is strong enough that the entire tensor layer is built with it rather than with a code generator. Figure 1 lays the stack out. Four decisions shaped it; each is stated below along with what it buys. (i) Single-language vertical integration. Tensor storage, kernels, autograd, modules, optimizers, distributed training, data decoding (including image codecs), tokenization, and monitoring are one codebase with no foreign-function seams except vendor GPU libraries. Applications ship as single static binaries. (ii) Explicit, deterministic memory. All allocations flow through caller-supplied allocators; tensor storage is 64-byte aligned; lifetimes are explicit. During backpropagation, saved activations can be freed incrementally as the tape drains; on the case study of §6 this reduces peak training memory per GPU from 18.5 to 13.5 GiB (−26.8%) with bitwise identical gradients. (iii) Compile-time specialization. The tensor type is generic over its element: Tensor(T), where T ranges over twelve types, from booleans and the integer widths up to the four floating-point precisions (F16, BF16, F32, F64). Both the tensor and the backend dispatch are monomorphized when the program compiles. The consequence matters more than the mechanism: at run time nothing stands between model code and the kernels — no dispatch layer to cross, no interpreter, no graph VM to feed. (iv) Reference-default semantics. Where the dominant ecosystem has a default — initializer distributions, AMP autocast placement, optimizer hyperparameter defaults, loss-scaler constants, RNG algorithms — numbat adopts it as its default. The counter-based Philox generator [31] is bit-compatible with the reference’s GPU uniform sampling (verified by known-answer tests), and the initializer suite reproduces torch.nn.init distributions exactly. Section 7 shows why this principle is load-bearing: most convergence divergences we found were violations of it.
4
3.1
Backends and kernel generation
On the CPU, kernels are vectorized — single instruction, multiple data (SIMD) — and scheduled by a work-stealing thread pool; convolutions go through the Winograd transform [17] at the sizes where it pays, and an external BLAS (Basic Linear Algebra Subprograms) library can be linked in for the matrix products, though none is required. The CUDA backend talks to the driver API directly. Vendor libraries are used where profiling says they win — cuBLAS throughout, cuDNN for particular convolutions — and fused kernels are generated at three tiers: a curated set of hand-written fused kernels; a runtime code-generator that compiles arbitrary elementwise chains via NVIDIA’s runtime compiler (NVRTC), cached by operation signature; and a fusiongraph compiler in the spirit of Triton [34] that lowers a typed operator directed acyclic graph (DAG) into two kernel templates (pure elementwise with broadcast indexing; reduction with fused pre/post epilogue, covering the normalization family). On a production connectionist-temporalclassification (CTC) speech-recognition encoder, graph fusion reduced per-batch latency from 258 ms to 23 ms. The ROCm, Vulkan, Metal, and WebGPU backends follow the same layering. Two of them are vendor-neutral, and they bracket an engineering progression. The Vulkan backend — GLSL compiled ahead of time to SPIR-V — is correctness-complete across the operator set with automatic differentiation on both forward and backward, and reaches the device’s tensor cores through half-precision cooperative-matrix multiplication. The newer single-source kernel plane goes further: device kernels are written in Zig and compiled to SPIR-V by the same compiler that builds the framework, then dispatched over the Vulkan runtime that the display driver already provides. The compiled kernel pack is embedded in the library, so the GPU path requires no vendor toolkit, no SDK and no shader compiler on the target machine, at build time or at run time — the machine needs a display driver and nothing else. The WebGPU backend is single-precision and inference-oriented, and the Metal backend is an early stub.
3.2
Training subsystem
Optimization. Stochastic gradient descent (SGD) and Adam(W) carry reference-default semantics throughout, down to details that turn out to be load-bearing: parameter groups (the case study needs the reference’s three-group structure, §7), fused multi-tensor update paths, and gradient accumulation that preserves nominal-batch semantics. Gradient clipping is worth its own sentence. The global-norm form is a single fused kernel — one squared-norm accumulation across all gradients, one device-to-host readback — where a naive port performed 243 per-tensor synchronizations; on the case study that is the difference between 46–71 ms and 7 ms per step. Mixed precision. Automatic mixed precision (AMP) follows reference autocast semantics [24]: F16 compute for convolutions and matrix products, F32 master weights, and — critically (§7) — F32 batch normalization (BN) [15] and loss, with dynamic loss scaling. Distributed data parallelism (DDP). One process per GPU, following Li et al. [19]; gradients are reduced via a host-staged, sharded all-reduce over POSIX shared memory with pinned staging buffers and futex-based barriers, producing rank-order-deterministic, bitwise-reproducible sums. An overlap engine that reduces buckets concurrently with backpropagation, and a transport over the NVIDIA Collective Communications Library (NCCL), are implemented and gated off by default on small-core hosts (§8.2 reports the measured reason honestly). Data pipeline. Background prefetch loaders perform decode and augmentation in worker threads. All image decoding is in-tree and reference-parity tested; a notable byproduct of parity testing was the discovery of a quantization-table ordering bug in our JPEG decoder that silently perturbed every decoded pixel (mean absolute error 9–20 intensity levels vs. libjpeg) — exactly the class of
5
silent divergence the protocol of §6 exists to catch. After correction and vectorization the decoder sustains 1.88× its previous throughput and the loader fully overlaps GPU compute on the case study. Monitoring and gating. Every run streams its metrics, as CSV and JSONL, to nbmonitor — numbat’s own experiment tracker, self-hosted behind its own accounts and API keys. Functionally it does what the hosted trackers do: live charts, run search and comparison, per-GPU telemetry down to power and temperature, logs, media, sweeps. The difference is where it runs. Training telemetry never leaves the organization, which for clinical work is a requirement rather than a preference. The more important consumer of the same stream is not a person at all: a trajectory gate sidecar polls each epoch and can pause or terminate a run that leaves its reference band (§6.3), because a dashboard someone glances at twice a day cannot stop a run that went wrong at 3 a.m. Checkpoints are atomic and carry optimizer, EMA, epoch and RNG state, so a run resumed mid-schedule continues as if it had never stopped.
3.3
Interoperability and model zoo
A framework that cannot read the ecosystem’s files is an island, so numbat reads and writes the formats models actually arrive in: safetensors, GGUF with its k-quantizations, NPY, PyTorch pickle checkpoints, and the medical and audio formats its applications need (NIfTI-1 among them). Its tokenizer suite is byte-identical to the reference tokenizer library on validation corpora — byteidentical rather than merely compatible, because §7 is a catalog of what "merely compatible" costs. For deployment, numbat defines .nbq, its native self-contained model container — a novel format developed at CloudKites AI Lab, for which a patent application is in preparation. A single .nbq file carries everything a production application loads: the (optionally quantized) weight buffers already in their device-ready layout, the tokenizer, the model configuration, and an authenticated, self-describing manifest recording each tensor’s dtype, layout, and role together with the target backend and minimum device capabilities — so a runtime can decide compatibility, with a precise reason on mismatch, before reading any weight data. The payload is page-aligned and laid out contiguously in load order, so a model starts with zero repacking or conversion: memory-mapped on CPU (zero-copy) or transferred to the accelerator in a single host-to-device copy, with per-tensor device pointers recovered by offset arithmetic. The model zoo counts more than a hundred complete implementations (102 at the time of writing) spanning large language models, vision–language models, vision (classification, detection, segmentation), optical character recognition (OCR), speech recognition and synthesis, speaker and audio analysis, video, and a dedicated medical and clinical group (nnU-Net, MedSAM2-class segmenters, electrocardiogram models, endoscopy detectors); LLaMA-3, Qwen-3, Whisper, SAMclass segmenters, DINOv3 ViT and the YOLO family are examples, not an inventory. Each implementation is the architecture rewritten in Zig, loading the originally published weights and checked against the original’s output. This breadth is exercised in production: numbat is the sole inference engine of emu, a freely distributed, local-first multimodal desktop application for Windows and Linux2 that runs language, vision–language, OCR, text-to-speech, and speech-to-text models entirely on the user’s machine (CPU or NVIDIA GPU) — the single-library deployment model of §4 operating in the field. Weight-level interoperability is what the methodology stands on: numbat-trained weights are exported to safetensors and evaluated by third-party stacks (§6.3). 2
https://huggingface.co/cloudkites/emu
6
4
The numbat SDK
Adopting a new framework should not require adopting a new toolchain, nor should the framework’s internals become part of every downstream build. The numbat SDK packages the framework as one ready-to-use library for building end-to-end machine-learning applications: training, evaluation and inference live in the same stack, so the code that fits a model in a research prototype is the code that serves it in production. The library is self-contained and cross-platform (Linux, macOS and Windows; x86-64 and ARM64; shared and static artifacts, plus a WebAssembly build for the browser), and the whole platform matrix is cross-compiled from a single development machine — a property of the Zig toolchain rather than of any release infrastructure. This answers the deployment-side costs catalogued in §1 directly: no interpreter in the serving path, no multi-gigabyte environment to reproduce, no separate export toolchain between the stack that trains a model and the stack that ships it, and no two-language rewrite between prototype and production. The SDK is structured in three layers (right column of Fig. 1). 1. Core: the framework of §3, built as prebuilt shared and static libraries per platform, with a single C header. 2. C ABI: one library, libnumbat, exposing inference and training through what is now more than 1,400 nb_* entry points — the tensor and module core, and the domain planes of §5, which account for most of the surface. Only opaque handles, runtime-tagged dtypes and status codes cross the boundary: no framework headers, no memory-layout assumptions. That is what keeps the interface stable. The ABI is versioned and strictly additive — at revision 3 at the time of writing, with the two most recent releases adding 282 entry points between them without breaking a single existing caller — so applications built today keep working as the framework evolves underneath them. Errors are returned as status codes with structured, thread-local detail; all handles are caller-owned with explicit destructors. 3. Bindings (deliberately thin): Zig, C, Python, Rust, Go and TypeScript, all over the same ABI. Each presents a PyTorch-shaped API (numbat ≈ torch, numbat.nn, numbat.optim) so that ecosystem experience transfers directly. The Python binding depends on the Python standard library and nothing else — deliberately not NumPy — and the other bindings likewise pull no third-party packages; the zero-dependency property of the framework extends to everything a consumer links. What crosses the boundary is the working vocabulary of the ecosystem, not a reduced export subset. Tensors come in seven runtime dtypes with casting and full autograd control (detach, retain_grad, no-grad mode); scaled dot-product attention is differentiable through the boundary, with causal and additive masks and grouped-/multi-query heads, on CPU and CUDA. The nn modules run from convolutions through pooling, normalization, embeddings and activations — enough to assemble ResNet- and transformer-class architectures from any binding — and training is served by the torch-parity optimizer and scheduler families, gradient clipping, EMA weight averaging, seeding, torch.nn.init-parity initializers, and losses from cross-entropy through the margin and ranking families. Data loading, module save/load, and I/O for safetensors, NumPy, NIfTI, images and audio round it out. A language-model path runs end to end through the same boundary: a GGUF-packaged model loads, tokenizes and generates text, with a complete sampling toolkit, from any of the six languages. And the models in the SDK’s zoo are not opaque handles. Each is composed in the binding’s own language from those same nn primitives, so a developer reads — and can change — the architecture in the language they already use. 7
Cross-binding numerical agreement is enforced at two levels. A shared known-answer test fixes a training task whose loss trajectory must be identical from all six languages, and a parity harness sweeps the operator surface. At trajectory level, the companion study provides the strongest evidence: four implementations of one fine-tuning specification — PyTorch, numbat driven natively, and the SDK driven from Python and from Zig — ended a full training epoch within 0.15% of one another in held-out cross-entropy, with 42 paired evaluations differing by 0.134% on average [35]. Notably, the same study found that four of the seventeen faults it exposed were reachable only from a language whose memory model differed from the others’; the six-language surface is a testing asset, not only a convenience. The SDK is assembled into versioned releases — the prebuilt platform libraries, the header, the binding sources, documentation and worked examples — and every release is verified by installing the built artifact and interrogating it, a step §5 motivates. GPU support is an opt-in build of the same library, with vendor runtimes loaded dynamically. Beyond application software. The same properties carry the SDK into robotics and longrunning agentic systems: real-time vision, speech and language models behind one C ABI that links directly into C/C++/Rust control stacks; a single static library with no interpreter or garbage collector, whose allocator-controlled memory keeps latency predictable on embedded compute; and cross-compilation to heterogeneous on-device hardware from one codebase. Because the one library also trains, on-device adaptation — fine-tuning a perception model on locally collected data without a round trip to a training cluster — is supported by the same binary. emu, a freely distributed local-first desktop application built solely on numbat (§3), exercises the deployment model in the field.
5
Domain planes: requirements as executable gates
numbat’s application focus is medical and clinical AI, and in that setting most of the engineering is not the model. It is the domain machinery around it: reading the hospital’s imaging and messaging formats, coding findings to licensed clinical terminologies, and assembling the evidence a regulator will ask for. The stack carries this machinery in planes: standalone packages, each depending on the framework but buildable and testable on its own, covering medical imaging and geometry (including the DICOM wire protocols), genomics and multi-omics (format-parity-tested against the reference C implementations), clinical language and terminology, evidence and evaluation, duplex clinical voice, and a machine-readable surface for AI agents that build against the stack. Together the planes account for most of the SDK’s entry points. What makes the planes relevant to a software-engineering audience is not the domain coverage but the method by which each was built, which we believe generalizes. Requirements are executable. Each plane’s design document enumerates numbered requirements, and each requirement is discharged by an acceptance gate: an executable check with a verdict and an exit code, run from the build system like a test. A development phase is complete when its gate passes, and not before. Gates print the clauses they deliberately do not discharge — requirements awaiting clinical data, elapsed time or absent hardware — so a passing gate cannot be misread as covering them. Across the six planes the suites currently comprise 98 gates and 1,944 individual checks. Failures are typed. Each plane ships a diagnostic registry: stable numbered codes, each carrying a severity class and a typed repair plan, so a caller branches on which requirement was violated rather than parsing a sentence that will eventually be reworded. The registry’s coverage — does
8
every registered code have a real check behind it? — is measured from a ledger written at the raise sites, not from a hand-maintained list. This distinction earned its keep: during development the measured-coverage gate failed three times, each on a code that was registered, documented and reachable from no check at all. A hand-kept list would have reported full coverage at each of those moments. Interfaces refuse. Where a domain rule exists, the API enforces it rather than documenting it. The terminology plane refuses to build an index over a licensed vocabulary unless the caller supplies a license reference, and refuses before reading a single concept row. The evidence plane has no code path that yields a bare metric value: a measurement carries the digest of the run manifest that produced it or is explicitly marked untraceable, and quoting an untraceable measurement is refused — across the C ABI as well as natively, so a consumer in another language cannot route around it. A claim gate can block a release outright when the evidence on file does not match the kind of claim being made, naming in the refusal the study design that would support it. The shipped artifact is part of the test surface. One incident shaped the release procedure. The evidence plane’s ABI reported that it ran fifteen acceptance gates; the suite runs sixteen. The ABI kept its own hand-maintained copy of the gate table, the copy had lost an entry, and the pure-C conformance test had been written against the same copy — so it asserted fifteen, and passed. Every check in the repository was green; the defect was found by installing the released archive, loading the library and asking it how many gates it had. Both tables now derive from one source under a compile-time length check, and every release ends by interrogating the built artifact. The episode independently corroborates the companion study’s finding that the faults that matter often live outside the code paths that testing effort concentrates on [35].
6
Case study: YOLO-NB-M from scratch on COCO
6.1
Task and model
We train YOLO-NB-M3 , an 80-class one-stage anchor-free detector of the YOLOv8m architecture class: CSP backbone, PAN-FPN neck, decoupled head with DFL box regression (reg_max = 16), SiLU activations [6], detection strides {8, 16, 32} (8,400 candidate locations at 6402 ), task-aligned assignment [7], and a binary-cross-entropy (BCE) classification + CIoU [39] + DFL [20] composite loss; 25.9M parameters. Training data is COCO 2017 train2017 (118,287 images); evaluation is val2017 (5,000 images) under the COCO mean-average-precision (mAP) protocol [21]. All weights start from random initialization with seed 0; no pretrained weights of any provenance are used at any point.
6.2
Recipe
Table 1 lists the complete recipe. It reproduces the reference implementation’s from-scratch schedule, including behaviors that are coded rather than configured in the reference trainer — loss scaling by world size, the three-way optimizer parameter grouping, per-step global-norm clipping at 10.0, unconditional per-step EMA with ramped decay, batch-scaled weight decay (λ · b w/nbs with nominal batch nbs = 64), warmup interpolation per iteration, and mosaic shutdown for the final ten epochs. Section 7 documents what happened when any of these was missed. 3
Named by its creators: YOLO for the real-time detector family the architecture belongs to, NB for the numbat framework, M for medium scale. Third-party product names appear in this paper nominatively, to identify the architecture class and the reference implementation used for benchmarking; the model, its weights, and its name are independent work.
9
Table 1: Complete training recipe (identical to the released run manifest). Optimization
SGD, Nesterov momentum 0.937; lr0 = 0.01 with linear decay to 0.01 · lr0 ; 3-epoch per-iteration warmup (momentum 0.8 → 0.937; bias-group LR from 0.1); weight decay 5×10−4 scaled by effective batch (48/64 ⇒ 3.75 × 10−4 ); three parameter groups (decayed weights; undecayed normalization gains; undecayed biases with elevated warmup LR); global-norm clip 10.0 every step; 500 epochs, 2,464 steps/epoch, 1.232 × 106 iterations.
Precision
AMP: F16 compute for conv/matmul; F32 batch normalization, loss, and master weights; dynamic loss scaling (init 216 , floor applied; §7). EMA decay 0.9999 with warmup ramp, applied every step, including BN running statistics.
Augmentation
Mosaic 1.0 (off for final 10 epochs); MixUp 0.1 [38]; copy-paste 0.1 [9]; HSV jitter (0.015, 0.7, 0.4); affine scale 0.9, translate 0.1; horizontal flip 0.5; random erasing 0.4 [40]; letterbox to 6402 .
Parallelism
Data-parallel, one process per GPU, world 3; per-GPU batch 16 (effective 48); bitwisedeterministic sharded all-reduce; rank-invariant epoch shuffling with per-rank augmentation streams (§7); seed 0.
Hardware
3× NVIDIA RTX 3090 (24 GB), power-capped 200 W each (graphics clock ≤1500 MHz), consumer host (6 physical cores); executed as ≈8-hour segments with automatic checkpointresume (≈25 cycles) as a host-reliability mitigation.
6.3
The reproduction protocol
The reference implementation — Ultralytics 8.3.75 on PyTorch 2.12.1 [16, 27] — is treated as an executable specification: its configuration surface, numerical conventions, and observable behavior define the target. This is the differential-oracle arrangement of the companion study [35] with the roles fixed: the incumbent plays the oracle, because the claim under test is that the independent stack reproduces it. No reference source code is incorporated into numbat; parity is achieved by independent implementation, validated at five levels. When any level disagrees, numbat is changed to match the reference, never the reverse. L1 — Operator. Every differentiable operator passes a numeric gradient check on CPU (tolerance 10−4 ) and a CPU-vs-CUDA forward/backward equivalence check. Initializer distributions and the Philox RNG are verified by known-answer tests against the reference (GPU uniform sampling is bit-identical). L2 — Module. Composite components — the detection loss with task-aligned assignment, and batch normalization under autocast — pass known-answer tests against reference outputs on fixed inputs, on both CPU and CUDA. L3 — Step. Fixed-data probes compare optimization kinetics: single- and multi-step loss decrease on identical batches, gradient equivalence between 1-rank and N -rank execution (bitwise, by construction of the deterministic all-reduce), and AMP-vs-F32 step-for-step agreement (4.2% mean absolute loss deviation over 285 steps after the fixes of §7). A kinetics decomposition attributes any trajectory difference to per-step progress vs. step count, which localized several divergences below. L4 — Trajectory. Before committing multi-day compute, the reference implementation itself is run from scratch on the same machine at matched recipe and regime (3-GPU DDP, effective batch 48, official augmentation) for 30 epochs, producing a reference trajectory (Fig. 2a). The production run is then supervised by an automated trajectory gate: a sidecar polls the run every 10 minutes and pauses or terminates it if the validation metric leaves a 20% tolerance band under the reference curve, plateaus, or diverges. Two operational rules emerged and are now part of the 10
Table 2: Ten silent recipe divergences surfaced by the reproduction protocol. Every fix adopted the reference behavior. “Level” = protocol level that detected it (§6.3). #
Component
Divergence → consequence
Level
1
Weight decay
L4
2
AMP autocast
3
Initialization
4
Loss scaler under DDP
5
Mosaic labels
6
Optimizer groups
7
Gradient clipping
8
EMA cadence
9
Evaluation BN
10
Augmentation RNG
Value taken from a released checkpoint’s arguments was already batch-scaled (batch 128); at effective batch 48 this doubled regularization. Fix: re-derive from the base default via λ bw/nbs. BatchNorm executed in the F16 shadow instead of F32 → degraded convergence and elevated overflow-skip rate. Fix: F32 BN and loss, F16 conv/matmul (reference autocast placement). √ Convolutions used Kaiming-normal (gain 2) instead of the refer√ ence’s Kaiming-uniform (a= 5) [11], a 2.45× wider distribution; in the normalization-free detection head this inflated initial classification-logit spread ≈14× and the resulting corrective gradients overflowed F16. With divergence #2/#3 present, per-rank overflow flags OR-reduced across 3 ranks collapsed the dynamic loss scale toward zero (silent learning freeze). Fix: scale floor; standard dynamics restored once #2/#3 were fixed. Labels were not clipped to the mosaic canvas, producing phantom boxes in padding regions. Normalization gains were placed in the elevated-warmup bias group; the reference warms them from zero in their own group. Absent, vs. the reference’s global-norm 10.0 [26] every step (a coded, not configured, behavior). EMA skipped on loss-scaler-rejected steps; the reference updates unconditionally. Validation ran train-mode BN (batch statistics) on an EMA model whose BN buffers were never averaged → noisy, low-biased metric that both hides progress and fakes failure. Fix: eval-mode running statistics, EMA including buffers. All ranks drew identical augmentation streams (effective augmentation diversity ÷3). Fix: rank-folded batch seeds with rank-invariant shuffling. A matched 5-epoch probe measured the single fix at +0.051 mAP over the window vs. +0.005 without it.
L2/L4
L3
L4
L3 L4 L4 L4 L5
L3/L4
protocol: (i) a reference band is valid only for the exact recipe and regime it was recorded under — an earlier band recorded with default (weaker) augmentation produced a spurious “growing lag” verdict against the strong-augmentation run; and (ii) short probes of the reference must pin its optimizer explicitly, since its automatic optimizer selection silently overrides the configured learning rate and momentum. L5 — Outcome. Reported accuracy comes from the official COCO protocol only: numbat weights are exported to safetensors, loaded by the reference stack’s own validator, and scored end-to-end with pycocotools — eliminating any possibility that numbat’s metric implementation flatters the result. numbat’s internal streaming evaluator, used solely for gating, was audited against this pipeline and reads 0.011–0.014 lower than the official score at converged checkpoints.
7
Divergence catalog
An early, un-gated attempt motivated the protocol: it silently plateaued at roughly one fifth of the target metric for 43 epochs (≈1.5 machine-days) before a human noticed — no crash, no NaN, loss decreasing. Post-mortem attributed it to a combination of the divergences below. All were subsequently caught by the protocol, fixed by adopting the reference behavior, and re-validated; Table 2 catalogs them. 11
(a) trajectory-gate window
(b) full 500-epoch schedule 0.5 published reference 0.502
mAP50–95 (val2017)
0.4
0.3
0.2 gate tolerance (20% below ref.) reference (same machine, 30 ep) numbat (internal evaluator)
0.1 0.0
0.4956 @ ep 500
0.4
0.3
0
10
epoch
20
30
0.2 0.1 0.0
numbat (internal, per epoch) official protocol (exported ckpts) 0
100
200
epoch
300
400
500
Figure 2: Convergence of the from-scratch run. (a) The trajectory-gate window: numbat’s per-epoch validation metric tracks the same-machine reference trajectory (recorded with the reference implementation at matched recipe and regime) along the entire 30-epoch window; the automated gate polled 1,403 times over the completed run with zero trajectory violations. (b) The full schedule: internal per-epoch metric (line; reads 0.011–0.014 low, §6.3 L5) and official-protocol scores of exported checkpoints (diamonds), against the published endpoint for the architecture class.
Three observations generalize beyond this case study. First, divergences compose: #3 (initialization) amplified #2 (autocast placement), which triggered #4 (scaler collapse) — three individually plausible implementations combining into a silent training freeze, while the F32 control self-healed within seven steps and masked the chain. Second, distinct root causes share one symptom. Four different divergences presented identically as “slightly below the reference band from epoch 1,” which is why a leveled protocol that can localize (operator? step kinetics? recipe? metric?) terminates debugging that a single end-metric cannot. Third, effective batch dominates short-horizon kinetics: at effective batch 16, both frameworks make near-zero early progress under this recipe; the reference’s gradient accumulation to a nominal batch of 64 is part of the recipe, not a tuning nicety — probes that ignore it mislead.
8
Results
The run is complete: all numbers below are final, measured on the released epoch-500 weights or recorded during the completed 500-epoch schedule.
8.1
Accuracy
Table 3 and Figure 2 summarize accuracy. The final (epoch-500) EMA checkpoint — the released weights — scores 0.4956 mAP50–95 (0.6610 mAP50 ) on COCO val2017 under the official protocol, scored by the reference stack’s validator; the published endpoint for the architecture class at this scale and schedule is 0.502, placing the final score within 1.3% (relative) of it. The best official score observed during training was 0.4970 at epoch 469 (recorded in the released score trend). Notably, the final ten mosaic-free epochs (§6), which typically contribute a final fraction of a point in this recipe family, produced no further gain in this run: the official score moved 0.4970 (469) → 0.4962 (489) → 0.4956 (500), differences at the scale of single-evaluation noise. We report the epoch-500 number as the headline because it corresponds to the released checkpoint; the epoch-469 value is 12
Table 3: COCO val2017 accuracy (official protocol; exported weights scored by the reference validator + pycocotools). The published endpoint is the reference implementation’s reported from-scratch result for the architecture class at this scale.
training loss (per-image)
YOLO-NB-M final, released weights (this work) YOLO-NB-M best during training (this work) Published endpoint (architecture class) Same-machine reference trajectory (gate baseline)
6
mAP50–95
mAP50
epochs
0.4956 0.4970 0.502 0.391
0.6610 0.6620 — —
500/500 469 500 30
box (CIoU regression) cls (BCE) dfl (distribution focal)
5 4 3 2 1 0
100
200
300
epoch
400
500
Figure 3: Training-loss components over the schedule (per-image, smoothed): CIoU box regression (box), binary-cross-entropy classification (cls), and distribution focal loss (dfl).
reproducible from the released trend. Early-trajectory equivalence is direct: over the 30-epoch gate window numbat’s curve lies on the same-machine reference trajectory (Fig. 2a), no lower than 97.9% of the reference value at any comparison epoch and up to 1.24× above it in the earliest epochs, before correcting for the internal evaluator’s known −0.011 to −0.014 offset — i.e., at or above the reference after correction.
8.2
Throughput, memory, energy
Single GPU: parity. At the production configuration on one RTX 3090, phase-synchronized stage timing (Fig. 4a) shows numbat ahead on forward (50.8 vs. 58.2 ms) and optimizer step (2.6 vs. 5.5 ms), at parity on backward (113.0 vs. 110.6 ms) and loss (12.3 vs. 13.4 ms, after replacing a scalar-atomic reduction with a block reduction: 48→12 ms), with data loading fully overlapped by both systems. End-to-end asynchronous step time is 185.9 ms (numbat) vs. 192.8 ms (reference) uncapped, and 245.5 vs. 256.9 ms under the run’s permanent 200 W caps — with numbat additionally carrying live metric streaming (Fig. 4b). We consider single-GPU training performance at parity, established end-to-end at the production configuration rather than on a favorable sub-benchmark. Three-GPU DDP: honest gap. At world size 3 the reference sustains 15.6 min/epoch (154 img/s) vs. numbat’s initial 27.8 min/epoch (87 img/s). Profiling attributes the gap to the serialized hoststaged all-reduce, which also absorbs the full inter-rank arrival skew inside the step, whereas the reference overlaps bucketed NCCL reductions with backpropagation [19]. Five optimizations — pinned staging (host-device copies ≈halved), sharding the reduction across ranks (60–106→19–26 ms), futex barriers replacing spin-waits, the fused gradient-norm kernel (§3), and folding the loss-scaler overflow flag into the gradient payload (one fewer barrier round) — raised throughput to ≈108 img/s 13
(b) end-to-end, 200 W caps
300
reference numbat
250
ms/step
data wait host→device forward loss backward optimizer
256.9
245.5
ref.
numbat
200
images/s
(a) single-GPU stages, bs16 AMP
150 100 50
0
20
40
60
80
0
100
ms/step (phase-synced)
(c) 3-GPU DDP
175 154 150 125 108 100 87 75 50 25 0 ref. nb v0 nb opt.
Figure 4: Same-machine training throughput vs. the reference at the production configuration (batch 16/GPU, AMP, 6402 ). (a) Phase-synchronized single-GPU stage times. (b) End-to-end step time under the run’s permanent 200 W power caps; the numbat measurement includes live metric streaming, the reference runs with auxiliary outputs disabled. (c) Three-GPU data-parallel throughput: reference, numbat before and after the communication optimizations described in the text. Table 4: Environment of record. GPUs Host numbat toolchain Reference stack Dataset Seed
3× NVIDIA GeForce RTX 3090, 24 GB, 200 W cap, ≤1500 MHz Intel Core i5-10600K (6C/12T), Debian 13 Zig 0.17.0-dev (pinned); CUDA 12.9; driver 610.43.02 Ultralytics 8.3.75; PyTorch 2.12.1+cu126; pycocotools COCO 2017 (train 118,287 / val 5,000) [21] 0 (data order, augmentation streams, initialization)
with bitwise-identical gradient math (Fig. 4c). Both remaining levers are implemented but disabled by default on this host class, a negative result we report deliberately: on the 6-physical-core host, both the backward-overlapped reduction and the NCCL transport starve the augmentation workers (data wait 0→60–155 ms/step), a net loss; they await validation on larger-core hosts. Convergence is unaffected throughout: gradient reduction is bitwise deterministic in all configurations. Memory. Peak training memory is 13.5 GiB/GPU after incremental activation freeing (−26.8% from 18.5), vs. ≈7 GiB for the reference at the same configuration; the residual factor ≈1.9× is attributed to AMP shadow parameter copies and allocator high-water behavior, and is an open engineering item. Energy and reliability. Under the 200 W caps, aggregate GPU board power is bounded by 0.6 kW; at the sustained 25–28 min/epoch the 500-epoch schedule completed in 9.6 days of wall clock including cool-down pauses, bounding GPU energy at ≤1.4×102 kWh. The run executed as ≈30 automatic checkpoint-resume segments with full state continuity (parameters, optimizer, EMA, RNG); the trajectory gate (Fig. 2a) polled 1,403 times over the run with zero violations and zero human interventions after launch.
9
Reproducibility and release
We release: the trained weights (a safetensors export of the final EMA checkpoint — the exact artifact scored in §8), per-epoch training and validation metrics, the official-protocol score trend
14
of exported checkpoints, training-curve exports, and a manifest recording every hyperparameter of Table 1 with the environment of Table 4, under a research-use-only (non-commercial) license.4 The claim structure of this paper is independently verifiable without access to numbat: the released safetensors load in the reference stack, and the reported numbers regenerate from its validator and pycocotools. The released weights are an original work: training started from random initialization; no pretrained weights of any provenance were used; and the training system shares no source code with any third-party machine-learning framework. The reference implementation served only as an executable behavioral specification and benchmark during development (§6.3) and as an independent scorer of the exported weights. Verification of every reported number goes through the released artifacts and the reference stack alone, and requires no access to numbat in any form.
10
Limitations
Single seed. The headline run is one seed (0); a variance study at 500 epochs × 3 GPUs was outside our compute budget. The same-machine reference baseline is likewise single-seed, so the equivalence claim is trajectory- and endpoint-level, not a distributional statement. Reference trajectory length. The same-machine reference run covers 30 epochs; beyond it, equivalence rests on the official published endpoint and the official-protocol scoring of our exported checkpoints. Multi-GPU scaling. The DDP gap of §8.2 is real on small-core hosts; the implemented overlap and NCCL paths remain to be validated on larger hosts. Memory. numbat currently uses ≈1.9× the reference’s training memory at this configuration. Scope. The protocol is demonstrated on one detector family; the framework’s LLM/ASR/segmentation stacks run in production but have not yet been taken through the full five-level protocol. Availability. The framework’s source is proprietary, and the terms under which the SDK is distributed are a commercial matter outside this paper’s scope. The paper is written so that neither is needed: every reported score regenerates from the released weights, the reference stack’s validator and pycocotools, and the claim structure stands or falls on those public artifacts.
11
Conclusion
We presented numbat, a self-contained machine-learning stack, and the engineering discipline that holds it to account: reference-default semantics so recipes transfer unchanged, an additive C ABI so applications outlive framework releases, domain requirements encoded as executable acceptance gates, and a five-level reproduction protocol that treats the incumbent implementation as an executable specification. The validation is the strongest test we know for a training stack: reproducing a mature, competitive recipe’s outcome from scratch, with the evidence scored by the incumbent’s own tooling — and, in the companion study, the same two-stack arrangement run in the opposite direction, as a fault-finding instrument [35]. Between them the exercises yield two transferable artifacts, the divergence catalog and the protocol itself, and one existence proof: training-outcome parity with the dominant ecosystem is achievable in a fully independent stack, on consumer hardware. Future work extends the protocol to the framework’s language-model and speech stacks, closes the multi-GPU overlap and memory gaps, and carries the acceptance-gate method into the domain planes still under construction, where we expect it to keep earning its keep the way it has so far: by refusing to let a passing build stand in for a discharged requirement. 4
https://huggingface.co/cloudkites/yolo-nb-m. Commercial licensing: [email protected].
15
List of abbreviations ABI AMP BCE BF16 BLAS BN CIoU CSP CTC DAG DDP DFL EMA F16/F32/F64
application binary interface automatic mixed precision binary cross-entropy bfloat16 floating point Basic Linear Algebra Subprograms batch normalization complete intersection over union cross-stage partial (backbone) connectionist temporal classification directed acyclic graph distributed data parallelism distribution focal loss exponential moving average 16/32/64-bit floating point
FPN IoU JIT LLM mAP nbs NCCL NVRTC OCR PAN RNG SDK SDPA SGD SIMD VLM WGSL DICOM Digital Imaging and Communications in Medicine
feature-pyramid network intersection over union just-in-time (compilation) large language model mean average precision nominal batch size NVIDIA Collective Communications Library NVIDIA runtime compiler optical character recognition path-aggregation network random-number generator software development kit scaled dot-product attention stochastic gradient descent single instruction, multiple data vision–language model WebGPU shading language
References [1] Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. TensorFlow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2016. [2] Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, et al. PyTorch 2: Faster machine learning through dynamic Python bytecode transformation and graph compilation. In 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), pages 929–947, 2024. [3] Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. YOLOv4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:2004.10934, 2020. [4] Xavier Bouthillier, César Laurent, and Pascal Vincent. Unreproducible research is reproducible. In International Conference on Machine Learning (ICML), PMLR 97, pages 725–734, 2019. [5] James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Yash Katariya, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye WandermanMilne, and Qiao Zhang. JAX: Composable transformations of Python+NumPy programs, 2018. https://github.com/jax-ml/jax. [6] Stefan Elfwing, Eiji Uchibe, and Kenji Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural Networks, 107:3–11, 2018. [7] Chengjian Feng, Yujie Zhong, Yu Gao, Matthew R. Scott, and Weilin Huang. TOOD: Taskaligned one-stage object detection. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021. [8] Georgi Gerganov and contributors. llama.cpp: LLM inference in C/C++, 2023. https: //github.com/ggml-org/llama.cpp.
16
[9] Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D. Cubuk, Quoc V. Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. [10] Awni Hannun, Jagrit Digani, Angelos Katharopoulos, and Ronan Collobert. MLX: Efficient and flexible machine learning on Apple silicon, 2023. https://github.com/ml-explore/mlx. [11] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In IEEE International Conference on Computer Vision (ICCV), 2015. [12] Kaiming He, Ross Girshick, and Piotr Dollár. Rethinking ImageNet pre-training. In IEEE/CVF International Conference on Computer Vision (ICCV), 2019. [13] Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. In AAAI Conference on Artificial Intelligence, 2018. [14] George Hotz and the tiny corp. tinygrad, 2023. https://github.com/tinygrad/tinygrad. [15] Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), 2015. [16] Glenn Jocher, Ayush Chaurasia, and Jing Qiu. Ultralytics YOLOv8, 2023. Software, version 8.x. https://github.com/ultralytics/ultralytics. [17] Andrew Lavin and Scott Gray. Fast algorithms for convolutional neural networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. [18] Chuyi Li, Lulu Li, Hongliang Jiang, Kaiheng Weng, Yifei Geng, Liang Li, Zaidan Ke, Qingyuan Li, Meng Cheng, Weiqiang Nie, et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv preprint arXiv:2209.02976, 2022. [19] Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. PyTorch distributed: Experiences on accelerating data parallel training. Proceedings of the VLDB Endowment, 13 (12):3005–3018, 2020. [20] Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In Advances in Neural Information Processing Systems 33, 2020. [21] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: Common objects in context. In European Conference on Computer Vision (ECCV), 2014. [22] Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
17
[23] Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. [24] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. Mixed precision training. In International Conference on Learning Representations (ICLR), 2018. [25] ONNX Runtime developers. ONNX Runtime, 2018. https://onnxruntime.ai. [26] Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International Conference on Machine Learning (ICML), 2013. [27] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, 2019. [28] Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Florence d’Alché Buc, Emily Fox, and Hugo Larochelle. Improving reproducibility in machine learning research (a report from the NeurIPS 2019 reproducibility program). Journal of Machine Learning Research, 22(164):1–20, 2021. [29] Joseph Redmon and Ali Farhadi. YOLOv3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. [30] Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. [31] John K. Salmon, Mark A. Moraes, Ron O. Dror, and David E. Shaw. Parallel random numbers: As easy as 1, 2, 3. In International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2011. [32] Zhiqiang Shen, Zhuang Liu, Jianguo Li, Yu-Gang Jiang, Yurong Chen, and Xiangyang Xue. DSOD: Learning deeply supervised object detectors from scratch. In IEEE International Conference on Computer Vision (ICCV), 2017. [33] Nathaniel Simard, Louis Fortier-Dubois, Dilshod Tadjibaev, Guillaume Lagrange, and contributors. Burn: A next-generation tensor library and deep learning framework, 2024. https://github.com/tracel-ai/burn. [34] Philippe Tillet, H. T. Kung, and David Cox. Triton: An intermediate language and compiler for tiled neural network computations. In 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (MAPL), 2019. [35] Thang Tran and Lan Dang. Cross-stack validation of language-model training: A clinical fine-tuning case study. arXiv preprint arXiv:2608.24267, 2026.
18
[36] Chien-Yao Wang, Hong-Yuan Mark Liao, Yueh-Hua Wu, Ping-Yang Chen, Jun-Wei Hsieh, and I-Hau Yeh. CSPNet: A new backbone that can enhance learning capability of CNN. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020. [37] Chien-Yao Wang, Alexey Bochkovskiy, and Hong-Yuan Mark Liao. YOLOv7: Trainable bag-offreebies sets new state-of-the-art for real-time object detectors. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. [38] Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations (ICLR), 2018. [39] Zhaohui Zheng, Ping Wang, Wei Liu, Jinze Li, Rongguang Ye, and Dongwei Ren. Distance-IoU loss: Faster and better learning for bounding box regression. In AAAI Conference on Artificial Intelligence, 2020. [40] Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In AAAI Conference on Artificial Intelligence, 2020.
19