Conceptio › Archive › arXiv CS
arXiv CSopen access

Nested Parallel von Neumann Architecture and Nested BSP

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

N ESTED PARALLEL VON N EUMANN A RCHITECTURE AND N ESTED BSP AR X IV P REPRINT

arXiv:2609.16787v1 [cs.DC] 15 Sep 2026

Liao Heng Huawei Technologies Co., Ltd.

A BSTRACT Large-scale AI computing is no longer a contest of “one stronger processor,” but of how an army of processors under one command can still be one computer. This paper offers two interlocking extensions. First, extend BSP to Nested BSP. The Turing machine describes computation as a single tape, in sequence [1]. A million processors need not a longer tape, but a battle plan nested layer within layer: at every layer, parallel work, barrier, exchange and aggregate, then the next phase. Every “parallel advance” inside a layer repeats the same four steps. Nested BSP extends classic BSP by nesting it recursively, a computing paradigm for million-scale parallelism, under one rule: every node at every layer is a peer. Second, extend von Neumann to the Nested Parallel von Neumann Architecture, and Unified Bus is its interconnect. Von Neumann taught us how to build one stored-program computer [2]. The false extrapolation of eighty years was that wiring many computers into a network yields one larger computer. A second habit ran deeper: nearly every design assumes a master that commands and slaves that obey—host over device, CPU over accelerator, center over edge. The Nested Parallel Architecture extends that idea rather than discarding it. Two nesting dolls must fit: Nested BSP in software, and the Nested Parallel von Neumann Architecture from package to autonomous zone, joined by one memory-semantic bus end to end, with full peer equality: physically sparse, logically tight. It pairs with Huawei’s τ Scaling law [6]: τ governs how each layer folds time, while the Architecture governs how the nested parallel computer stands, layer by layer, peer by peer. In summary, the paper extends BSP to Nested BSP and extends von Neumann to the Nested Parallel von Neumann Architecture. τ folds time, peer-equal parallelism nests layer by layer—many processors, still one computer.

1

Introduction: The Lesson von Neumann Never Taught

The question this paper asks is a single one: when processors number a hundred thousand, or a million, do we still know how to design? How to build a computer, von Neumann made clear [2]. For eighty years we have followed that path, making processors stronger generation after generation. But how to turn so many processors into one army, that he never said. He was talking about one machine, not a campaign of a million. Chip contests used to look like two martial artists sparring: the harder punch wins. An AI campaign is not fought that way. A million soldiers take the field. What matters is not who hits hardest, but whether the army holds formation. At a hundred thousand or a million, victory turns on one thing: how to bind them into an army, and once bound, whether it is still one computer. Anyone who leads troops knows: an army holds only with a battle plan and a chain of command that can reach the ranks. Without a clear plan and without orders that get through, numbers alone are loose sand. And when a million must move at once, those orders are hard to deliver.

Nested Parallel von Neumann Architecture and Nested BSP

2

AR X IV P REPRINT

Nested BSP: A Battle Plan for a Million

The HPC community hit this wall long ago. Leslie Valiant and Bill McColl proposed BSP and cut a campaign into phases [3, 4]. In each phase, units advance in parallel, and no one waits on anyone else. When they reach their marks, the whole army barriers on one line. After the barrier, they exchange what must be exchanged and aggregate what must be aggregated. Once the handoff is clean, they enter the next phase together and advance in parallel again. A problem of that scale becomes tractable. We now formally introduce Nested BSP, which extends classic BSP by nesting it recursively. The idea is simple. The outermost layer is BSP for the whole army. And every “advance in parallel” inside that layer is itself a smaller army; within its own scope it runs the same four steps again: parallel work, barrier, exchange and aggregate, next phase. Inside that sits a still smaller layer, and within that, yet another. Layer within layer, recursion all the way down. What Nested BSP extends beyond classic BSP is one hard constraint: at every layer, the units are peers—no master that must own every barrier, no slave that may only obey. Parallelism is nested, authority is not. Huawei’s SuperNode cluster design rests on Nested BSP as its theoretical foundation. The hardware properties described so far—one protocol end to end, memory semantics, full peer equality, copper near and optics far—all exist so that this nested structure can stand in the physical world. That hardware nest is the Nested Parallel von Neumann Architecture: from package to autonomous zone, concurrent units at every layer, peers on one bus. First, the software nest, illustrated in Fig. 1.

DP

DP

PP

EP PP

PP

FSDP

CP EP

EP

EP

EP

TP

Figure 1: Software nesting dolls: Nested BSP. From outermost to innermost, the levels are data parallel (DP), pipeline parallel (PP), expert parallel (EP), fully sharded data parallel (FSDP), context parallel (CP), and tensor parallel (TP). Every layer runs the same playbook: compute in parallel, then barrier and reduce. In plain terms: map and reduce. A problem of that scale is broken into a problem every layer can manage. Orders pass down layer by layer; only then does the army hold. Notice one thing: inside each larger doll there is not one smaller doll, but many. One DP holds many PPs; one PP holds many EPs; one EP holds many FSDP shards. That multiplicity is the source of parallelism. 2

Nested Parallel von Neumann Architecture and Nested BSP

AR X IV P REPRINT

The body of this army is also a nest of dolls, illustrated in Fig. 2, layer within layer, each with its own boundary. Counting from the inside out: CORE, chiplet, chip package, board, rack, SuperNode, data hall, autonomous zone, data center, and outermost, inter data center.

autonomous zone

autonomous zone

data hall

SuperNode data hall

data hall

rack

board SuperNode

SuperNode

SuperNode

SuperNode

chip package

Figure 2: Hardware nesting dolls: the body of the machine, from chip package to autonomous zone. One autonomous zone holds many data halls, one hall holds many SuperNodes, and one SuperNode holds many racks, all the way down to the package.

The software nest above and this hardware nest must line up, layer by layer. This paper does not go from the innermost atom to the outermost campus. We focus on the middle stretch of hardware system design: from the chip package to the cluster that fills a data center. Further in is how to train one soldier; further out is how campuses talk to campuses. The battlefield where a hundred thousand or a million become one army sits in these middle layers: package, board, rack, SuperNode, data hall, autonomous zone. Orders must pass down without a single layer dropped. Between layers: is it still one plan, one command set, one computer? Unified Bus does exactly that: it makes the split and join of every layer line up, so that command does not break at the boundaries [5]. On the τ Scaling law [6], we note this: at the system level, it maps to these two nests of dolls. At every Nested BSP layer and every physical scale, Huawei’s τ Scaling law does the same thing: it folds time. The playbook at every layer is the same: work that would have been done one piece after another is spread across many units that can work at once on that layer, and that layer’s time constant folds shorter. Once inside the package, once on the board, once in the rack, once in the SuperNode, once again in the data hall. The software nest is Nested BSP, and the hardware nest is the Nested Parallel von Neumann Architecture, with more concurrent, peer-equal units at every layer: more in the package, more on the board, more in the rack, more in the hall. The two nests correspond layer by layer, and time folds at each. One layer alone folds only so much, six layers fold multiplicatively. That is how the time of the whole campaign truly comes down. 3

Nested Parallel von Neumann Architecture and Nested BSP

AR X IV P REPRINT

So the two nests must fit with no gap. Where a layer fails to match, the step scatters: in the protocol, in the software, in the network between racks. No matter how many parallel units are placed, they cannot deliver. This nest follows the same recursive, fractal logic as the physical world: farther out, larger scale, longer distance, lower bandwidth, higher latency; farther in, smaller scale, shorter distance, higher bandwidth, lower latency. We are continuing von Neumann’s idea: building this hierarchically nested parallel computing system. The principle does not change. What changes is how large the computer is, how many layers it nests, and how far a single command can reach. Distant memory must feel as natural as memory at hand. Chips are soldiers, and the computer is the whole army—from package to autonomous zone—that answers to one command set.

3

Six Things for Unified Bus Hardware: Turn the Cliff into a Slope

The principle is clear. We now turn to the hardware that realizes it, and to where past technology fell short. Begin by setting the target. Farther out means longer distance, lower bandwidth, higher latency: that is physics, and it is accepted. The real past failure was not a “gentle slope”—it was a “cliff.” One step past the chassis, bandwidth dropped by an order of magnitude and latency rose several-fold. Inside the rack it behaved like one machine; outside the rack, everyone went their own way. The job of the critical technology is to pave that cliff into a slope, so that command passes down without a single layer dropped. First: one protocol, the same inside the rack and outside it. The past used two. Inside the box, a bus was fast and nanosecond-class, but born for short reach and unable to leave the chassis. Outside the box, a network could reach far, but every boundary was a toll booth—unpack, inspect, re-pack. One language was used inside the rack and another between racks, every boundary required a transfer. In large clusters, more than 80% of energy goes to moving data, and a large share of that is lost at these transfers. Unified Bus merges “bus” and “unified” into a single protocol from the package all the way to the autonomous zone, end to end, with no transfers. Second: make distant memory feel like memory at hand. Moving data used to mean climbing layer by layer to the application, then converting across protocols. A TCP/IP round trip cost tens of microseconds, mostly in software. RDMA improved that sharply, but at heart it was still “send a message, wait for a reply.” Barriers and reductions all stuck here, and the larger the scale, the worse the stick. We switched to memory semantics: native load and store, with consistency handled in hardware. One communication round trip fell from tens of microseconds to around a hundred nanoseconds, approximately five hundred times; per-chip I/O bandwidth reaches the 7.2 Tbps class. Third: tear down master–slave, the foundation of peer equality. In the past, the host commanded and the device obeyed, namely everything went through a center. That master–slave pattern is not an accident of one product line; it is the default pattern of traditional computer design: the CPU as master with accelerators as slaves, the host as master with devices as slaves, and a control plane that owns initiation while everyone else waits. The larger the army, the more the center clogs. Adding soldiers does not add strength—one plus one is less than two. This design puts CPU, NPU, memory, storage, and NICs on the same bus as peers: anyone can initiate, anyone can respond. Barrier and aggregation at every Nested BSP layer need not report back to a center for every act. Without that equality, the nested battle plan collapses back into a queue at the master; with it, the Nested Parallel von Neumann Architecture can scale without a single throat. Fourth: copper near, optics far, but the language does not change. Copper has hit its wall. Push up the data rate per wire, and the wire gets thicker and its reach get shorter; bundle thousands of copper cables and they become too thick to install. So first-generation Unified Bus was designed for both electrical and optical: copper nearby to protect latency, optics farther out to protect scale, with the protocol unchanged when the medium changes. On the optical side we chose near-package optics (NPO). Of more than twenty decibels of loss on a full electrical path, NPO removes nine to eleven in one step; co-packaged optics (CPO) can take only about three more. What is gained: optical engines that can be built, tested, and replaced as single modules. Latency at the ten-nanosecond class and overall cost more than 40% below CPO. This path is now an OIF project, pursued by dozens of partners [7]. Fifth: do not cram the soldiers into one tent. The old instinct was to pack to the limit: larger chips, fuller racks. Compute grows with area, but I/O bandwidth and power grow only with perimeter: the mismatch worsens as one packs. Heat is more concrete: a megawatt rack may 4

Nested Parallel von Neumann Architecture and Nested BSP

AR X IV P REPRINT

need hundreds of square meters just for cooling; a gigawatt hall may hold only a thousand racks yet occupy a square kilometer, with water pipes stretched a kilometer. Floor space in the hall is cheap; chips in the rack are expensive. So we do not chase megawatt racks. We let the army stand open, then bind it with optics into one machine: stand sparsely, compute tightly, physically sparse, logically tight. Sixth: in each generation, fight only a few hard battles. Pick twenty physical limits at once, each at 90% odds, and the joint success rate is near zero. Pick five, and one can still ship. Choosing what to take on and what to defer to a later form is itself part of design.

4

Down to Practice: Three Things That Must Be Done Right

Principle and path are set. What must be got right in the current build comes down to three things. First: switches must be high radix. Everyone agrees on this. The more exits at a junction, the fewer handoffs across the army. For the same scale, higher radix means fewer intermediate layers, and the latency and power saved are real. Second: compute chips must also be high radix. This one is often overlooked. The old view was that the chip should just compute; one or two outbound links are enough, leaving the rest to the network. In an army of a million, that is like forcing a whole battalion’s dispatches through a single door. The chip itself must be a junction. More exits buy three gains: First, higher total outbound bandwidth. Second, more paths: if one breaks, another remains, and in ten-thousand-GPU clusters, link and module failure is the norm, not the surprise. Third, and most important, one-hop coverage equals switch ports times chip ports. Both sides are multipliers. Double the chip’s radix and one doubles the one-hop SuperNode scale. The larger that one-hop reach, the shorter the path for barrier and aggregation: exactly the innermost, most frequent Nested BSP layers fall inside that reach. Raising the product of these two radices sets the parallelism at every nested hardware layer. Third: use high-density NPO to bring light to the very edge of the compute chip. The closer electro-optical conversion sits to the chip edge, the shorter the copper segment that remains. What this buys is a result that sounds contradictory but is natural: physically sparse, logically tight—stand sparsely, compute tightly. Logically tight means one-hop reach, memory semantics, and that the whole SuperNode remains one computer. Physically sparse means chips need not crowd into one iron box; they can spread out. Across this six-layer nest, physical space can open and relax freely; nothing needs to be crushed into a tiny volume. Look at the scale, illustrated in Fig. 3: from inside to outside, five orders of magnitude. The critical ring is the innermost. In the past, the electro-optical boundary sat around one meter: optics only after leaving the chassis. That meant the whole stretch from package through board to rack was carried by copper alone. Copper is unforgiving: push the rate up, and the wire grows thicker and reaches shorter; the board fights insertion loss, the rack fights density, and every step meets a physical limit. NPO moves that boundary inward by two orders of magnitude, to ten millimeters at the package edge. Copper travels only the last few centimeters; everything beyond is handed to light. Once the boundary moves inward, every outer layer relaxes. The board no longer fights insertion loss for dozens of copper cables, and the rack need not pack full or pile to a megawatt. Racks can simply pull apart, since light does not care about that distance. Each layer outward can expand tenfold in scale, and components can sit more sparsely. And through that entire expansion, logically it remains one computer. So we need not challenge the limit of spatial density. Cooling, power, and reliability: these three hurdles no longer have to advance against physical limits at the same time. “Stand sparsely, compute tightly” lands on this one move. 5

Nested Parallel von Neumann Architecture and Nested BSP

AR X IV P REPRINT

Past:~boundary~at~1~m Optics~only~after~leaving~the~rack Board~and~rack~carried~by~copper~alone

SuperNode

data~hall

~10~m

~100~m

rack

data~center ~1~km

~1~m Boundary~moves~inward by~two~orders~of~magnitude

board ~100~mm

Now:~boundary~at~10~mm NPO~converts~at~the~package~edge Copper~travels~only~the~last~few~cm

Same~computer,~scale~expanded~~10⁵×

chip~package ~10~mm

10~mm

100~mm

1~m

10~m

100~m

1~km

Physical~distance~(log~scale)

Figure 3: Design freedom from NPO: package at ten millimeters, board at ten centimeters, rack at one meter, SuperNode at ten meters, data hall at a hundred meters, and data center at a kilometer. Moving the electro-optical boundary inward from 1 m to 10 mm relaxes every outer layer—physically sparse, logically tight.

5

Conclusion: The Nested Parallel von Neumann Architecture—Many Processors, Still One Computer

Extend BSP to Nested BSP. The Turing machine is a single-tape, sequential picture of computing. A campaign of a million processors needs the same four steps, nested layer within layer: parallel work, barrier, exchange and aggregate, next phase. Nested BSP extends classic BSP layer within layer under peer equality. At every layer, Huawei’s τ Scaling law folds time: one layer alone folds only so much, but six layers multiply, and the time of the whole campaign truly comes down. Extend von Neumann to the Nested Parallel von Neumann Architecture, and Unified Bus is the interconnect that lands it. Von Neumann defined “one computer”; the Nested Parallel von Neumann Architecture extends that definition to how an army of processors remains one computer: software and hardware nesting dolls aligned, memory semantics end to end, every node a peer on the bus, physically sparse, logically tight. It pairs with the τ Scaling law: τ is the law of time folding; the Nested Parallel von Neumann Architecture is the extended, peer-equal architecture that sets master–slave aside. Put together, a SuperNode can stand more than eight thousand nodes, with aggregate memory bandwidth of 6.7 PB/s, full interconnect of 400 Tbps, and a full-army barrier under ten microseconds. For users, this means that adding chips adds compute, chip strength is no longer spent on moving data, and the cost per token comes down. Follow this nested structure outward and scale can keep growing. A single Super AI computer system is currently being deployed at the 256K-node class. It is one machine, not a network of two hundred thousand machines glued together but one computer with one Nested BSP battle plan and one Nested BSP / Unified Bus command set through to the end. The Unified Bus Protocol specification is made openly available [8]. The standard is designed to be larger than any single company: whoever connects to it obtains the same latency and the same linearity. In summary, this paper extends BSP to Nested BSP and extends the von Neumann single-machine architecture to the Nested Parallel von Neumann Architecture, connected end to end by Unified Bus. The result is that many processors remain one computer, and every node is a peer. 6

Nested Parallel von Neumann Architecture and Nested BSP

AR X IV P REPRINT

References [1] A. M. Turing, “On computable numbers, with an application to the Entscheidungsproblem,” Proceedings of the London Mathematical Society, vol. s2-42, no. 1, pp. 230–265, 1936. [2] J. von Neumann, “First draft of a report on the EDVAC,” University of Pennsylvania, Philadelphia, PA, USA, Tech. Rep., 1945. [3] L. G. Valiant, “A bridging model for parallel computation,” Communications of the ACM, vol. 33, no. 8, pp. 103–111, 1990. [4] W. F. McColl, “Scalable computing,” in Computer Science Today, ser. Lecture Notes in Computer Science. Springer, vol. 1000, pp. 46–61, 1995. [5] H. Liao et al., “UB-Mesh: A hierarchically localized nD-fullmesh data center network architecture,” IEEE Micro, vol. 45, no. 5, pp. 20–29, 2025. [6] T. He, “A time scaling theory for multi-layer electronic systems,” Science China Information Sciences, vol. 69, Art. no. 183401, 2026. [7] Optical Internetworking Forum (OIF), “OIF 2026Q2 Meeting Press Releases,” [Online]. Available: https: //www.oiforum.com. [8] Unified Bus, “The Unified Bus official website.” [Online]. Available: https://www.unifiedbus.com/en

7

Record · ID 919331 · SHA-256 2fe5d6c9805510be
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.