PREPRINT
1
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
arXiv:2605.25655v1 [cs.DC] 25 May 2026
Yao Lu, Zhongzhi Luan, Senior Member, IEEE, Gen Li, Jiaxing Qi, Shiqing Ma, Bin Han, Shizhe Shang, Hailong Yang, Member, IEEE, and Depei Qian, Senior Member, IEEE
Abstract—Large language model (LLM) inference is limited by high computational cost and memory bandwidth demands, making deployment on heterogeneous many-core processors challenging. Taking the MT-3000 processor used in the Tianhe supercomputer as an example, its limited main memory bandwidth and distributed memory hierarchy exemplify these bottlenecks, making it difficult to directly migrate existing GPU-based inference frameworks. To address this, we propose THInfer, a hardwareaware inference framework that maximizes data locality under bandwidth-constrained conditions through hardware-software co-design and parallel strategy optimization. The framework incorporates three key technologies: (1) a high-performance operator library for the VLIW-SIMD architecture, providing hand-optimized FP16 kernels that achieve up to 70% of peak performance per cluster; (2) a density-driven computation graph fusion and unified kernel scheduling mechanism, combined with a staged pipelined attention fusion method for co-design; (3) a Prefill–Buffer–Decode (P–B–D) pipeline and bounded buffer management strategy, which supports hybrid parallelism while enabling efficient multi-cluster collaboration and scaling through two-level communication integrating MPI and hthreads. Experiments on the Llama model series show that THInfer improves throughput on the 7B model by 62%–73% over DeepSpeed on 2×V100S, and by 67%–84% over the A800 GPU. The 13B and 30B models also demonstrate comparable or even better performance. Moreover, THInfer maintains stable performance on the 70B model, whereas typical GPU-based frameworks fail. Overall, THInfer significantly enhances throughput, reduces latency, and improves scalability, providing a feasible technical pathway and system solution for efficient and scalable LLM inference on heterogeneous many-core architectures. Index Terms—LLM Inference, Tianhe New-Generation Supercomputer, MT-3000, Parallel processing, Distributed systems.
I. I NTRODUCTION In recent years, large language models (LLMs) such as Llama [1], Qwen [2], and DeepSeek [3] have achieved remarkable breakthroughs in natural language processing and multimodal reasoning, owing to their powerful semantic modeling capabilities. As model sizes surge to hundreds of billions of parameters, the computational and memory demands during inference have grown exponentially, imposing extremely high requirements on the computing platform’s arithmetic capability, memory bandwidth, and energy efficiency. Existing inference frameworks (e.g., vLLM [4], TensorRT-LLM [5], and DeepSpeed [6]) improve efficiency through dynamic Yao Lu, Zhongzhi Luan, Gen Li, Jiaxing Qi, Shiqing Ma, Bin Han, Shizhe Shang, Hailong Yang, and Depei Qian are with the Sino-German Joint Software Institute, Beihang University, Beijing 100191, China. E-mail: {luyuan, luan.zhongzhi}@buaa.edu.cn. (Corresponding author: Zhongzhi Luan.)
GPU cache (SRAM)
≈20MB
GPU VRAM (HBM)
≈32GB
Main Memory (DDR)
≈200GB
Disk
≈TB
MT-3000 cache (SRAM)
72MB
MT-3000 GSM (HBM)
24MB
Main Memory (DDR)
≈80GB
Disk
≈TB
≈TB/s
≈TB/s
≈30GB/s
≈120GB/s
≈2GB/s
≈2GB/s
Fig. 1. GPU vs. MT-3000 Memory Hierarchy
batching, memory optimization, and quantization. However, their designs heavily rely on high-bandwidth unified memory architectures (such as the 900 GB/s memory bandwidth of NVIDIA V100), making them difficult to adapt to specific high-performance computing systems. Therefore, this paper targets the Tianhe New-Generation supercomputers, aiming to develop an efficient and flexible inference framework for LLMs. The system employs independently developed MT3000 multi-core digital signal processor (DSP), featuring a heterogeneous many-core architecture with “16 general-purpose CPUs + 4 acceleration clusters”, leveraging instruction-level parallelism and software-controlled memory to achieve highefficiency, low-power computation. Although this platform offers abundant heterogeneous computing resources, current applications are largely confined to traditional scientific computing domains such as molecular dynamics [7]. Due to differences in architecture and programming models, its resource utilization remains low in deep learning applications, resulting in issues such as load imbalance and hardware idling. The massive parallel computing capability of this processor provides a hardware foundation for distributed inference of models with tens of billions of parameters. Meanwhile, its unique hardware architecture presents three core challenges for developing an efficient LLM inference framework: • Unique Hardware Characteristics. The unique heterogeneous many-core architecture of the MT-3000 leads to a mismatch with traditional computing paradigms. It necessitates fine-grained task partitioning for computeintensive operators (e.g., GEMM) and memory-intensive operators (e.g., Softmax) across and within accelerator clusters to achieve load balancing and instruction pipeline optimization. • Limited Memory Bandwidth. The actual main memory bandwidth per processor is only about 120GB/s, which is significantly lower than the memory bandwidth of GPUs with comparable computing power (a comparison of their memory hierarchies is shown in Fig. 1). Additionally, the
PREPRINT
limited on-chip storage capacity makes improving data reuse to alleviate bandwidth constraints a key aspect of performance optimization. • Scalability bottlenecks. The orders-of-magnitude gap between the roughly 20 GB of main memory available per cluster and the storage required by models with tens of billions of parameters causes the parallelization strategies of existing distributed inference frameworks to suffer scalability limits due to cross-cluster communication overhead. To address these challenges, we propose THInfer––an LLM inference framework for the Tianhe system. Its core innovation lies in establishing a deep collaborative optimization mechanism between hardware characteristics and model architecture. The framework adopts a co-design approach across three levels: operators, algorithms, and system architecture, addressing the challenges as follows: At the operator level, by adapting to hardware features, hand-written assembly kernels based on dataflow analysis are developed for high-frequency operators such as linear and attention layers, increasing singlecluster computational efficiency to 70% of the theoretical peak. FP16 optimization is employed to reduce memory usage, thereby mitigating hardware mismatch. At the algorithmic level, by abstracting the model’s computational graph, a “lowdensity–high-density–low-density” operator fusion strategy is adopted, embedding low-density operators like RoPE into high-density kernels. Furthermore, a phased MT Attention mechanism is proposed to align with the storage and broadcasting characteristics of AM/SM/GSM, reducing I/O access pressure. Unified kernels are introduced to minimize launch overhead, effectively addressing limited memory bandwidth and insufficient computational efficiency. At the system architecture level, a three-stage synchronized pipeline based on Prefill–Decode separated scheduling, namely P–B–D (Prefill–Buffer–Decode), is constructed. This integrates a selective batching mechanism and combines KV cache asynchronous migration with computation overlap to prevent request accumulation and out-of-memory (OOM) errors. A constraintbased hybrid parallel strategy is employed during both Prefill and Decode stages: intra-cluster synchronization is achieved via hthread, while inter-cluster communication relies on MPI, forming a two-level parallel mode that collaboratively enables efficient aggregation of cross-device memory bandwidth, thereby overcoming scalability bottlenecks. Experimental results demonstrate that the framework achieves near-linear throughput scaling on commonly used LLMs such as Llama, while significantly reducing end-to-end inference latency. Under peak performance-aligned experimental settings, the end-to-end throughput of this framework is overall comparable to or better than the V100/A800 baselines based on DeepSpeed, providing an efficient and feasible solution for industrial-scale LLM inference on the Tianhe New-Generation supercomputers. The main contributions of this paper can be summarized as follows: •
High-Performance Operator Library for MT-3000. Developing dataflow analysis-based high-performance
2
operators and FP16 kernels tailored for the MT-3000 architecture, co-optimizing computational efficiency and memory usage; • Density-Aware Deep Computation Graph Fusion. Proposing a density-driven fusion strategy to embed lowdensity operators into high-density kernels, and designing an MT Attention mechanism adapted to memory hierarchy for reduced I/O and scheduling overhead; • Adaptive Parallel Scheduling Mechanism. Constructing a P–B–D three-stage pipeline at micro-batch granularity to safely decouple prefill and decode phases. Combined with hybrid parallelism and communication optimization, it supports efficient scaling for models with tens of billions of parameters; • Systematic Evaluation Framework. Establishing a multi-dimensional performance evaluation system that validates the superiority and robustness of the proposed approach in terms of latency, throughput, and resource utilization on the Llama series of models.
II. BACKGROUND AND MOTIVATION A. Tianhe New-Generation Supercomputer The Tianhe new-generation supercomputer is built on the independently developed MT-3000 DSP, employing an innovative CPU-DSP hybrid architecture for hardware design [8]. This processor achieves an energy efficiency ratio of 45.4 GFLOPS/W in double-precision mode, representing a 62% improvement compared to the NVIDIA V100 [9]. Its prototype system delivers a peak performance of 12 Petaflops with a Linpack efficiency of 80%, demonstrating significant largescale scalability. The MT-3000 processor is divided into a general-purpose region and an accelerator region: the general-purpose area includes 16 CPUs equipped with two levels of private caches; the acceleration area consists of 4 autonomous clusters, each containing 24 control cores and 384 accelerator cores, with communication achieved through a hierarchical interconnection network, as shown in Fig. 2a. Each DSP core is composed of 1 control core and 16 accelerator cores, supporting 1024bit SIMD instructions and VLIW architecture, integrating a Scalar Processing Unit (SPU), Vector Processing Unit (VPU), and equipped with dedicated storage resources (64KB scalar memory (SM) and 768KB vector memory (AM)). The structure of a single DSP is shown in Fig. 2b. The MT-3000 adopts a three-tier memory hierarchy: onchip AM/SM cache directly interfaces with registers, the 6MB GSM within the cluster supports cross-core data sharing, while off-chip DDR memory enables efficient data transfer via DMA, with bandwidth ranked as AM/SM ¿ GSM ¿ DDR. DMA supports asynchronous transfer and broadcast mechanisms. The communication system adopts a hybrid programming model: within clusters, dynamic multi-core task scheduling is achieved through the hthreads heterogeneous threading library, and inter-cluster collaboration is based on the MPI protocol [10] for data synchronization.
PREPRINT
3
Cluster 0
Cluster 1
General Zone
Acceleration Zone 32 GB D D R
D S P 0
D S P 1
D S P 23
GSM
CPU
CPU
CPU
L1
L1
L1
L1
L2
L2
L2
L2
CPU
CPU
CPU
CPU
L1
L1
L1
L1
L2
L2
L2
L2
CPU
CPU
CPU
CPU
L1
L1
L1
L1
Cluster 2 Acceleration Zone 32 GB D D R
D S P 0
D S P 1
D S P 23
GSM
DSP
Acceleration Zone
CPU
D S P 0
D S P 1
D S P 23
32 GB D D R
GSM Cluster 3
L2
L2
L2
L2
CPU
CPU
CPU
CPU
L1
L1
L1
L1
L2
L2
L2
L2
Control Core
SPU
VPU
VMAC1 VMAC2
SPE
VMAC3
SM 64KB
SVR Broadcast
V P E 0
V P E 1
V P E 15
Acceleration Zone D S P 0
D S P 1
D S P 23
32 GB D D R
AM 768KB
GSM
(a) Overview of MT-3000 Architecture
(b) DSP Internal Structure
Fig. 2. Illustration of the MT-3000 system.
B. Classical LLMs Since the introduction of the Transformer architecture [11], its powerful contextual modeling and cross-task generalization capabilities have driven a paradigm shift in artificial intelligence. In natural language processing (NLP), BERT [12], based on bidirectional masked language modeling, broke through semantic representation bottlenecks with 340 million parameters. The GPT series started with GPT-2’s [13] 1.5 billion parameters, demonstrating that large-scale autoregressive models could significantly improve text generation quality. GPT-3 [14] reached 175 billion parameters, exhibiting emergent abilities and few-shot learning for the first time. The Llama [1], [15], [16] series introduced Grouped Query Attention (GQA), achieving leading multi-task performance within a parameter range of 7 billion to 405 billion. DeepSeek-V3 [17] employs a Mixture of Experts (MoE) architecture, dynamically activating 37 billion parameters with an effective capacity of 671 billion, making it suitable for complex reasoning tasks such as mathematics and programming. The Qwen [2], [18] series supports multimodal tasks through cross-modal alignment, including visual question answering and document analysis. Research shows a significant positive correlation between model scale and performance—scaling parameters from hundreds of millions to hundreds of billions not only enhances generation quality but also catalyzes emergent capabilities. C. LLM Inference Optimization Techniques To enhance the inference efficiency of LLMs, various optimization techniques have been proposed. In model compression and quantization, methods such as GPTQ [19] and AWQ [20] achieve aggressive 3–4 bit quantization, reducing resource requirements; SparseGPT [21] combines sparsification and quantization to simultaneously decrease computational and memory overhead. Pruning and distillation techniques (e.g., LLM-Pruner [22], MiniLLM [23]) compress model scale through structured pruning and knowledge distillation. For attention mechanism optimization, Flash Attention [24] reduces computational complexity via tiling and I/O optimization; models like Longformer [25], Linformer [26], and Reformer [27] leverage sparsification, low-rank approximation, or hashing techniques to lower complexity to linear or near-linear. To address the KV cache memory bottleneck, PagedAttention [4] improves throughput via paged management; GQA [28] reduces the number of key-value
heads; SCISSORHANDS [29] prunes redundant tokens; and StreamingLLM [30] compresses memory using a combination of window attention and sink token mechanisms. At the system level, scheduling strategies such as ORCA [31] enhance throughput and concurrency through iteration-level scheduling and selective batching. Together, these techniques advance LLMs toward greater efficiency and practicality, laying the foundation for edge computing and real-time applications. D. LLM Inference Frameworks Transformer inference faces significant computational and memory bottlenecks, for which various frameworks have proposed efficient solutions. vLLM mitigates memory fragmentation through PagedAttention, improving throughput by 2–4× [4]; TensorRT-LLM leverages TensorRT for deep optimization, supporting dynamic batching and quantization for low-latency inference on NVIDIA GPUs [5]; TGI employs continuous batching and tensor parallelism as a high-performance inference backbone in the Hugging Face ecosystem [32]; DeepSpeed Inference optimizes multi-GPU and heterogeneous computing to enable real-time inference for trillion-parameter models [6]; FlexGen enhances heterogeneous resource scheduling with 4-bit quantization for high-batch offline tasks [33]; and DistServe decouples Prefill and Decode stages through bandwidth-aware placement and pipeline scheduling, combining continuous batching with cross-request KV reuse to improve goodput under TTFT/TPOT constraints [34]. These advancements collectively address key challenges in scalable and efficient LLM deployment. III. M ETHODS A. THInfer Framework The THInfer framework adopts a systematic co-optimization approach, constructing a multi-level LLM inference system with deep hardware integration. Its overall architecture is illustrated in Fig. 3. Operator-Level Optimization: As the foundation of system performance, this layer focuses on unlocking the extreme performance of key operators (e.g., GEMM). Through handwritten assembly kernels, fine-grained dataflow design, and FP16 precision support, it fully leverages the computational
PREPRINT
4
Graph-level Algorithmic
High-performance Operator Library
P-B-D three-level synchronous pipeline
Density-driven Fusion Strategy
Hardware-aware Dataflow
Selective Batching
MT attention
Hand-written Assembly Kernels
Constraint-based Hybrid Parallelism
Unified Kernel Strategy
FP16 Operators Support
Hardware Abstraction Off-chip Bandwidth
Adaptive Parallel Scheduling
Off-chip Memory
THInfer
LLM
On-chip Memory Register File Scalar Memory
Array Memory On-chip Bandwidth
Global Shared Memory
Fig. 3. THInfer: Co-Designing Operators, Graph-Level Algorithms, and System-Level Adaptive Parallelism for MT-3000 LLM Inference
FP32peak = 4.05 TFLOPS), the following analysis can be performed. For FP32, the computational intensity is
5
Input
RMSNorm
Linear
RoPE
Linear
RoPE
1
Matmul
Scale
Add
SoftMax
3
Matmul
Linear
RMSNorm
Linear
4
Linear
Silu
Mul
2
Linear
Linear
Add
Output
I=
FLOPs M KN = . Dbytes 4(M K + KN + M N )
(1)
The roofline intensity threshold is Fig. 4. Computational Flowchart of LLM inference
F P 32peak 4.05 × 1012 = = 135FLOPs/Byte. B 30 × 109 (2) The attainable performance is Ithreshold =
potential of the single-cluster VLIW-SIMD architecture, effectively translating theoretical throughput into actual performance. Algorithm-Level Optimization: Building upon the operator layer, we break through the traditional operator-centric scheduling paradigm. By employing a density-driven fusion strategy, multiple fine-grained operators are fused into unified macro-operators. Moreover, we innovatively propose MT Attention to replace Flash Attention, significantly reducing redundant memory accesses and kernel launch overhead while improving on-chip memory utilization. System-Level Adaptive Parallel Optimization: At the highest level, we restructure the entire inference pipeline from a system-wide perspective. A three-stage P-B-D synchronous pipeline is designed, integrating selective batching and hybrid parallel strategies. This ensures safe decoupling of the prefill and decoding phases while enabling efficient coordination and scaling of multi-cluster resources through techniques such as asynchronous KV-Cache migration and computationcommunication overlap. B. High-performance Operator Design Based on the abstract operator graph of LLMs (Fig. 4), operators are categorized into high-density (e.g., GEMM) and low-density (e.g., Softmax/RoPE) types according to their arithmetic intensity [35]. The Roofline model indicates that high-density operators are often compute-bound, while lowdensity operators are mostly I/O-bound [36]. Accordingly, double buffering is used for low-density operators to overlap computation and data transfer, while dataflow and kernel-level optimizations are applied to high-density operators to improve computational utilization. 1) GEMM Operator Analysis: Consider matrix multiplication Y = X × W , where X ∈ RM ×K , W ∈ RK×N , Y ∈ RM ×N . Combining hardware parameters (memory bandwidth B = 30 GB/s, FP32 peak multiply-accumulate performance
Pmax = min (F P 32peak , B · I) .
(3)
Due to the unique on-chip memory structure of the MT3000 processor, this paper employs the outer product method to optimize the GEMM operation [37]. Under the capacity constraints of the SM and AM in FP32 precision, the following conditions must be satisfied. 4B × M × K ≤ 64 KB.
(4)
4B × K × N + 4B × M × N ≤ 768 KB.
(5)
Under these constraints, we obtain an optimal set of configuration parameters: M = 128, K = 128, N = 768. The computational intensity I at this point is 192/13, which is significantly lower than the threshold Ithreshold , indicating that the performance is primarily limited by memory bandwidth. 2) Dataflow Design and Memory Access Optimization: For the GEMM operator, when M is large, we partition the input matrix X ∈ RM ×K along the M dimension. The entire process consists of three levels: first, X is divided into blocks Xg [Mg , Kg ] and cached into GSM leveraging its highbandwidth characteristics; subsequently, it is decomposed into P pieces of X2 [M2 , K2 ] loaded into SM, and W2 [K2 , N2 ] is broadcast to the AM of each DSP; finally, at the kernel level, X2 and W2 are further decomposed into X1 [M1 , K1 ] and W1 [K1 , N1 ] to compute the submatrix Y1 [M1 , N1 ]. This process is described in detail in the next section. The workflow is illustrated in Fig. 5. To reduce the amount of DMA data transfer, a matrix W broadcasting method is employed for acceleration, with the output matrix Y set as a stationary target. To maximize parallelism between computation and data transfer, this paper designs a triple-buffering mechanism that decouples the input, output, and computation processes of Y2 , thereby achieving parallelization of computation and data I/O. However, since Y2 is generated through multiple iterations of
PREPRINT
5
X
W
Xg DMA to GSM
Xg
P*XAg 2
W2
Broadcast to AM
DMA to AM
W1 W2 Y1 Y2
X1 Load to Register
Kernel
P*Y2
Y2
DMA to SM
X2
Y
X1
Load to Register
W1
Load to Register
Y1
Fig. 5. Data Flow Graph for GEMM Operator Computation: Three-level dynamic tiling (DDR→GSM→AM) with triple buffering for Linear operators, optimizing dataflow via W broadcasting and static output Y to reduce DMA transfers.
X2 and W2 , it tends to cause uneven transfer loads and induce tail effects (Fig. 6). To address this, a dynamic tiling strategy is introduced to further decompose Y2 into finer-grained Y2 tile and embed its transfer operations into the inner computation loop of each X2 and W2 iteration, as illustrated in Fig. 7. Based on the storage capacity constraints of the MT-3000, the optimization parameters need to meet the capacity constraints of SM, AM, and GSM. 2 × 4B × M2 × K2 ≤ 64KB,
(6)
2×4B×K2 ×N2 +3×4B×M2 ×N2 +4B×N ≤ 768KB, (7) 4B × Kg × Mg ≤ 6MB.
instruction, with the instruction pipeline scheduling outlined in Table I. Here, VMAC denotes floating-point multiply-accumulate units, SMAC represents vector multiply-accumulate units, SLDST and VLDST signify scalar and vector load/store units respectively, and SIEU stands for scalar integer execution units. Notably, in this work, every two 32-bit data elements of W1 are combined into 64-bit entities, such that each row of the right matrix processed by each accelerator core contains four double-precision values. The functionality and cycle counts of each instruction are summarized in Table II. 4) FP16 operator support: Leveraging the half-precision vector multiply-add instruction (VFMULAH16) of the MT3000, we implemented a high-performance FP16 operator library, which delivers approximately twice the peak performance of FP32 at the same frequency. A mixed-precision strategy of ”FP16 storage with FP32 accumulation” is adopted: critical normalization and reduction operations (such as Softmax and RMSNorm) are computed at FP32 intermediate precision to prevent numerical underflow. The conversion between FP16 and FP32 is efficiently handled by native vector instructions with negligible overhead. For GEMM, we optimized the parameter configuration while maintaining the FP32 kernel structure to reduce intermediate data movement and enhance bandwidth utilization.
(8)
To enhance computational efficiency, the kernel size is set to M1 = 6 and N1 = 128, with N1 = 128 chosen to better align with the Multi-Head Attention (MHA) mechanism. The optimal configuration derived from parameter exploration is given by Mg = 720, Kg = 2,048, M2 = 30, with K2 and N2 both set to 256. When the dimension M is small, we adopt a strategy of broadcasting the input matrix X and perform parallel computation along the N dimension to achieve load balancing among multiple DSPs. 3) VLIW-SIMD instruction-level optimization: To fully exploit the computational potential of the MT-3000 processor and address the suboptimal performance of the compiler-generated outer product kernel, this section proposes an optimization method for matrix multiplication based on hand-written assembly. Each accelerator core of the MT-3000 is equipped with three parallel floating-point multiply-accumulate units (VMACs). By offloading computational tasks to 16 accelerator cores via SIMD instructions, each core can efficiently handle matrix operations of size 6×K2 ×8. The detailed optimization strategies are as follows: A two-row two-column outer product computation pattern is employed. For the left matrix X1 , scalar loading, bit packing, and vector broadcasting are achieved through an instruction chain consisting of sldh → sbale2 → svbcast, while the right matrix W1 is directly loaded into vector registers using the vldw instruction. Vectorized floating-point multiplyaccumulate operations are then executed via the vfmulas32
C. Scheduling Strategy for Computational Graphs 1) Computational Graph Partitioning and Fusion Rules: Based on the operator graph in Fig. 4, we propose a densitydriven fusion method: fusion units are constructed in a “[lowdensity]–high-density–[low-density]” pattern, with each unit containing at least one high-density operator, which can incorporate multiple low-density operators either before or after it. The fusion boundaries are determined by two constraints: first, the data dependency distance—ensuring that data production and consumption form a closed loop within the unit to eliminate off-chip memory access; second, Onchip memory capacity—statically setting an upper limit based on the remaining capacity of AM, SM, and GSM to prevent working set overflow. Regions 1, 3, 4, and 5 in Fig. 4 serve as typical examples. Taking the fusion of Linear and RoPE as an example, we embed RoPE into the Linear kernel, allowing results to be consumed on-the-fly in registers or on-chip caches, thereby further reducing I/O operations and kernel launch overhead. Implementation details are provided in Algorithm 1. The algorithm uses VPU load to load the RoPE rotation angle θ into the vector register VR θ; the scalar register R row stores the row index, which is broadcast to the vector VR row. It combines the instructions vm sinf32 u35 and vm cosf32 u35 to implement complex number rotation, and VR mix is generated through VEC NEG and bale2lh instructions. 2) MT Attention: Staged Pipeline-Based Fusion Optimization: Focusing on the matrix dimensions of Qi and Ki in attention mechanisms, we propose a hardware-aware phased fusion optimization strategy. For small sequence lengths (S),
PREPRINT
6
Kg/K2 times T3
computer
Kg/K2 times
T3
T3
T3
T3
T3
T1
Transfer time of X2
T2
Transfer time of W2
T3
Computing Time
T4
Transfer time of Y2
T4_tile Transfer time of Y2_tile T4_round_0_in
Transfer
T1
T2
T1
T2
T1
T2
T1
T2
T4_round_1_in
T1
T2
T1
T2
T1
T2
T4_round_0_out
T4_round_2_in
Fig. 6. Before Tiling: Unbalanced transmission load of Y2 input and output Kg/K2 times T3
computer
Kg/K2 times
T3
T3
T3
T3
T3 T4_tile_round_0_out
Transfer
T4_round_0_in
T1
T2
T1
T2
T1
T2
T1
T2
T1
T2
T4_tile_round_1_in
T1
T2
T1
T2
T4_tile_round_2_in
Fig. 7. After Tiling: Y2 input and output are tiled to balance the transmission load TABLE I I NSTRUCTION P IPELINE S CHEDULING FOR M ATRIX M ULTIPLICATION K ERNEL VMAC
SMAC
SLDST
VLDST
SIEU
vfmulas32 X1 [0, 1, 2][0], W1 [0][0] vfmulas32 X1 [0, 1, 2][0], W1 [0][1] vfmulas32 X1 [0, 1, 2][0], W1 [0][2] vfmulas32 X1 [0, 1, 2][0], W1 [0][3] vfmulas32 X1 [3, 4, 5][0], W1 [0][0] vfmulas32 X1 [3, 4, 5][0], W1 [0][1] vfmulas32 X1 [3, 4, 5][0], W1 [0][2] vfmulas32 X1 [3, 4, 5][0], W1 [0][3] vfmulas32 X1 [0, 1, 2][1], W1 [1][0] vfmulas32 X1 [0, 1, 2][1], W1 [1][1] vfmulas32 X1 [0, 1, 2][1], W1 [1][2] vfmulas32 X1 [0, 1, 2][1], W1 [1][3] vfmulas32 X1 [3, 4, 5][1], W1 [1][0] vfmulas32 X1 [3, 4, 5][1], W1 [1][1] vfmulas32 X1 [3, 4, 5][1], W1 [1][2] vfmulas32 X1 [3, 4, 5][1], W1 [1][3]
– – svbcast X1 [0][1] svbcast X1 [1][1] svbcast X1 [2][1] svbcast X1 [3][1] svbcast X1 [4][1] svbcast X1 [5][1] – – svbcast X1next [0][0] svbcast X1next [1][0] svbcast X1next [2][0] svbcast X1next [3][0] svbcast X1next [4][0] svbcast X1next [5][0]
– – sldh X1next [0][0] sldh X1next [1][0] sldh X1next [2][0] sldh X1next [3][0] sldh X1next [4][0] sldh X1next [5][0] – – sldh X1next [0][1] sldh X1next [1][1] sldh X1next [2][1] sldh X1next [3][1] sldh X1next [4][1] sldh X1next [5][1]
vldw W1 [1][2, 3] – – – – – – vldw W1next [0][0, 1] vldw W1next [0][2, 3] – – – – – – vldw W1next [1][0, 1]
– sbale X1 [0][1] sbale X1 [1][1] sbale X1 [2][1] sbale X1 [3][1] sbale X1 [4][1] sbale X1 [5][1] – – sbale X1next [0][0] sbale X1next [1][0] sbale X1next [2][0] sbale X1next [3][0] sbale X1next [4][0] sbale X1next [5][0] –
Note: The left matrix X1 is loaded through sldh→sbale2→svbcast chain for scalar loading, bit-packing, and vector broadcasting. The right matrix W1 is directly loaded via vldw.
Algorithm 1 Linear–RoPE Fusion (compact)
TABLE II I NSTRUCTION C HARACTERISTICS Instruction
Cycle Count Description
vfmulas32
6
svbcast sldh sbale2
4 7 1
vldw
9
FP32-precision vector floating-point multiplyaccumulate Broadcast scalar register value to vector register Load single FP32 value into scalar register Pack lower 32 bits of two scalar registers into one scalar register Load 16×64-bit values into vector register
intermediate attention scores are cached in on-chip memory through operator fusion to eliminate redundant memory accesses. When S exceeds the SM capacity threshold, we dynamically adjust row-block sizes (M rows) using an outer product method and utilize GSM to cache attention scores, enabling block reuse and deep computational synergy. The optimization workflow, shown in Fig. 8, employs row-granular scheduling per attention head, with computation for a single head formulated as shown in Eq. 9. Attention(Qi , Ki , Vi ) = softmax
T
Qi Ki √ d
Vi
(9)
1: θid ← (n1 + n2 ) mod dim / 32 (floor) 2: for col = 0 to N1 /32 − 1 do 3: VPU load θ[θid ++] → V Rθ [col] 4: for row = 0 to M1 − 1 do 5: SPU load row+m0 +m2 +Pid · M2 +m1 → Rrow 6: Broadcast Rrow → V Rrow 7: for col = 0 to N1 /32 − 1 do 8: VEC NEG V Ry [row][col] → V Rneg [col] 9: bale2lh V Ry [row][col], V Rneg [col] → V Rmix [col] 10: vec muli V Rθ [col], V Rrow → V RIP [col] 11: vm sinf32 u35 V RIP [col] → V Rsin [col] 12: vm cosf32 u35 V RIP [col] → V Rcos [col] 13: end for 14: for col = 0 to N1 /32 − 1 do ▷ loop unrolled 15: muli V Ry [row][col], V Rcos [col] → V Rres [col] 16: end for 17: for col = 0 to N1 /32 − 1 do ▷ loop unrolled 18: Mula V Rmix [col], V Rsin [col], V Rres [col] → V Rres [col] 19: end for 20: end for 21: VPU store V Rres [0 : N2 /32] → Y2 [m1 +row][n1 : n1 +N1 ] 22: end for
It is divided into three stages: Stage 1: Partition Qi into row blocks (M rows) and load them into SM. Load KiT into AM in column blocks (N
PREPRINT
7
Inner Loop
KTj[dxN]
Qi[Mxd]
Vj[d'xd]
BroadCast to AM
Oi[Mxd]
Scoreij=QiKTj
Inner Loop
Outer Loop
Stage 1
computer/storage on AM
Outer Loop
DMA to SM
Attention weightsi Stage 2: scale && softmax Transfor to SM
Oi'=WijVj
Wij[Mxd']
DMA to AM Accumulate & DMA to DDR
Stage 3
Fig. 8. MT Attention: Staged Pipeline-Based Fusion Optimization Strategy
columns). After iteration, obtain the Attention Scores, with some stored in GSM and the rest in AM. Stage 2: Scale the M rows of Attention Scores obtained in Stage 1 and perform the softmax operation to get Attention Weights. Accelerate the Softmax using a hardware-customized vector reduction algorithm, with detailed steps shown in Fig. 9. Repeat steps 4-5 to ultimately obtain the reduction result of the sum inside 32 floating-point vectors. Stage 3: Perform GEMM operations on the M rows of Attention Weights and Vi to obtain M rows of the final output Oi for a single head. By iterating through these three stages multiple times, we can obtain the final result for a single head. Through this scheduling approach, the transmission of Attention Scores and Attention Weights is completely hidden, reducing I/O overhead and lowering the on-chip space complexity to O(M ). The proposed MT Attention method and Flash Attention exhibit significant differences in core I/O behavior on the MT-3000 platform. Flash Attention requires keeping the K/V blocks fixed in AM and repeatedly loads Q blocks and output O blocks via DMA. In contrast, MT Attention keeps the Q/O blocks fixed in on-chip memory, executes in parallel across multiple DSPs, and iteratively loads K/V blocks via broadcasting. Assuming each loaded K/V block contains Mc rows and each Q/O block contains Mr rows. The I/O comTABLE III I/O OPERATION COUNT COMPARISON BETWEEN MT ATTENTION AND F LASH ATTENTION I/O Type
Flash Attention
MT Attention
Q DMA O DMA K Broadcast V Broadcast
S/Mc 2S/Mc 1 1
1 1 S/(Mr × 24) S/(Mr × 24)
plexity comparison between the two methods is shown in Table III. By slightly increasing the number of broadcast operations, MT Attention significantly reduces the number of DMA operations. Since the actual broadcast bandwidth is higher than that of DMA (with broadcast latency being about 0.9 times that of DMA), this strategy effectively reduces the overall I/O overhead and latency in Attention computation.
3) Unified Kernel Scheduling and Load Balancing: This study employs the Hthreads heterogeneous threading library to implement parallel scheduling for the intra-cluster DSP acceleration array. In the current scheduling mechanism, the total execution time for a single kernel is defined as: Tsingle = Tcreate group + Tlaunch group + Texec group ,
(10)
where Tcreate group is the thread group creation time, Tlaunch group is the thread group launch time, and Texec group is the actual execution time. For a computation graph containing N operators where each requires launching a separate thread group, the total overhead becomes: Ttotal = N · Tsingle .
(11)
Under the MHA mechanism, a short input sequence length readily induces load imbalance across multiple DSPs. This problem originates from two primary factors: the coarse task granularity in per-head scheduling mode impedes effective workload partitioning and leads to low resource utilization, while the need for independent kernel initialization and termination per attention head introduces substantial repeated scheduling overhead (Tcreate group + Tlaunch group ). To address these issues, we propose a unified kernel scheduling strategy that operates at the granularity of the complete computation graph. This approach performs thread group creation and launch in a single operation, effectively eliminating multiple kernel launch overheads while enabling fine-grained task distribution across attention heads, thereby significantly alleviating load imbalance under short-sequence input scenarios. D. Adaptive Parallel Scheduling 1) Selective Batching: To enhance system throughput under bandwidth constraints, THInfer adopts a selective batching strategy. This approach is based on weight sharing analysis: except for the Q, K, V matrices in the Attention mechanism, other weights (such as those in Linear and LayerNorm layers) can be fully shared within a batch. Thus, traditional batching is applied to these compute-intensive operators. In contrast, the Attention component is processed on a per-task basis due to its I/O-intensive nature and dynamically independent KV cache. This hybrid strategy aims to meticulously balance the trade-off between latency and throughput. Let the effective batch size be B, the input sequence length be S, and the output sequence length be N . • Latency (Tlatency ) is defined as the total time required to process the prompt and generate all B × N tokens. • Throughput is defined as B × N/Tlatency (tokens/s). 2) Resource-Aware P-B-D Three-Level Synchronous Pipeline: During inference, the Prefill and Decode stages exhibit significant bottleneck heterogeneity. While a fully asynchronous decoupling strategy could improve system throughput, it may lead to KV cache accumulation and subsequent GPU out-of-memory (OOM) errors when Decode latency is high. To address this issue, THInfer designs a Prefill-Buffer-Decode (P-B-D) three-level synchronous pipeline (Fig. 10), which achieves inter-stage decoupling and parallel execution at a micro-batch granularity. By
PREPRINT
8
vector reduction VR0
0,0
0,1
0,2
0,3
0,30
0,31
VR1
1,0
1,1
1,2
1,3
1,30
1,31
VR30
30,0
30,1
30,2
30,3
30,30
30,31
VR31
31,0
31,1
31,2
31,3
31,30
31,31
0,0
1,0
0,2
1,2
0,30
1,30
VR1
0,1
1,1
0,3
1,3
0,31
1,31
VR30
30,0
31,0
30,2
31,2
30,30
31,30
VR31
30,1
31,1
30,3
31,3
30,31
31,31
0,0+1
1,0+1
0,2+3
1,2+3
VR8
8,0+1
9,0+1
8,2+3
9,2+3
0,30+31
1,30+31
8,30+31
9,30+31
VR0 4
VR8
vstdw0m16&& VR16 vstdw1m16&& VR24 vldw
9,0+1
24,6+7
0,8+9
1,8+9
8,8+9
9,8+9
24,14+15 25,14+15
0,16+17
0,0+1
1,16+17
1,0+1
8,16+17
8,0+1
9,16+17
24,22+23 25,22+23
25,6+7
0,24+25
1,24+25
8,24+25
9,24+25
24,30+31 25,30+31
6,0+1
7,0+1
14,0+1
15,0+1
30,6+7
7,8+9
14,8+9
15,8+9
30,14+15 31,14+15
7,16+17 14,16+17 15,16+17
30,22+23 31,22+23
VR16 16,0+1 17,0+1 16,2+3 17,2+3
16,30+31 17,30+31
VR24 24,0+1 25,0+1 24,2+3 25,2+3
24,30+31 25,30+31
VR6
7,2+3
6,30+31 7,30+31
VR6
VR14 14,0+1 15,0+1 14,2+3 15,2+3
14,30+31 15,30+31
VR14
6,8+9
VR22 22,0+1 23,0+1 22,2+3 23,2+3
22,30+31 23,30+31
VR22 6,16+17
30,30+31 31,30+31
VR30 6,24+25
7,24+25 14,24+25 15,24+25
30,30+31 31,30+31
1 bale VR0
VR0
6,0+1
7,0+1
6,2+3
VR30 30,0+1 31,0+1 30,2+3 31,2+3
3 group 2
add
31,6+7
5 add
VR0
0,0+1
1,0+1
0,2+3
1,2+3
0,30+31 1,30+31
VR0
0,0+1+8+9+16+17+24+25
1,0+1+8+9+16+17+24+25
8,0+1+8+9+16+17+24+25
VR2
2,0+1
3,0+1
2,2+3
3,2+3
2,30+31 3,30+31
VR2
2,0+1+8+9+16+17+24+25
3,0+1+8+9+16+17+24+25
10,0+1+8+9+16+17+24+25
VR4
4,0+1+8+9+16+17+24+25
5,0+1+8+9+16+17+24+25
12,0+1+8+9+16+17+24+25
VR30 30,0+1 31,0+1 30,2+3 31,2+3
30,30+31 31,30+31
VR6
6,0+1+8+9+16+17+24+25
7,0+1+8+9+16+17+24+25
14,0+1+8+9+16+17+24+25
Fig. 9. Hardware-Aware Reduction Algorithm: (1) Use the bale instruction to uniformly pack the high and low 32 bits of the register; (2) Add every two registers as a group; (3) Group the 16 added registers; (4) Use shuffle operations vstdw0m16 and vstdw1m16 to move the data from each group of registers to AM, and then load it back into the registers; (5) Add each group of registers to obtain the partial sum of each vector.
incorporating a bounded buffer and a backpressure control mechanism, it restricts how far the Prefill stage can advance ahead, thus structurally avoiding unbounded cache growth issues similar to those in DistServe. Buffer
Prefill
Decode Token 0
Micro Batch
Token 1
Dec PP0 Dec PP1
the Tensor Parallelism degree be T P , the Pipeline Parallelism degree be P P , the micro-batch size be Bmicro , and the Data Parallelism degree be DP . The total batch size is B. • Memory Overhead Modeling The per-cluster memory footprint during the Prefill stage is:
Dec PP2
Batch
Paramsper cluster =
Dec PP3
enqueue
Pre PP0
Pull KV && First-Token
Model size , P PP
(12)
Pre PP1
KVper cluster = Bmicro P · Smax · 2 · Dsize · Demb .
Pre PP2
The per-cluster memory footprint during the Decode stage is:
Fig. 10. P-B-D Three-Level Synchronous Pipeline
Paramsper cluster =
This pipeline divides the cluster into three logical pools based on functionality: • Prefill Pool: Handles incoming requests, performs prefill computation, and generates the first token along with its corresponding initial KV cache. This pool employs Data Parallelism (DP) and Pipeline Parallelism (PP) strategies to fully exploit GEMM computational throughput. • Buffer Pool: Acts as an intermediate staging area, maintaining a bounded buffer with a capacity of Buf max on idle nodes, specifically for temporarily storing KV cache. • Decode Pool: Synchronously fetches data at the microbatch granularity, maps the corresponding KV cache in a pull-based manner, and employs Tensor Parallelism (TP), PP and DP strategies to distribute the storage and bandwidth pressure of model weights and KV cache. To ensure system stability, a bounded buffer and a backpressure control mechanism are introduced: when buffer utilization reaches its upper limit, backpressure will be applied to the Prefill stage, pausing the acceptance of new requests, thereby preventing KV cache overflow. The Decode stage pulls KV data on-demand upon readiness, effectively shortening the cache lifecycle and suppressing peak memory usage. 3) Hybrid Parallel Strategy Based on Resource Constraints: To address the challenge that the memory requirements of models with tens of billions of parameters far exceed the capacity of a single cluster, THInfer performs systematic modeling and configuration optimization of hybrid parallelism strategies based on hardware resource constraints. Let the total model parameter size be Model size, the number of layers be L, the maximum context length be Smax , the word embedding dimension be Demb , and the byte size of the data be Dsize . Let
(13)
KVper cluster =
Model size , T P · P PD
Bmicro D · Smax · 2 · Dsize · Demb · L . TP
(14) (15)
The constraint condition is: Paramsper cluster +KVper cluster +Overhead < 20 GB. (16) •
Communication Overhead Modeling The primary communication overhead comes from the All-Reduce operation within the TP group during the Decode stage: CommT P = B · 2 · Demb · Dsize · (T P − 1),
(17)
KV cache is asynchronously migrated using MPI nonblocking communication: upon completion of a single layer computation in the Prefill stage, the corresponding KV cache is transmitted asynchronously. The Prefill Pool can simultaneously proceed with the computation of the next layer, thereby overlapping communication with computation. • Performance Modeling and Parallel Configuration Search The computation time for a single layer on a single cluster without TP is: Tsingle clu = Tnorm + Tself attention + Tnorm + Tffn .
(18)
After applying TP, the single layer computation time becomes: Tclu = Tnorm + +
Tself attention + Tnorm TP
Tffn + 2 · Tall reduce + 2 · Tlaunch group . TP
(19)
PREPRINT
9
The benefit of TP depends on the trade-off between the reduction in computation time (primarily Tself attention and Tffn ) and the introduced additional overhead. To facilitate practical deployment, we systematically perform a parallel configuration search on the Decode size based on DPD = 1: first, under the constraints of memory usage and communication overhead, we search for a T PD that balances latency and throughput; second, we determine P PD within the memory constraints, thereby obtaining the target batch size B = P PD × micro batch; the size of the Buffer Pool Bp is determined by the KV cache size corresponding to batch size B; the P PP on the Prefill side is directly computed based on the minimum memory overhead constraint; under the condition that the Decode Pool occupies T PD × P PD computing clusters, the cluster overhead of a parallel pool is: Clupool = DPP · P PP + T PD · P PD + BP .
(20)
The overall throughput is approximately: Throughput ∝
1 Clutotal . · Clupool max(Tp /DPP , Td )
(21)
Thus, maximizing throughput is equivalent to minimizing: min {Clupool · (max(Tp /DPP , Td ))} , (22) DPP
where Tp and Td represent the latency of the Prefill stage and Decode stage, respectively, when DP = 1. These can be estimated for a given batch size through simulation or stress testing. The optimal DPP configuration that satisfies the memory and bandwidth constraints is then solved. IV. E XPERIMENTS A. Experimental Setup Hardware: Experiments were conducted on MT-3000 nodes, with comparisons made against Tesla V100SPCIE-32GB and NVIDIA A800 80GB PCIe. • Model: Evaluation was performed using the Llama2 series of models, with parameter sizes ranging from 7B to 70B. Although other models were not tested, the techniques employed by the THInfer framework are equally applicable to other LLMs based on the Transformer architecture (such as Qwen, DeepSeek, etc.). • Workload: The experiments focused on high-throughput text generation tasks. A synthetic prompt dataset, padded to a uniform length, was used, with each prompt requiring the generation of 128 tokens. Two prompt lengths were selected for testing: 512 and 1024. The evaluation metric was generation throughput. • Baseline: The baseline systems included DeepSpeedInference’s [6] tensor parallelism scheme and Hugging Face Accelerate’s [32] pipeline parallelism scheme. • Implementation: THInfer is developed in pure C++ with inline assembly, its codebase comprising roughly ten thousand lines. •
B. End-to-End Throughput Tests To comprehensively evaluate system performance, end-toend throughput tests were conducted. The baseline performance of the hardware platforms involved in the comparison is shown in Table IV. TABLE IV H ARDWARE P LATFORM P ERFORMANCE S PECIFICATIONS Device MT-3000 V100S-PCIE-32GB A800 80GB PCIe
FP16 Performance
Bandwidth
32.4 TFLOPS 130 TFLOPS 312 TFLOPS
120 GB/s 1134 GB/s 1935 GB/s
To ensure a fair comparison, the experiments followed a peak-performance alignment principle: 8 MT-3000 devices were compared against 2 V100S cards, and 10 MT-3000 devices were compared against 1 A800 card. Since DeepSpeed and Accelerate demonstrated similar performance in a singlecard environment, only DeepSpeed was used as the baseline for A800. To further investigate performance under bandwidthaligned conditions, additional configurations of 18 and 16 MT3000 devices were tested to match the bandwidth levels of 2×V100S and 1×A800, respectively. The experimental results are shown in Tables V and VI. The experimental results demonstrate that under the peakperformance-aligned configuration, MT-3000 achieves significant advantages or delivers comparable performance across various model sizes and input lengths. For the 7B model, MT-3000 achieves a 62%–73% improvement over V100S (using DeepSpeed) and a 67%–84% improvement over A800, underscoring its capability to efficiently leverage GEMM computation during the prefill phase via handcrafted kernels and operator fusion. In the 13B model, MT-3000 outperforms competing systems on long-sequence (1024) tasks, indicating that its computational advantage is more pronounced in compute-intensive scenarios. For the 30B and 70B models, the performance lead of MT-3000 further expands. Notably, in the 70B model, GPU-based solutions encountered out-ofmemory (OOM) errors, whereas MT-3000 sustained efficient inference through hybrid parallelism and memory expansion mechanisms, demonstrating exceptional scalability. Under the bandwidth-aligned configuration, the system throughput shows near-linear improvement (e.g., approximately 2.25× for 7B, 13B, and 30B models, and about 2× for the 70B model), indicating that aggregated DDR bandwidth effectively alleviates the KV cache read/write bottleneck during the decoding phase. Combined with deep computational graph fusion, selective batching, and the P-B-D three-level pipeline mechanism, THInfer successfully synergizes the high computational efficiency of the prefill phase with the bandwidth scalability of the decoding phase, achieving a seamless transition from core-level compute optimization to systemlevel I/O expansion. In summary, THInfer excels in long-context and largemodel scenarios, with performance advantages becoming more significant as input length and model size increase. This validates that its hardware-software co-design effectively addresses various bottlenecks in Transformer inference. The
PREPRINT
10
TABLE V T HROUGHPUT C OMPARISON : MT-3000 VS . V100S (T OKENS / S )
7B 13B 30B 70B
Accelerate
DeepSpeed THInfer (Peak-Aligned) THInfer (BW-Aligned)
512
1024
512
1024 512
1024
512
1024
323 168 1.9
173 93 0.82
481 273 7.69
263 145 4.07
OOM
OOM
OOM
OOM
781 241 81 64
456 169 46 41
1755 543 186 129
1026 381 105 81
TABLE VI T HROUGHPUT C OMPARISON : MT-3000 VS . A800 (T OKENS / S )
7B 13B 30B 70B
DeepSpeed
THInfer (Peak-Aligned)
THInfer (BW-Aligned)
512
1024
512
1024
512
1024
584 345 116
310 184 60
OOM
OOM
975 301 105 64
570 211 61 41
1560 483 162 129
912 339 89 81
bandwidth-aligned experiments further confirm the system’s strong scalability, providing a stable and efficient solution for inference of billion-parameter models. C. Ablation Study To evaluate the effectiveness of each optimization strategy, we conducted an ablation study. The experimental setup was as follows: a prompt length of 512, generating 128 tokens, using a 12-cluster configuration. The baseline version (A0) did not employ any optimizations and used only single-node deployment with multi-node data parallelism; we then progressively introduced operator optimization (A1), computational graphlevel algorithm optimization (A2), batching (A3 series), and the P–B–D three-stage pipeline scheduling (A4). A3 and A4 correspond to system-level optimizations, and the experimental results are shown in Table VII. As shown in Table VII, the ablation study comprehensively validates the performance contribution of each optimization module within THInfer. From A0 to A4, all models exhibit a stepwise increase in throughput, with models of different scales showing distinct responses to the optimization strategies. Operator optimization (A1), leveraging FP16 handcrafted assembly kernels to fully utilize hardware compute capabilities, increased the throughput of the 30B model from 0.07 token/s to 2.47 token/s, a 35-fold improvement; the benefit of this optimization amplifies with increasing model scale, indicating that large-scale GEMM computations are more reliant on low-level compute optimizations. Computational graph scheduling optimization (A2) further provided a 22% performance gain (30B model: 2.47 → 3.02 token/s), primarily achieved by fusing operators like MHA and RoPE to reduce redundant memory access and kernel launch overhead. Batching and parallelism strategies (A3 series) yielded significant improvements. Using only selective batching with PP (A3-1) achieved a throughput of 14 token/s for the 30B model, an approximately 4.6x improvement over A2. Introducing tensor parallelism (A3-2) further increased throughput for all models, but the magnitude of gain varied significantly: the 7B model increased from 124 to 312 token/s (+152%), while
TABLE VII T HROUGHPUT COMPARISON UNDER DIFFERENT OPTIMIZATION STRATEGIES ( TOKENS / S ). Group Operator Opt. Graph Sched. Batching P–B-D Key Features A0
×
×
×
×
A1
✓
×
×
×
A2
✓
✓
×
×
A3-1
✓
✓
✓
×
A3-2
✓
✓
✓
×
A4
✓
✓
✓
✓
7B
13B 30B
Single-node deployment with 1.02 0.21 0.07 multi-node data parallelism Enable FP16 high-performance 21 5.80 2.47 operators Enable computation graph op- 26 7.22 3.02 timization Enable selective batching (PP 124 40 14 only) Enable selective batching (TP 312 63 27 + PP) Enable P–B–D three-stage 321 101 34 pipeline
the 13B and 30B models grew by approximately 58% and 93 % respectively, suggesting that as model scale increases, All-Reduce communication gradually becomes a scaling bottleneck. The P-B-D three-level pipeline scheduling (A4) delivered additional performance improvements across all model sizes, with the magnitude of enhancement becoming more substantial as model scale increased: a 3% improvement for the 7B model (312 → 321 tokens/s), a 60% increase for the 13B model (63 → 101 tokens/s), and a 26% gain for the 30B model (27 → 34 tokens/s). These results indicate that the P-B-D three-level pipeline effectively optimizes resource scheduling and data transfer efficiency, proving particularly advantageous for medium to large-scale models characterized by high decoding load and frequent KV cache transfers. D. Effectiveness of the Linear Operator This section tests the operational efficiency of the core LLM operator, Linear, on a single cluster of MT-3000. The experiment covers various scales, with sequence length (S) ranging from 1024 to 8192, input dimension (Input dim) fixed at 4096, and output dimension (Output dim) ranging from 128 to 4096. The test results are shown in Table VIII. TABLE VIII MT-3000 S INGLE -C LUSTER L INEAR O PERATOR L ATENCY T ESTING R ESULTS (FP32/FP16, U NIT: MS ) Precision Output dim
256
512
3072
4096
FP32
S=1k S=2k S=4k S=8k
1.12 1.74 3.20 6.23
1.80 2.78 3.52 5.79 7.93 3.06 4.23 5.06 8.20 11.51 4.87 8.10 9.14 15.45 21.51 9.50 13.95 17.23 29.28 41.45
768
10.04 14.66 27.76 53.77
FP16
S=1k S=2k S=4k S=8k
0.52 0.66 1.59 3.05
0.64 1.05 2.19 4.45
0.90 1.45 3.28 5.81
1024
1.62 1.95 3.62 7.67
2048
2.03 2.93 3.88 3.38 5.05 6.53 6.80 9.54 12.52 12.93 18.47 24.17
The performance calculation is shown in Eq. (23). P er =
S Inputdim Outputdim 1000 × × × GFLOPS. 1024 1024 1024 times (23)
Based on the performance evaluation of the MT-3000 singlecluster Linear operator (Fig. 11), the computational utilization increases with larger computational scales, primarily due to improved computational density and parallel efficiency
PREPRINT
11
90 80 70
Peak: 8.10 TFLOPS
S=1024 S=2048 S=4096 S=8192
8000 7000 6000
100 80 70
2500
60
5000
60
2000
50
4000
50
1500
40
3000
40
30
1000
20
500 0
256
512
768
1024
2048
Output Dimension
3072
4096
1000
0
0
3.0
20 10 256
512
768
1024
2048
Output Dimension
3072
4096
0
Fig. 11. Performance utilization and theoretical peak analysis for the Linear operator on MT-3000.
MT Attention VS NO Optimize(Head=32) MT Attention VS Flash Attention (Head=32) MT Attention VS NO Optimize(Head=64) 3.0x MT Attention VS Flash Attention (Head=64) 2.7x
30
2000
10
3.5
90
Speedup
3000
100
Performance Ratio (%) Performance (GFLOPS)
Performance (GFLOPS)
3500
Performance Ratio (%)
Peak: 4.05 TFLOPS
S=1024 S=2048 S=4096 S=8192
4000
2.5x
2.5
2.3x
Method
Head
Before Optimization
Head=32 29.20 67.54 81.77 183.45 247.50 272.79 422.49 527.43 Head=64 57.98 137.46 161.18 367.72 493.82 542.11 842.89 1058.14
512
1024
1536
2048
2560
3072
3584
4096
After MT Attention
Head=32 9.72 Head=64 18.20
25.13 49.29
32.45 63.20
79.00 107.61 119.39 191.98 156.56 212.08 236.59 381.52
248.26 496.23
After Flash Attention Head=32 14.02 Head=64 27.69
39.21 77.40
52.37 120.51 156.03 162.15 255.17 103.96 240.31 311.34 323.74 509.70
320.74 640.81
from large-scale. This allows the VLIW-SIMD computing units to operate near saturation, effectively hiding memory access latency. The FP16 precision significantly enhances computational efficiency: under the configuration of Input dim=4096, Output dim=4096, and S=8k, the FP16 implementation achieves 5686.34 GFLOPS, reaching 70.20% of the theoretical peak performance. This represents a 122% improvement compared to the FP32 performance of 2556.05 GFLOPS (63.11% utilization). The advantage stems from native hardware support for FP16 multiply-accumulate instructions (VFMULAH16) and reduced data volume, which optimizes memory access and alleviates bandwidth pressure on the data stream. The experiments also validate the effectiveness of the DDR-GSM-AM dataflow design and hand-tuned kernel design, which maximize vector register reuse through strategies such as loop unrolling and vector broadcasting. The peak utilization rate shows that the best-case scenario (FP16, Input dim=4096, Output dim=4096, S=8k) achieves 70.20% utilization, with the remaining bottleneck primarily attributed to DDR bandwidth limitations. E. Scaling Dot-Product Attention Operator Optimization Test This section evaluates the MT Attention optimization strategy for the scaled dot-product attention operator. Table IX compares the performance before and after optimization. The experimental results indicate that the phased fusion optimization of MT Attention significantly reduces the latency of the Scaled Dot-Product Attention operator. In long-sequence scenarios (e.g., S=4096, Head=64), performance improves by over 53%, achieving a speedup of 2.5–3.5 times compared to the unoptimized version. Compared to FlashAttention, MT Attention is better adapted to the MT-3000 architecture, reducing latency by 26%–36% in most tests. Its advantage stems from effectively reducing DMA transfers and redundant I/O through broadcast mechanisms and a phased pipeline, fully leveraging the parallel potential of the heterogeneous
2.3x
2.2x
2.1x
2.0 1.5 1.4x
TABLE IX L ATENCY COMPARISON BEFORE / AFTER ATTENTION OPTIMIZATIONS FOR SCALED DOT- PRODUCT ATTENTION ( BATCH SIZE = 4, HEAD DIM = 128; UNIT: MS ).
2.3x
1.0
500
1.6x
1000
1.6x
1500
1.5x
2000
1.4x
2500
Sequence Length
1.4x
1.3x
1.3x
3000
3500
4000
Fig. 12. Speedup of MT Attention on MT-3000
many-core architecture. This optimization, synergizing with the system-level pipeline design, further validates the critical role of hardware-software co-design in enhancing the inference performance of large models. V. C ONCLUSION This paper presents THInfer, an LLM inference framework designed for the Tianhe New-Generation Supercomputer. It effectively addresses the key challenges of deploying models with tens of billions of parameters on a many-core architecture with limited memory bandwidth and distributed memory hierarchy. Through hardware-software co-design and cross-layer optimization, three core contributions are achieved: a handoptimized VLIW-SIMD operator library with FP16 support, boosting computational efficiency to 70.2% of the theoretical peak; a density-driven operator fusion strategy and the MT Attention mechanism, significantly reducing I/O and kernel scheduling overhead; and a P–B–D pipeline with hybrid parallelism, enabling efficient decoupling of decoding and prefill phases while supporting scalable and stable inference even on 70B models. Experiments demonstrate that THInfer significantly outperforms GPU-based baseline systems in throughput, validating its effectiveness, scalability, and robustness under practical workloads. This work provides an essential technical pathway and practical foundation for efficient LLM inference on heterogeneous high-performance architectures. ACKNOWLEDGMENT This work was supported by the National Key R&D Program of China under Grant 2023YFB3001900, as well as the National Natural Science Foundation of China (NSFC) under Grants U23B2020 and 62322201. R EFERENCES [1] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023. [2] J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang et al., “Qwen technical report,” arXiv preprint arXiv:2309.16609, 2023.
PREPRINT
[3] D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025. [4] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th symposium on operating systems principles, 2023, pp. 611–626. [5] NVIDIA, “TensorRT-LLM,” https://github.com/NVIDIA/ TensorRT-LLM. [6] R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley et al., “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2022, pp. 1–15. [7] J. Chen, C. Liu, Z. Luana, M. Gong, Q. Li, and D. Qian, “Largescale parallelization and optimization of lattice qcd on tianhe new generation supercomputer,” in 2023 IEEE International Conference on High Performance Computing & Communications, Data Science & Systems, Smart City & Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys). IEEE, 2023, pp. 499–506. [8] K. Lu, Y. Wang, Y. Guo, C. Huang, S. Liu, R. Wang, J. Fang, T. Tang, Z. Chen, B. Liu et al., “Mt-3000: a heterogeneous multi-zone processor for hpc,” CCF Transactions on High Performance Computing, vol. 4, no. 2, pp. 150–164, 2022. [9] M. Khalilov and A. Timoveev, “Performance analysis of cuda, openacc and openmp programming models on tesla v100 gpu,” in Journal of Physics: Conference Series, vol. 1740, no. 1. IOP Publishing, 2021, p. 012056. [10] D. W. Walker and J. J. Dongarra, “Mpi: a standard message passing interface,” Supercomputer, vol. 12, pp. 56–68, 1996. [11] A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems, 2017. [12] M. V. Koroteev, “Bert: a review of applications in natural language processing and understanding,” arXiv preprint arXiv:2103.11943, 2021. [13] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019. [14] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020. [15] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023. [16] A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [17] A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan et al., “Deepseek-v3 technical report,” arXiv preprint arXiv:2412.19437, 2024. [18] Q. Team, “Qwen2 technical report,” arXiv preprint arXiv:2407.10671, 2024. [19] E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “Gptq: Accurate post-training quantization for generative pre-trained transformers,” arXiv preprint arXiv:2210.17323, 2022. [20] J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,” Proceedings of Machine Learning and Systems, vol. 6, pp. 87–100, 2024. [21] E. Frantar and D. Alistarh, “Sparsegpt: Massive language models can be accurately pruned in one-shot,” in International Conference on Machine Learning. PMLR, 2023, pp. 10 323–10 337. [22] X. Ma, G. Fang, and X. Wang, “Llm-pruner: On the structural pruning of large language models,” Advances in neural information processing systems, vol. 36, pp. 21 702–21 720, 2023. [23] Y. Gu, L. Dong, F. Wei, and M. Huang, “Minillm: Knowledge distillation of large language models,” arXiv preprint arXiv:2306.08543, 2023. [24] T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems, vol. 35, pp. 16 344–16 359, 2022. [25] I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The longdocument transformer,” arXiv preprint arXiv:2004.05150, 2020.
12
[26] S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self-attention with linear complexity,” arXiv preprint arXiv:2006.04768, 2020. [27] N. Kitaev, Ł. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” arXiv preprint arXiv:2001.04451, 2020. [28] J. Ainslie, J. Lee-Thorp, M. De Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai, “Gqa: Training generalized multi-query transformer models from multi-head checkpoints,” arXiv preprint arXiv:2305.13245, 2023. [29] Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava, “Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,” Advances in Neural Information Processing Systems, vol. 36, pp. 52 342–52 364, 2023. [30] G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” arXiv preprint arXiv:2309.17453, 2023. [31] G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for {Transformer-Based} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), 2022, pp. 521–538. [32] Hugging Face, “Text Generation Inference,” https://github.com/ huggingface/text-generation-inference. [33] Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” in International Conference on Machine Learning. PMLR, 2023, pp. 31 094–31 116. [34] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “{DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24), 2024, pp. 193–210. [35] V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient processing of deep neural networks: A tutorial and survey,” Proceedings of the IEEE, vol. 105, no. 12, pp. 2295–2329, 2017. [36] S. Williams, A. Waterman, and D. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Communications of the ACM, vol. 52, no. 4, pp. 65–76, 2009. [37] K. Yu, X. Qi, P. Zhang, J. Fang, D. Dong, R. Wang, T. Tang, C. Huang, Y. Che, and Z. Wang, “Optimizing general matrix multiplications on modern multi-core dsps,” in 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 2024, pp. 964– 975.
Yao Lu born in 1998, PhD candidate with Beihang University, Beijing China. His main research interests are parallel computing and high-performance computing.
Zhongzhi Luan (Member, IEEE) is an Associate Professor with the School of Computer Science and Engineering, Beihang University. He has been engaged in teaching and researching distributed computing, high-performance computing, and computer architecture for many years. His main research interests include resource management in distributed environments, supporting methods and technologies for application development on heterogeneous computing systems, performance analysis, and optimization of high-performance computing applications.
PREPRINT
13
Gen Li received the B.S. degree in computer science and technology from the School of Computer Science at Beihang University, Beijing, China, where he is currently pursuing the M.S. degree. His main research interest is high-performance computing.
Jiaxing Qi received the B.S. degree in software engineering from the Industrial and Commercial College, Hebei University, Baoding, China, in 2017. And the M.S. degree in software engineering from Hebei University, Baoding, China, in 2020. He is currently pursuing the Ph.D. degree at the School of Computer Science and Engineering, Beihang University, Beijing, China. His main research interests include text analytics, and natural language processing.
Shiqing Ma received the B.S. degree in computer science and technology from the School of Computer Science at Beihang University, Beijing, China, where he is currently pursuing the M.S. degree. His main research interest is high-performance computing.
Bin Han received the B.S. degree in computer science from the School of Computer Science at Beihang University, Beijing, China, where he is currently pursuing his PhD degree. His main research interests include high-performance computing and task scheduling systems.
Shizhe Shang received the B.S. degree in information and computing science from Beihang University, Beijing, China, where he is currently pursuing the Ph.D degree. Besides, he attended a doubledegree program in Applied Math with Centrale Nantes, Nantes, France. His main research interests include high-performance computing, performance analysis and machine learning.
Hailong Yang (Member, IEEE) received the Ph.D. degree from the School of Computer Science and Engineering, Beihang University, in 2014, where he is currently working as an Associate Professor. He has been involved in several scientific projects, such as performance analysis for big data systems and performance optimization for large scale applications. His research interests include parallel and distributed computing, HPC, performance optimization, and energy efficiency. He is a Member of China Computer Federation.
Depei Qian received the master’s degree from the University of North Texas, Denton, Texas, in 1984. He is a professor with the School of Computer Science and Engineering, Beihang University, China. He served as the chief scientist of China National High Technology Program on high performance computing for 20 years. His research interests include innovative technologies in distributed computing, high performance computing, and computer architecture. He is an academician of Chinese Academy of Science, and also a fellow of the China Computer Federation (CCF).