arXiv:2604.24088v1 [cs.DC] 27 Apr 2026
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training Man Liu
Xingchen Liu
Xingjian Tian
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences Hangzhou, China [email protected]
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences Hangzhou, China [email protected]
Bing Lu
Shengkai Lyu
Shengquan Yin
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
University of Science and Technology of China Hefei, China [email protected]
Wenjing Huang
Zheng Wei
Hairui Zhao
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Guangming Tan
Dingwen Tao
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Institute of Computing Technology, Chinese Academy of Sciences Beijing, China [email protected]
Abstract
CCS Concepts
Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under frequent communication and introduce significant computational overhead during compression. To this end, we propose TACO (Tensor-parallel Adaptive COmmunication compression), a robust FP8-based framework for compressing TP intermediate tensors. First, we employ a data-driven reshaping strategy combined with an Adaptive Scale–Hadamard Transform to enable high-fidelity FP8 quantization, while its Dual-Scale Quantization mechanism ensures numerical stability throughout training. Second, we design a highly fused compression operator to reduce memory traffic and kernel launch overhead, allowing efficient overlap with communication. Finally, we integrate TACO with existing state-of-the-art methods for Data and Pipeline Parallelism to develop a compression-enabled 3Dparallel training framework. Detailed experiments on GPT models and Qwen model demonstrate up to 1.87× end-to-end throughput improvement while maintaining near-lossless accuracy, validating the effectiveness and efficiency of TACO in large-scale training.
• Software and its engineering → Message passing; • Theory of computation → Data compression.
This work is licensed under a Creative Commons Attribution 4.0 International License. HPDC ’26, Cleveland, OH, USA © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2640-8/2026/07 https://doi.org/10.1145/3806645.3807584
Keywords Tensor parallelism, quantization, distributed training, communication compression, large language model. ACM Reference Format: Man Liu, Xingchen Liu, Xingjian Tian, Bing Lu, Shengkai Lyu, Shengquan Yin, Wenjing Huang, Zheng Wei, Hairui Zhao, Guangming Tan, and Dingwen Tao. 2026. TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training. In The 35th International Symposium on High-Performance Parallel and Distributed Computing (HPDC ’26), July 13–16, 2026, Cleveland, OH, USA. ACM, New York, NY, USA, 14 pages. https://doi.org/10.1145/3806645.3807584
1
Introduction
The rapid scaling of large language models (LLMs) to tens of billions, hundreds of billions, and even trillion-parameter scales has driven the adoption of increasingly sophisticated distributed training strategies [11, 22, 24, 44]. Among them, 3D parallelism — comprising data parallelism (DP), tensor parallelism (TP), and pipeline parallelism (PP)—has emerged as the dominant paradigm for training ultra-large models [10, 57]. However, as model scale grows, training performance becomes increasingly constrained by communication rather than computation. Recent system studies show
HPDC ’26, July 13–16, 2026, Cleveland, OH, USA TP Comm. Other Comm.
36.0% 45.1% 33.8%
27.6%
34.2% 31.2%
Others
43.9%
30.3%
30.2% 27.3%
34.6% 25.8%
25B 39B Llama
18B 30B GPT