Conceptio
›
parallel-computing
Topic
parallel-computing
Knowledge-graph topic
· documents ABOUT parallel-computing across the archive
1,734
Documents about parallel-computing
Documents about parallel-computing
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems
#204749
arXiv CS
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
#204751
arXiv CS
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
#204752
arXiv CS
LatentBox: An Efficient Latent-First Storage System for AI-Generated Images
#204757
arXiv CS
Near-Resolution of the Tradeoff Conjecture in Distributed Proof Labeling Schemes
#204763
arXiv CS
Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning
#204764
arXiv CS
AI-Driven Multi-Region Provisioning for Cloud Services Using Spot Fleets
#216778
arXiv CS
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
#216788
arXiv CS
Hypergraph Partitioning on GPU with Distinct Incident Hyperedges and Size Constraints
#216803
arXiv CS
DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving
#224466
arXiv CS
Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics
#224470
arXiv CS
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
#224478
arXiv CS
Reducing Internal State in Eigenvalue-Only Divide-and-Conquer Tridiagonal Eigensolvers
#229465
arXiv CS
Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
#229474
arXiv CS
CARM Tool: Cache-Aware Roofline Model Automatic Benchmarking and Application Analysis
#238563
arXiv CS
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
#238565
arXiv CS
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
#238576
arXiv CS
EFaaS: A Quantum-Classical Serverless Entangled Scheduler for Hybrid Variational Algorithms
#238587
arXiv CS
TAPAAL SMC: Statistical Model Checking of Stochastic Timed-Arc Petri Nets
#246506
arXiv CS
Post-Deterministic Distributed Systems: A New Foundation for Trustworthy Autonomous Infrastructure
#246512
arXiv CS
Fail-Closed Lowering of Resident KV Claims onto LLM Serving Runtimes
#246519
arXiv CS
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
#246523
arXiv CS
Magnum.np.distributed: Accelerating Finite Difference Micromagnetic Simulations with Multiple GPUs
#246525
arXiv CS
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
#246529
arXiv CS
ScanWeaver: Compiler-Driven Parallelization of Affine Recurrences via Associative Scan Lowering
#246531
arXiv CS
Augur: Pre-Execution Energy Prediction for Workflow Tasks in Heterogeneous Clusters
#246532
arXiv CS
HeLoCo: Efficient asynchronous low-communication training under data and device heterogeneity
#246534
arXiv CS
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
#246540
arXiv CS
Energy-Efficient Aggregation and Minimum-Degree Spanning Trees in Radio Networks
#246542
arXiv CS
Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference
#259421
arXiv CS
Graph Traversal on Tensor Cores: A BFS Framework for Modern GPUs
#259426
arXiv CS
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location
#259433
arXiv CS
OpenAgenet/OAN: Technical Architecture for Trust-Governed Agent Identity and Discovery
#259446
arXiv CS
AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis
#267609
arXiv CS
Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM Serving
#267618
arXiv CS
When More Cores Hurts: The Vector Database Scaling Paradox in HPC
#267619
arXiv CS
APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing
#267623
arXiv CS
Vivace: Exact Temporal OLAP over Interval Histories via Independent Serverless Execution
#271781
arXiv CS
Temporal Conductance and Bounds on the Voter Model for Dynamic Networks
#271788
arXiv CS
On the Limits of Performance Portability in Directive-Based GPU Programming
#271800
arXiv CS
Eidola: Modeling Multi-GPU Network Communication Traffic in Distributed AI Workloads
#271803
arXiv CS
Near-Optimal Distributed 2-Ruling Sets on Graphs with Low Arboricity
#271808
arXiv CS
Optimizing Cloud Deployment: Blending of IaaS and FaaS for Microservice Architecture
#271811
arXiv CS
PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity
#280171
arXiv CS
Quantifying the Impact of Lossy Compression on Neural Generative Surrogate Modeling
#280172
arXiv CS
Green SARC: Predictive Cost and Carbon Governance for Agentic AI Systems
#280173
arXiv CS
SDVDiag: Multimodal Causal Discovery for Online Diagnosis in Software-defined Vehicles
#280174
arXiv CS
NEURON-Fabric: CXL-Side Low-Bit Gradient Aggregation for Distributed Training
#280182
arXiv CS
RouteBalance: Fused Model Routing and Load Balancing for Heterogeneous LLM Serving
#282762
arXiv CS
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators
#282773
arXiv CS
← Previous
Page 11 of 20
Next →
Topic record
· derived from the Conceptio knowledge graph (shared subject terms across the corpus)
Conceptio Open Knowledge Archive — topic hubs link to canonical document pages with full provenance.