Hybrid GPU–CPU Retrieval for Personalized Search at Ultra-Large Scale Hao Fu Jichao Sun Baiting Zhu Qiaoling Liu Yan Shi Cheng Lu Liu Liu Yubo Wang Xin Yao Xiangyu Niu Xu Dong Wenhan Lyu Chiyao Shen Yinjie Huang Minglei Chen Shuai Ding Li Fan Xiao Kong Meta Platforms, Inc. Menlo Park, California, USA
arXiv:2609.21281v1 [cs.IR] 18 Sep 2026
Abstract Embedding-based retrieval on user-generated content at the trilliondocument scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the personalization–scale paradox: hosting the full serving inventory in GPU memory is too resource intensive, while CPU compute cannot execute the same interaction-heavy model on the latency-critical path. We present a hybrid GPU–CPU co-serving system that resolves the paradox through orchestration rather than a new model class. A high-depth GPU pathway fuses retrieval and interaction preranking over a curated online pool on the order of a billion documents, while a high-breadth CPU pathway searches an independently selected online inventory roughly twenty times larger with lightweight personalized scoring. Either or both pathways can run per request; candidates are deduplicated before shared downstream ranking. The system is deployed in production. A full-system A/B test against the legacy CPU-only configuration improves model-scored relevance and substantive engagement, while separate pathway experiments show positive value at their own deployment scopes. Retrieval logs show that the pathways contribute structurally distinct candidates, production serving measurements characterize their latency, and a matched capacity plan quantifies the economic rationale for assigning modeling depth to GPUs and inventory breadth to CPUs. Together, these results validate a practical, independently evolvable depth–breadth architecture for ultra-large-scale personalized search.
CCS Concepts • Information systems → Search engine architectures and scalability; Retrieval models and ranking; • Computer systems organization → Heterogeneous (hybrid) systems.
Keywords Personalized Search, Embedding-based Retrieval, Hybrid Architecture, GPU-CPU Co-serving, Vector Search, Scalability ACM Reference Format: Hao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Minglei Chen, Shuai Ding, Li Fan, and Xiao Kong. 2027. KDD ’27, San Jose, CA, USA 2027.
Hybrid GPU–CPU Retrieval for Personalized Search at Ultra-Large Scale. In Proceedings of The 33rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’27). ACM, New York, NY, USA, 10 pages.
1
Introduction
Embedding retrieval over user-generated content depends on query semantics and evolving user interests. Discovery queries benefit from non-linear interactions with behavioral history, while navigational and specific queries require rare items from a broader inventory. At a multi-trillion-document source corpus, this creates a personalization–scale paradox: GPUs support deep interaction but cannot economically hold the full online inventory in highbandwidth memory (HBM); CPUs provide affordable capacity but cannot run the same interaction model at scale. User-generated content intensifies this tension: the same query can require different candidates as interests evolve, while new posts arrive and expired, deleted, or policy-ineligible items disappear. Serving must support changing user-conditioned scores and broad, frequently refreshed coverage; measuring only full-precision model quality or approximate nearest-neighbor (ANN) recall misses one side of the objective. A discovery query may require behavioral history to prefer seafood over fast food, whereas an obscure restaurant-review query succeeds only if that sparse document remains retrievable. Modeling depth and inventory breadth consume different resources: interactions increase per-candidate compute, while rare content increases memory and scan cost. A homogeneous tier must compromise one side of this frontier. We resolve this conflict with heterogeneous co-serving. A GPU pathway retrieves and pre-ranks over a curated inventory on the order of a billion documents, while a CPU pathway uses lightweight embedding matching over an inventory roughly twenty times larger, including popular content. Either or both can run per request and be disabled independently for capacity control or rollback. In a separate matched broad-vector plan, annualized accelerator capacity cost is roughly four times the CPU cost. We study candidate generation and GPU interaction pre-ranking. Downstream full ranking, query-understanding training, generative retrieval, and learned request selection are out of scope; lexical retrieval appears only in measurements at the shared-ranker input. Our deployed architecture uses established two-tower and DeepFMstyle components, validated through a full-system A/B test, pathway experiments, serving measurements, and source-attributed retrieval logs. Concretely, this paper makes three contributions:
KDD ’27, August 1–5, 2027, San Jose, CA, USA
• Heterogeneous co-serving system. We jointly serve an independently selected billion-scale GPU inventory and a CPU inventory roughly twenty times larger through separately versioned publication, deadlines, rollback, and a shared aggregation interface. • Depth–breadth pathway design. The GPU pathway fuses two-tower retrieval with interaction scoring. The CPU pathway combines a dedicated embedding index, eager term-at-atime scoring, bulk scanning, and finer clustering; component cost reductions are reported against their own baselines rather than multiplied together. • Production validation. Separate pathway experiments, source-attributed retrieval logs, and the full-system A/B test establish pathway-level value, candidate diversity, and system-level value, respectively. We additionally report a conservative cross-day-dependence-robust GSRR interval, serving measurements, and a computational-cost comparison between GPU and CPU broad-inventory plans.
2 Related Work 2.1 Personalized Retrieval and Interaction Deep retrieval spans DSSM [7], two-tower models [3], and graphbased industrial recommendation [1, 24]. Language models have also improved representations and retrieval [21, 25]. These systems generally retain separable query–document scoring for efficient candidate generation. ColBERT [10], generative DSI [18], and the target-aware generate–rank design of GRank [17] increase expressiveness but make large-scale serving more demanding. We apply non-linear interaction scoring inside an accelerator-resident retrieval pass [23] before candidates leave the GPU.
2.2
Industrial Scaling and Heterogeneous Systems
Industrial examples include PinSage [24], Facebook search [6], Alibaba commodity retrieval [20], and Que2Search [12]. They scale through hard-negative mining, real-time intent modeling [4], and multi-tower architectures. SVFusion [14] and FusionANNS [19] divide one ANN call across CPUs and GPUs. Our system instead operates distinct model families and independently selected inventories that meet at aggregation.
2.3
Efficient Indexing and Tiered Serving
IVF-PQ [8], HNSW [13], and DiskANN [16] optimize the cost– recall frontier of a single ANN engine. Disk-resident SPANN [2] and SPFresh [22] extend that frontier through tiered posting lists and online index maintenance, while GPU libraries such as Faiss [9] and cuVS [15] provide high scan throughput but remain bounded by HBM capacity. These advances are complementary to our contribution: either pathway could adopt a new index or kernel without changing the co-serving and candidate-aggregation contracts. Unlike SVFusion and FusionANNS, which co-process one ANN request, our independently selected billion- and tens-of-billionsscale inventories meet at aggregation. The GPU also pre-ranks candidates before aggregation. Section 6.5 compares broad-inventory capacity plans.
Hao Fu et al.
Table 1: Architectural comparison at each paper’s reported evaluation scale. Fused pre-rank denotes learned interaction scoring within the retrieval pass, before shared aggregation. System Facebook search [6] DiskANN [16] SPANN [2] SVFusion [14] FusionANNS [19] This work
3
Scale
Hardware
Fused pre-rank
100M-scale index 1B points 1B vectors Up to 1B vectors 1B vectors
CPU CPU CPU CPU+GPU CPU+GPU
No No No No No
Billion + tens-of-billions scale
CPU+GPU
Yes
System Architecture
Figure 1 separates the upstream source corpus from online retrieval: the multi-trillion figure describes the corpus before selection, not an ANN scan. After policy, quality, language, freshness, and duplication checks, hardware-specific selectors produce a GPU snapshot on the order of a billion documents and a CPU index roughly an order of magnitude larger. The CPU index is not a disjoint collection of low-popularity content; it also contains many popular documents eligible for the GPU pool. Because selection and publication are versioned independently, containment can vary.
3.1
Parallel Retrieval and Aggregation
Query understanding derives semantic, language, intent, and personalization signals from the query and available user context. Capacity and availability controls determine which pathways run; both can run for the same request. These controls support independent rollback and are not evaluated as a learned decision policy. Active branches execute in parallel. The GPU branch combines ANN retrieval with interaction pre-ranking, while the CPU branch searches its dedicated embedding index. Each branch has its own deadline and can return independently. The aggregator deduplicates by document identifier and preserves source attribution before the shared downstream ranker. The branches execute under independent deadlines rather than extending a serial critical path. Section 6.4 reports production latency for lexical retrieval and the CPU and GPU pathways; Section 4.3 separately reports the narrower per-accelerator modelserver benchmark.
3.2
Failure Isolation and Versioned Release
The fork–join boundary also isolates failures. Each branch has a deadline; aggregation proceeds with candidates returned in time, so one slow branch does not discard the other’s result. Releases are independent: the GPU publishes a matched model–index snapshot, while the CPU combines versioned centroids and a base index with live updates. Either pathway can be disabled without changing the aggregation interface.
4
High-Depth GPU Retrieval
The GPU pathway jointly trains retrieval and interaction scoring to prioritize modeling depth. This section follows its lifecycle from pool construction and multi-objective training to acceleratorresident retrieval and pre-ranking.
Hybrid GPU–CPU Retrieval for Personalized Search at Ultra-Large Scale
KDD ’27, August 1–5, 2027, San Jose, CA, USA
ONLINE REQUEST PATH Input Query
User Context
Query Understanding Capacity / Availability Routing
GPU ANN Retrieval Interaction Pre-Ranking
serving inventory
Candidate Aggregation Deduplication + Source Attribution
CPU Embedding Retrieval Lightweight Scoring
serving inventory
Shared Downstream Ranker OFFLINE INVENTORY PUBLICATION
Search-Value-Selected Billion-Scale GPU Inventory
Search-Value Selection
Broad CPU Inventory Tens-of-Billions Scale
Multi-Trillion Source Corpus
Broad-Coverage Selection
Figure 1: GPU–CPU co-serving over independently selected billion- and tens-of-billions-scale inventories from a multi-trilliondocument source corpus.
4.1
Construction of the High-Quality Pool
Because high-bandwidth memory cannot hold the full inventory, the GPU pool optimizes search value, not raw virality. Recommendation engagement is often dominated by entertainment with little query specificity, while local guides and tutorials may be essential to explicit search intents despite modest aggregate engagement. The selector protects this informative content while pruning viral items with low search utility.
Lifecycle-aware filtration. Content passes through three stages as it ages. At ingest, hard metadata gates enforce language, originality, and content type. During the first week, early-engagement filters remove spam and low-value items. From day 7 onward, taxonomy, search-usefulness, local-information, and temporal-decay models estimate long-term utility. Staging prevents mature-content thresholds before new items accrue evidence.
Adaptive retention. Decay is category-dependent: niche tutorials and local guides with high search value bypass engagement pruning, while news and trends decay aggressively. A highly selective evergreen pool retains under 0.1% of content older than one year. The policy keeps the billion-scale pool fresh without discarding low-frequency content that remains useful for explicit search intents.
4.2
Model Architecture and Joint Training
The two-tower retriever and interaction pre-ranker are jointly trained on search logs, using InfoNCE and hard negatives to align semantic relevance and engagement. Figure 2 combines user-query and document towers with a DeepFM-based interaction module [5]. The query tower forms a 512-dimensional dense input from a 384-dimensional semantic embedding and a 128-dimensional user-profile embedding. A residual multilayer perceptron processes this input with sparse categorical context. The document tower consumes dense, categorical, and content-embedding features and emits a 128-dimensional vector for nearest-neighbor retrieval. Joint training aligns retrieval and scoring distributions. Training uses three weighted objectives: contrastive InfoNCE for retrieval, Smooth L1 for relevance, and binary cross-entropy (BCE) for engagement, LTotal = 𝑤 1 LInfoNCE + 𝑤 2 LSmoothL1 + 𝑤 3 LBCE,
(1)
with hyperparameters 𝑤 1, 𝑤 2, 𝑤 3 . Let ℎ(𝑞, 𝑑) = exp(sim(𝑞, 𝑑)/𝜏). For a batch of size 𝐵 with anchor query 𝑞𝑖 , positive document 𝑑𝑖+ , and negative documents 𝐷 𝑗 drawn from every other session 𝑗 ≠ 𝑖, the two-tower stack uses the in-batch loss 𝐵 ℎ(𝑞𝑖 , 𝑑𝑖+ ) 1 ∑︁ Í Í . (2) LInfoNCE = − log + 𝐵 𝑖=1 ℎ(𝑞𝑖 , 𝑑𝑖 ) + 𝑗≠𝑖 𝑑 ∈𝐷 𝑗 ℎ(𝑞𝑖 , 𝑑) The interaction pre-ranker predicts relevance and engagement from impression labels or later-stage ranker distillation, using binary cross-entropy for engagement and Smooth L1 for relevance. In
KDD ’27, August 1–5, 2027, San Jose, CA, USA
Hao Fu et al.
Figure 3: GPU inference; sparse tables remain on CPU. Figure 2: Joint training combines InfoNCE retrieval with engagement and relevance objectives.
addition to the cross-session loss in Eq. (2), a within-session InfoNCE term treats engaged candidates as positives and relevant but non-engaged candidates as negatives. These harder session-local negatives already match the query. Í A value model computes 𝑆 final = 𝑡 𝜆𝑡 𝑝ˆ𝑡 . The per-task weights balance relevance and engagement objectives and can vary by traffic segment. Consequently, one trained model can serve heterogeneous traffic and test new operating points without retraining. Appendix A.3 summarizes objective supervision and serving roles. Training lifecycle. The pre-ranker trains continuously from daily refreshed search logs. An automated pipeline publishes model weights and the candidate index together, so both reflect the same corpus view. Scheduled retraining is independent of capacity configuration, allowing an experiment to change the active operating point without rebuilding the model.
4.3
The serving stack follows SilverTorch [23] and keeps ANN search, inference, and pre-ranking accelerator-resident, avoiding the CPU– GPU transfer that would otherwise separate retrieval from scoring. Figure 3 shows the interaction pre-ranker filtering retrieved candidates before the value model combines its predictions. Candidate vectors and interaction features remain on the accelerator through the compute-intensive stages. The compute backbone resides in high-bandwidth memory, while sparse embedding tables are fetched on demand from host memory over fifth-generation PCI Express (PCIe Gen5). This hybrid memory hierarchy lets model parameters exceed accelerator-memory capacity without moving latency-critical dense kernels off the GPU. Fused 8-bit integer kernels combine quantization and distance computation in one operator, reducing intermediate materialization and memory bandwidth during coarse candidate selection. The deployment uses AMD MI300X accelerators. In a production load test, one card sustained roughly 150–200 queries per second at approximately 30–40 ms 99th-percentile model-server latency. These per-accelerator figures exclude network, aggregation, and downstream ranking.
5 Index freshness. The daily cadence matches the application’s freshness requirements rather than a hardware limit. It also aligns the model checkpoint with the training-data view of the corpus. Other applications can choose a different cadence without changing the model or aggregation interface.
Accelerator-Resident Serving
Broad-Inventory CPU Retrieval
The CPU pathway complements GPU modeling depth with an orderof-magnitude larger inventory and a faster update loop. It trades away deep per-document interaction but preserves personalized semantic retrieval across a broad inventory that cannot be hosted economically in HBM.
Hybrid GPU–CPU Retrieval for Personalized Search at Ultra-Large Scale
5.1
KDD ’27, August 1–5, 2027, San Jose, CA, USA
CPU Retrieval Architecture
Pretrained centroids key inverted embedding lists. A query first selects nearby centroids for coarse pruning, then scans the selected lists and combines cosine similarity with lightweight lexical signals. A dedicated document-understanding service consumes both the real-time create/update stream and the static pool, dispatches content for embedding generation, and loads the resulting vectors into the serving index. Embedding generation is therefore decoupled from CPU retrieval even though the index itself resides in DRAM.
5.2
Dedicated Embedding Index
Embedding search was previously confined to subsets of lexical indices. Sharing that substrate forced vector scoring through termoriented abstractions and made its capacity contend with latencysensitive lexical workloads. A dedicated embedding index reduces this coupling and expands the pool by more than an order of magnitude. Document selection. The dedicated index holds tens of billions of public documents selected from the multi-trillion-document upstream corpus. Selection covers the principal languages in the target market, recent and historically useful content, and minimum quality requirements. Documents, embeddings, and a minimal lexical signal set are partitioned and replicated in dynamic random-access memory (DRAM) on commodity x86 servers. A batch builds the base index, while an online service applies updates. Distributed cluster training. The 10× larger pool requires proportionally more centroids to control list size. The previous indexingtime procedure trained 64k centroids separately per shard; extending the same serial workflow to 512k takes prohibitively long and blocks deployment. We decouple training from indexing: an independent distributed trainer runs daily across many machines, materializes a versioned centroid index, and publishes it for serving shards to ingest. Index servers no longer repeat clustering. Section 6.3 evaluates the 64k→512k operating point.
5.3
Efficient Serving: Eager Evaluation and Hit Scanning
Legacy document-at-a-time (DAAT) execution thrashes cache on the tens-of-billions-scale pool because vector scoring jumps among document identifiers scattered across posting lists. We use two complementary optimizations. Eager evaluation. We invert the query plan instead of interleaving nearest-neighbor scoring with Boolean constraints. The system first scans relevant nearest-neighbor clusters term-at-a-time (TAAT) with contiguous reads that maximize cache locality, computes a global Top-𝐾, and then uses that candidate list as a dynamic skip list for the remaining DAAT constraints. Heavy vector computation therefore completes before Boolean processing, converting scattered random reads into sequential scans and preventing computationally expensive distance calculations on documents that cannot enter the prefetched set (Figure 4). Hit scanning. The iterator-based hit abstraction incurs a virtualfunction call for every matched document. We bypass it by passing
Figure 4: Eager retrieval for three nearest neighbors with prefetch multiplier 2 scans clusters before applying other rules to the top-6.
raw vectors in bulk to 512-bit vector-instruction (AVX-512) listscanning kernels, shifting cycles from dispatch and pointer chasing to useful distance computation. Both optimizations run on the dedicated embedding index, so high-throughput vector scans do not contend with latency-sensitive lexical operations or inherit their iterator overhead.
5.4
Lightweight Personalized Modeling
The CPU model applies lightweight contextual personalization and multilingual semantics to the tens-of-billions-scale inventory. Its purpose is not to reproduce GPU interaction depth, but to retain a useful user-conditioned representation under a much lower percandidate compute budget. Unified two-tower model. User-query and document towers share a fine-tuned XLM-V backbone [11], trained with bidirectional InfoNCE to produce compact 384-dimensional multilingual vectors. An attention module fuses the dense query representation with sparse contextual features through a pooled embedding lookup. The shared backbone reuses document embeddings across contexts while keeping personalization query-side. Scale trade-offs. User and document vectors interact only by dot product, excluding the cross-feature interactions used by the GPU pathway. Features are limited to lightweight profile, coarse context, and history signals. These restrictions keep document vectors precomputable and make scans suitable for single-instruction, multiple-data execution. The deliberate loss of expressiveness buys
KDD ’27, August 1–5, 2027, San Jose, CA, USA
inventory breadth: the CPU pathway retrieves from tens of billions of documents, while later stages apply richer modeling only to the much smaller surviving set.
6
Evaluation
We ask four questions. RQ1: What is the end-to-end value of the high-depth GPU pathway? RQ2: Can the CPU pathway serve a tensof-billions-scale inventory efficiently and improve online quality? RQ3: Does the combined system improve production quality while contributing distinct candidates? RQ4: For a fixed broad-vector workload, what capacity and computational cost follow from GPU versus CPU placement?
6.1
Experimental Setup and Metrics
We evaluate the system through production A/B tests, serving benchmarks, and retrieval logs. The headline comparison is a fullsystem A/B test against the legacy CPU-only configuration. It uses persistent account-level assignment, split approximately equally between control and treatment. Daily assigned populations are on the order of millions of accounts per arm. The seven-day GPUpathway experiment and nine-day CPU-pathway evaluation with the GPU pathway disabled ran separately against contemporaneous production controls. Metrics. DCG@20 is model-scored relevance over the top 20 results, not a human-rater measure; we interpret it with behavioral outcomes rather than as independent ground truth. GSRR (Good Search Result Rate) is a retention-oriented metric of substantive engagement. For the eligible search sessions S, let E𝑠 be the events observed in session 𝑠, 𝑣 (𝑒) the strength of event 𝑒, and 𝜏type(𝑒 ) its event-type-specific threshold. We define 𝑔(𝑠) = ⊮ ∃𝑒 ∈ E𝑠 : 𝑣 (𝑒) ≥ 𝜏type(𝑒 ) , (3) ∑︁ 1 GSRR = 𝑔(𝑠). (4) |S| 𝑠∈S
Thus, a good session contains at least one qualifying event, such as a sufficiently long view or an explicit positive action. Statistical reporting. All reported metric changes are relative lifts, computed as 100(𝜇𝑡 /𝜇𝑐 − 1) from the treatment and control means. Statistical uncertainty is computed at the account-assignment level, so sessions from the same account are not treated as independent. For full-system GSRR, we give equal weight to each of the nine consecutive daily lifts. Persistent assignment also reuses accounts across days, so we do not√treat daily estimates as independent or divide their uncertainty by 9. Instead, for each day we take the larger side of its reported 95% interval and average the nine half-widths. By the Cauchy–Schwarz bound, the resulting half-width is conservative under arbitrary cross-day dependence, including perfect correlation. It yields the rounded 95% interval [+1.71%, +2.32%]; all nine daily intervals are also above zero. The ± values for the separate pathway experiments below are their two-sided 95% interval half-widths. All reported lifts and intervals are rounded to two decimal places. Over the same nine-day window, the hybrid treatment improved the equal-day mean DCG@20 by 4.51% and the equal-day mean GSRR by 2.01% relative to the legacy all-CPU configuration. The
Hao Fu et al.
Table 2: Production evaluation. Relative lifts use contemporaneous controls; experiments ran in separate windows and should not be compared across rows. For the hybrid row, both intervals are the conservative nine-day interval described in Appendix A.5. Configuration
DCG@20
GSRR
High-depth GPU Broad-inventory CPU
+0.73% ± 0.15% +3.16% ± 0.15%
+1.04% ± 0.13% +1.20% ± 0.11%
Hybrid architecture
+4.51% [+3.99%, +5.04%]
+2.01% [+1.71%, +2.32%]
separate pathway experiments establish positive online value for each pathway against its contemporaneous production control. Candidate-set analysis shows that they contribute structurally distinct results, while the full-system A/B test validates the deployed co-serving operating point. Together, these results demonstrate complementarity in the operational sense. Evaluation scope. The production tests evaluate complete pathways under real-time state, policy filters, and a shared downstream ranker rather than isolated ANN algorithms. Because each treatment changes multiple components, its online lift measures the complete treatment configuration; attribution to a kernel, inventory change, or model layer requires a corresponding isolated experiment.
6.2
GPU Pathway: End-to-End Modeling Value (RQ1)
A predecessor semantic-model deployment illustrates why offline model quality alone is insufficient. It improved offline relevance but regressed online viewing. Investigation identified three deployment mismatches: offline evaluation used exact nearest-neighbor search while serving used a quantized ANN configuration; offline document text was richer than the online representation; and a downstream filter remained calibrated to the previous model. The embedding, ANN configuration, online features, and downstream filters therefore form one deployment unit. The deployed GPU pathway changes accelerator serving, the ANN operating point, and interaction scoring together, so Table 2 reports an end-to-end comparison. The seven-day experiment measured +1.04% ± 0.13% GSRR and +0.73% ± 0.15% DCG@20. A production load test measured on the order of ten billion floating-point operations per query (roughly twice that at peak) at a per-card throughput of a few hundred queries per second. This compute profile makes interaction pre-ranking practical but does not isolate hardware, model, or ANN effects.
6.3
CPU Pathway: Scale and Efficiency (RQ2)
With the GPU pathway disabled, the nine-day CPU-only evaluation measured the lightweight CPU personalization model together with a relevance-filter update. It delivered +1.20% ± 0.11% GSRR and +3.16%±0.15% DCG@20 (Table 2). This is a complete CPU-pathway estimate: it does not isolate inventory breadth, the personalization model, the dedicated index, or the filter update.
Hybrid GPU–CPU Retrieval for Personalized Search at Ultra-Large Scale
KDD ’27, August 1–5, 2027, San Jose, CA, USA
Table 3: Production retrieval-branch wall time for posts over seven complete days. Call counts and latency percentiles are sampling-weighted and measured at the retrieval-branch boundary. Branch
Calls P50 (ms) P95 (ms) P99 (ms)
Lexical 3.43M CPU pathway 2.11M GPU pathway 0.94M
196 41 151
686 439 675
1,184 859 1,376
Table 4: Macro-averaged GPU–CPU dense-retrieval overlap over three days. GPU-side is |𝐺 ∩𝐶 |/|𝐺 |; CPU-side is |𝐺 ∩𝐶 |/|𝐶 |. Figure 5: Recall–candidate frontier for 64k and 512k centroid indices as the number of probed clusters varies from 1 to 1024. The 512k index reaches comparable recall with fewer scanned candidates.
Segment (queries) GPU-side CPU-side Jaccard
Dedicated indexing and eager execution reduced retrieval-stage compute by 89.45% (9.48×). Figure 5 separately shows the offline recall–candidate frontier for 64k and 512k centroid indices. A 15day online experiment then evaluated the selected 512k-centroid configuration with 256 probed clusters (𝑛 probe = 256): it reduced CPU use by 63.5% without a detected GSRR regression. We report the two reductions separately because their baselines differ and each measurement includes fixed overhead.
execution, not the historical A/B window or complete-search latency, and exclude query understanding, network transit outside the measured branch boundary, candidate aggregation, downstream ranking, and client rendering. The independent benchmark in Section 4.3 reports a tens-of-milliseconds P99 for the accelerator model server alone and is narrower still.
6.4
Overall Production Value and Candidate Diversity (RQ3)
The production A/B directly compares the deployed hybrid system with its legacy all-CPU predecessor. Over nine consecutive days the hybrid treatment improves the equal-day mean of modelscored DCG@20 by 4.51% (conservative 95% confidence interval [+3.99%, +5.04%]) and of GSRR by 2.01% ([+1.71%, +2.32%]; Table 2). Every daily interval of both metrics is above zero. These headline estimates validate the complete production operating point; they do not identify an interaction effect or assign the lift among hardware, inventory, models, eligibility controls, and aggregation. Production retrieval-branch latency. Table 3 reports seven days of production wall time from entry to exit of each post-retrieval branch at the top-level aggregator. We restrict log records to post retrieval and apply the logging system’s sampling weights. For the CPU pathway, we combine all dense retrieval work belonging to the same retrieval call before computing percentiles; the displayed values are therefore not averages of per-source percentiles. Retrieval for short-form video and other video surfaces is excluded. The CPU pathway is faster than lexical retrieval at all three reported percentiles. The GPU pathway is faster at P50 and P95, while its P99 is 16% higher. This is not a like-for-like kernel comparison: the GPU pathway applies interaction pre-ranking and passes approximately 224 candidates to the shared-ranker input, versus 59 from the dense CPU pathway (Table 5). The distributions therefore characterize two deployed operating points with different candidate budgets and scoring depth, rather than ranking one pathway’s computational efficiency. They describe current production branch
Head (85,617) Torso (17,057) Tail (11,405)
2.41% 2.08% 2.20%
1.15% 0.60% 0.52%
0.75% 0.44% 0.41%
Candidate-set analysis. For each search request in a three-day log window for which both dense pathways ran, let 𝐺 be the GPU retrieval set and 𝐶 the union of three dense CPU retrieval sets. We compute containment and Jaccard per request and average within query-popularity segments. Head is the top popularity decile, torso is deciles 2–5, and tail is deciles 6–10. Table 4 shows low overlap across approximately 114k eligible requests. Low overlap establishes candidate diversity, not user value by itself. The production system A/B supplies the overall user-value evidence, while the separate pathway launches show gains for the complete deployed pathways under their own controls. Composition at the shared-ranker input provides a second structural check: the pool averages approximately 781 deduplicated candidates per eligible request, with independently rounded group averages of 209 GPU-only, 47 dense-CPU-only, 500 lexical-CPU-only, and 26 cross-group candidates. These counts do not establish final Top20 exposure or per-source engagement attribution; Appendix A.2 gives definitions and robustness checks. Contribution at the shared-ranker input. Both dense pathways continue to contribute after their own scoring stages. On the same three-day slice, the GPU pathway retrieves an average of 594 candidates and passes 224 to the shared-ranker input. The deduplicated dense CPU union retrieves 1,352 and passes 59 (Table 5). Because their candidate budgets and filtering stages differ, the reach rates describe two candidate pipelines rather than rank pathway quality. GPU interaction scoring selects from the curated billion-scale pool under a precision-oriented budget, while CPU retrieval starts with a roughly twenty-times-larger inventory and applies a lighter filter. The production A/B tests supply the user-value evidence. The result is also stable to a longer window. A separate seven-day check against the same dense CPU union gives GPU-side overlap
KDD ’27, August 1–5, 2027, San Jose, CA, USA
Hao Fu et al.
Table 5: Dense-pathway candidate flow over three days. Counts are rounded averages over eligible content-search requests for which the GPU pathway ran; reach rates use unrounded counts. The CPU row is the deduplicated union of all dense CPU sources. Pathway
Retrieved
Ranker input
Reach rate
594 1,352
224 59
37.6% 4.4%
GPU pathway Dense CPU union
Table 6: Normalized capacity-plan inputs for a matched 𝑑 = 256 SQ8 workload with three regional replicas. Capacity-cost values use the CPU plan as the reference. Planning input Unit capacity cost Query throughput/unit Vector capacity/unit Units/region 3-region capacity cost
CPU host
Accelerator
1.0× 1.0× 1.0× 1.0× 1.0×
∼ 12× ∼ 3–4× ∼ 2–3× ∼ 1/3× ∼ 4×
of 2.32%, 2.07%, and 2.15% for head, torso, and tail, within 0.13 percentage points of Table 4.
6.5
Broad-Inventory Capacity Cost (RQ4)
The deployed CPU inventory is at tens-of-billions scale; this capacity plan is separate. It compares GPU and CPU placement for the same broad vector-search workload using 256-dimensional vectors encoded with scalar 8-bit quantization (SQ8). Each alternative uses three full regional replicas for disaster recovery and satisfies the same regional throughput requirement. Table 6 reports normalized planning inputs rather than absolute fleet counts. Storage, rather than the throughput floor, determines both unit counts. The accelerator stores roughly two to three times as many vectors per unit and therefore uses roughly one-third as many units per region. After normalizing the CPU plan to 1.0×, the rounded planning inputs give an accelerator unit-capacity-cost ratio of approximately 12× and a three-region accelerator/CPU capacity-cost ratio of roughly 4×. The comparison excludes shared ranking, networking, the deployed billion-scale GPU pathway, and utilization-dependent overhead. It also does not establish latency or retrieval-quality equivalence between the two implementations; it is an auditable capacity-plan result for the stated vector representation, not total capacity cost for the deployed hybrid system. Appendix A.4 details the plan’s scope and assumptions.
7
Discussion and Future Work
The stable interface in this deployment is the candidate contract: each pathway owns its inventory, model, eligibility, publication cadence, deadline, and rollback, while the aggregator owns source attribution and deduplication. The failed predecessor in Section 6.2 shows why an embedding alone is not a release unit: the online representation, ANN operating point, filters, model, and index must be versioned as a compatible bundle. Evidence also needs matched scope. The full-system A/B establishes overall value, pathway experiments estimate their complete treatment configurations, and
overlap logs test structural redundancy without substituting for causal evidence. Likewise, the capacity comparison fixes representation, inventory, throughput, and replication; it informs one placement decision rather than whole-system capacity cost. The pathway experiments change multiple components, so their gains cannot be assigned to hardware, ANN precision, model depth, or inventory size individually. Learned request-level routing is a direct next step, but unbiased evaluation requires randomized pathway assignment with quality, deadline, and capacity-cost logging. Finally, shared-ranker input diversity does not reveal sourceattributed Top-20 exposure or engagement; those outcomes require attributed impressions and actions.
8
Conclusion
We presented a deployed depth–breadth retrieval architecture that treats model expressiveness and inventory coverage as separate serving objectives. A GPU pathway performs interaction-heavy retrieval over a curated inventory, while a CPU pathway supplies broad semantic coverage. Independent selection, publication, deadlines, and rollback allow both pathways to co-serve behind a stable candidate contract and evolve without forcing one implementation to follow the other. Production experiments validate the combined operating point, pathway launches establish value at their own deployment scope, and retrieval logs show that the pathways contribute structurally distinct candidates. The broader lesson is that ultra-large-scale personalized retrieval need not force one hardware tier to optimize incompatible objectives: heterogeneous pathways can specialize, evolve independently, and be evaluated as production bundles while preserving explicit quality, latency, and capacity-cost trade-offs.
Ethical Considerations Deployment and controlled experiments followed institutional privacy and experimentation-review processes. They added no data collection or direct recruitment; models use existing access-controlled signals, and we report aggregates under retention controls. These aggregates do not establish subgroup parity; subgroup quality and creator exposure remain limitations.
GenAI Usage Disclosure GenAI assisted source discovery and editing; the authors verified all claims and remain responsible for the manuscript.
Hybrid GPU–CPU Retrieval for Personalized Search at Ultra-Large Scale
References [1] Fedor Borisyuk et al. 2024. LiGNN: Graph Neural Networks at LinkedIn. arXiv preprint arXiv:2402.11139 abs/2402.11139 (2024), 1–12. [2] Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-efficient billion-scale approximate nearest neighborhood search. In Advances in Neural Information Processing Systems (NeurIPS). Curran Associates, Virtual, 5199–5212. [3] Paul Covington et al. 2016. Deep neural networks for youtube recommendations. In RecSys. ACM, Boston, MA, USA, 191–198. [4] Mihajlo Grbovic et al. 2018. Real-time personalization using embeddings for search ranking at airbnb. In KDD. ACM, London, UK, 311–320. [5] Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine based Neural Network for CTR Prediction. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence. International Joint Conferences on Artificial Intelligence Organization, Melbourne, Australia, 1725–1731. doi:10.24963/ijcai.2017/239 [6] Jui-Ting Huang et al. 2020. Embedding-based retrieval in facebook search. In KDD. ACM, Virtual Event, CA, USA, 2553–2561. [7] Po-Sen Huang et al. 2013. Learning deep structured semantic models for web search using clickthrough data. In CIKM. ACM, San Francisco, California, USA, 2333–2338. [8] Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez. 2011. Product Quantization for Nearest Neighbor Search. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, 1 (2011), 117–128. doi:10.1109/TPAMI.2010.57 [9] Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021. Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data 7, 3 (2021), 535–547. doi:10. 1109/TBDATA.2019.2921572 [10] Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In SIGIR. ACM, Virtual Event, 39–48. doi:10.1145/3397271.3401075 [11] Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023. XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 13142–13152. doi:10.18653/v1/2023.emnlp-main.813 [12] Yiqun Liu et al. 2021. Que2Search: fast and accurate query and document understanding for search at Facebook. In KDD. ACM, Virtual Event, 3376–3384. [13] Yu A. Malkov and Dmitry A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 4 (2020), 824–836. doi:10.1109/TPAMI.2018.2889473 [14] Yuchen Peng, Dingyu Yang, Zhongle Xie, Ji Sun, Lidan Shou, Ke Chen, and Gang Chen. 2026. SVFusion: A CPU-GPU Co-Processing Architecture for LargeScale Real-Time Vector Search. Proceedings of the VLDB Endowment 19, 5 (2026), 1074–1087. doi:10.14778/3796195.3796216 [15] Junjie Qi, Gergely Szilvasy, Michael Norris, and Vishal Gandhi. 2025. Accelerating GPU indexes in Faiss with NVIDIA cuVS. https://engineering.fb.com/2025/05/08/ data-infrastructure/accelerating-gpu-indexes-in-faiss-with-nvidia-cuvs/. Meta Engineering blog; accessed 2026-07-26. [16] Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnaswamy, and Rohan Kadekodi. 2019. DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node. In Advances in Neural Information Processing Systems (NeurIPS). Curran Associates, Vancouver, BC, Canada, 13748–13758. https://papers.nips.cc/paper/2019/hash/ 09853c7fb1d3f8ee67a61b6bf4a7f8e6-Abstract.html [17] Yijia Sun, Shanshan Huang, Zhiyuan Guan, Qiang Luo, Ruiming Tang, Kun Gai, and Guorui Zhou. 2026. GRank: Towards Target-Aware and Streamlined Industrial Retrieval with a Generate-Rank Framework. In Proceedings of the ACM Web Conference 2026. ACM, New York, NY, USA, 7798–7808. doi:10.1145/3774904. 3792810 [18] Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, and Donald Metzler. 2022. Transformer Memory as a Differentiable Search Index. In Advances in Neural Information Processing Systems, Vol. 35. Curran Associates, Inc., Red Hook, NY, USA, 21831–21843. [19] Bing Tian, Haikun Liu, Yuhang Tang, Shihai Xiao, Zhuohui Duan, Xiaofei Liao, Xuecang Zhang, Junhua Zhu, and Yu Zhang. 2024. FusionANNS: An Efficient CPU/GPU Cooperative Processing Architecture for Billion-scale Approximate Nearest Neighbor Search. arXiv preprint arXiv:2409.16576 abs/2409.16576 (2024), 1–15. [20] Jizhe Wang, Pipei Huang, Huan Zhao, Zhibo Zhang, Binqiang Zhao, and Dik Lun Lee. 2018. Billion-scale Commodity Embedding for E-commerce Recommendation in Alibaba. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, London, UK, 839–848. doi:10.1145/ 3219819.3219869
KDD ’27, August 1–5, 2027, San Jose, CA, USA
[21] Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Improving Text Embeddings with Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics, Bangkok, Thailand, 11897–11916. doi:10.18653/v1/2024.acl-long.642 [22] Yuming Xu, Hengyu Liang, Jin Li, Shuotao Xu, Qi Chen, Qianxi Zhang, Cheng Li, Ziyue Yang, Fan Yang, Yuqing Yang, Peng Cheng, and Mao Yang. 2023. SPFresh: Incremental in-place update for billion-scale vector search. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP). ACM, Koblenz, Germany, 545–561. doi:10.1145/3600006.3613166 [23] Bi Xue et al. 2026. SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUs. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, New York, NY, USA, 2139–2149. doi:10.1145/3805712.3809755 [24] Rex Ying et al. 2018. Graph convolutional neural networks for web-scale recommender systems. In KDD. ACM, London, UK, 974–983. [25] Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Haonan Chen, Zheng Liu, Zhicheng Dou, and Ji-Rong Wen. 2025. Large Language Models for Information Retrieval: A Survey. ACM Transactions on Information Systems 44, 1 (2025), 1–54. doi:10.1145/3748304
A
Supplementary Definitions and Measurement Details A.1 Per-Pathway Serving Footprint Table 7: Serving scopes. The GPU benchmark is per accelerator and excludes end-to-end search stages.
Hardware Inventory
GPU pathway
CPU pathway
MI300X accelerator
Commodity x86 DRAM hosts Tens-of-billions-scale inventory Base backfill + live updates
Billion-scale selected pool Refresh Daily model + snapshot Measurement Roughly 150–200 Online CPU-use queries/s/card; measurements 99th-percentile latency ∼30–40 ms Excludes Network, aggregation, No comparable per-host downstream ranker latency reported
The scopes in Table 7 must not be conflated. The per-card tens-ofmilliseconds value characterizes the GPU model server. Table 3 measures the wider retrieval-branch execution block for each pathway over posts, including its in-process orchestration. Neither measurement includes the complete search request through downstream ranking and client rendering.
A.2
Overlap and Shared-Ranker Input Composition
For each search request 𝑖 for which both dense pathways ran, let 𝐺𝑖 be the GPU set and 𝐶𝑖 the union of the three dense CPU retrieval sets. We compute |𝐺𝑖 ∩ 𝐶𝑖 | |𝐺𝑖 ∩ 𝐶𝑖 | , 𝑂𝐶 (𝑖) = , (5) 𝑂𝐺 (𝑖) = |𝐺𝑖 | |𝐶𝑖 | |𝐺𝑖 ∩ 𝐶𝑖 | 𝐽 (𝑖) = , (6) |𝐺𝑖 ∪ 𝐶𝑖 | then average each ratio over requests in a popularity segment. Requests in which only one pathway runs are excluded because
KDD ’27, August 1–5, 2027, San Jose, CA, USA
Hao Fu et al.
the question is redundancy when both run. The main result uses a three-day window. A separate seven-day computation gives GPUside overlap of 2.32%, 2.07%, and 2.15% for head, torso, and tail, within 0.13 percentage points of the main window. Table 8: Deduplicated composition at the shared-ranker input per eligible search request. Counts are independently rounded; “Overlap” contains documents attributed to more than one group. Exclusive source group
Candidates
Share
GPU only Dense CPU only Lexical CPU only Cross-group overlap
209 47 500 26
27% 6% 64% 3%
Total
∼781
100%
These are candidate counts at the shared-ranker input, not final Top-20 impressions. They therefore support source diversity but not per-source exposure or engagement claims.
A.3
Training Objectives by Pathway Table 9: Training objectives and serving roles.
Path
Objective
GPU
Cross-session In-batch documents from other sessions; InfoNCE learns the ANN space. Within-session In- Engaged positives and relevant nonfoNCE engaged negatives. Binary cross- Engagement and relevance targets for inentropy + Smooth teraction scoring. L1 Weighted value Task weights select an operating point. model Bidirectional Produces reusable 384-dimensional docInfoNCE ument vectors.
GPU GPU
GPU CPU
A.4
Supervision and serving role
Capacity-Plan Assumptions
Table 6 concerns a 𝑑 = 256 SQ8 representation, not the deployed CPU model’s 384-dimensional representation. For this broad-vector workload, the accelerator provides roughly three to four times the per-unit query throughput and two to three times the per-unit vector capacity of a commodity CPU host. The resulting storagedominated plan uses roughly one-third as many accelerator units per region. These ratios are distinct from the interaction-model benchmark in Section 4.3. The CPU plan includes modest inventory headroom, making the reported accelerator/CPU cost ratio conservative relative to storage-minimal CPU provisioning. The source does not report utilization, a paired latency target, or a recall/quality-equivalence test, and the plan does not include the complete hybrid stack.
A.5
Measurement Procedure and Reproducibility
The primary result is the relative change from a production A/B against the legacy all-CPU control. The experiment uses persistent account-level assignment, split approximately equally between control and treatment, and includes millions of assigned accounts per arm per day. Both headline metrics are aggregated identically over the same nine consecutive days. Because the same accounts√can recur across days, we do not divide a daily standard error by 9. The Cauchy–Schwarz bound permits arbitrary cross-day covariance and upper-bounds the aggregate 95% half-width by the mean of the nine daily half-widths, taking the larger reported side for each day. For GSRR, this procedure yields the conservative 95% confidence interval [+1.71%, +2.32%] around an equal-day mean lift of +2.01%; for DCG@20 it yields [+3.99%, +5.04%] around +4.51%. Every daily interval of both metrics is above zero. The GPU-pathway experiment and GPU-disabled CPU-pathway evaluation ran in separate windows against their contemporaneous production controls and report their own two-sided 95% intervals; they are not pooled with the full-system test. Overlap is computed per eligible request from a later log window and averaged by popularity. The manuscript provides the equations, treatment scopes, aggregation procedures, and aggregate tables. Production configurations and privacy-sensitive logs remain unavailable.