ConceptioArchivearXiv CS
arXiv CSopen access

The Model in the Middle: Toward AI-Native Real-Time Communication

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributedsystemsprotocols
networking, internet, protocols, distributed systems

The Model in the Middle: Toward AI-Native Real-Time Communication Ziqian Liu

Minghao Li

Yiming Qiu

The University of Hong Kong

The University of Hong Kong

The University of Hong Kong

arXiv:2607.25792v1 [cs.NI] 28 Jul 2026

Abstract Full-duplex omni models are transforming human–AI interaction from turn-based exchanges into continuous multimodal conversations in which speaking, listening, and reasoning unfold concurrently. Rather than viewing the model as a replacement for a human endpoint, we argue for a new perspective: the model is a stateful computational middlebox inside a human-centered feedback loop, with network transport, model serving, and user playback jointly shaping how the interaction evolves. This perspective breaks the traditional boundaries among stages designed around local objectives. Rather than optimizing them in isolation, an AI-native realtime stack should allow the state of each stage to shape the actions of the others. We explore three cross-stage coordination opportunities: network-aware inference scheduling, execution-aware transport prioritization, and playback control that accounts for both network and model variability. We are building C ONFLUX to explore these ideas, and preliminary results show substantial improvements in response latency and playback deadline adherence under network degradation. More broadly, we call for an AI-native real-time communication stack that resolve the joint control problem spanning communication, computation, and playback.

1

Gemini Live

Hunyuan Voice

Aug. 2024

May. 2025

Jun. 2025

Apr. 2026

Jul. 2026

Jun. 2026

Apr. 2026

GPT-Live

JoyAI

Seeduplex

GLM-Realtime Qwen-3.5-Omni-Realtime

Figure 1: Real-time omni-model development. Transport

Client

Server Inference Framework

Audio

Video

Text

(Buffer, Schedule, Preprocess) Chunk

WebSocket WebRTC WebTransport

Omni model Modality Encoder LLM (AR) Modality Generator

Figure 2: Full-duplex omni-model workflow. consumes it, whether that receiver is human or AI. Full-duplex omni-model interaction calls for a fundamentally different perspective. As Figure 2 shows, a full-duplex model continuously updates its state from streaming chunks while producing corresponding output chunks from that evolving state, forming a latency-sensitive uplink–model–downlink path. As such, we argue that a full-duplex model is best understood as a stateful computational middlebox inside a human-centered feedback loop. The central shift is not that one receiver becomes AI, but that communication, computation, and user reaction now co-evolve as a single closed-loop system. This closed-loop structure breaks the traditional boundaries among network transport, model serving, and user playback. These components were designed around different local objectives: transport delivers media efficiently, model serving manages inference latency and throughput, and playback absorbs jitter to maintain smooth output. In full-duplex interaction, however, each component changes the operating conditions faced by the others. Network delay determines when input becomes available to the model; model serving determines when chunks enter the downlink; and the model’s modality, fidelity, and freshness requirements constrain what the network must deliver. Playback must then absorb variability introduced by both the network and the model. Full-duplex AI thus shifts the design focus from optimizing individual stages to jointly optimizing their interactions across the end-to-end path.

Introduction

Large language models are rapidly moving beyond turn-based question answering toward real-time, full-duplex multimodal interaction [12, 14, 29, 37, 42, 49]. Instead of submitting a complete query and waiting for a response, users can speak while sharing their camera or screen, interrupt the model mid-sentence, point to new visual evidence, and continuously refine their intent as the interaction unfolds. A growing landscape of systems, as shown in Figure 1, including Gemini Live [18], Qwen-3.5-Omni-Realtime [31], and GPT-Live [1], is beginning to support this experience. Their capabilities enable applications ranging from visual assistance [12, 21, 30] and real-time translation [11, 17, 34] to meeting agents [5, 20], interactive avatars [10, 19, 39], and embodied AI [13, 25]. More fundamentally, they turn model interaction into a continuous real-time process in which perception, reasoning, generation, and user feedback overlap in time [1, 43]. Conventional real-time communication [2, 3, 15] and LLM serving [23, 44, 48] both operate under a sender–receiver abstraction: one endpoint produces information and the other 1

Ziqian Liu, Minghao Li, and Yiming Qiu

Model Inference

Uplink

In this paper, we envision an AI-native real-time communication stack that accounts for the interactions among network transport, model serving, and user playback. Unlike recent systems that optimize communication via costly model context and task semantics [7, 40], we argue that broad crossstage coordination can be achieved through execution-level properties. Our central idea is that chunks provide a common coordination substrate across the end-to-end path. As a chunk progresses through transport, inference, and playback, it can carry and accumulate cross-stage signals such as the remaining latency budget, while preserving a stable identity. We therefore call for chunk-level interfaces that expose the system conditions and requirements needed for coordination. Such interfaces allow stages to jointly determine when each chunk should be transmitted, processed, and presented, without requiring them to understand what the chunk means. Our first observation is that network conditions can change the urgency and usefulness of model-serving work. Conventional serving schedulers make admission and batching decisions largely from server-local signals, such as queuing time, request age, and GPU efficiency. In full-duplex interaction, however, chunks reach the model after accumulating different amounts of network delay, leaving them with different remaining latency budgets. A delayed barge-in or turn-transition chunk may require immediate processing before the model produces more obsolete output, whereas a visual update may already be too stale to justify further queuing or computation. Model serving should therefore schedule chunks by their remaining value to the interaction, not merely by their arrival order or batching efficiency. Conversely, our second observation is that the model can expose how input chunks should compete for network resources. We argue that transport prioritization need not depend on online semantic interpretation [7, 40], because full-duplex chunks already differ in execution-level properties: some are urgent or irreplaceable, some require reliable delivery, while others can be degraded, deferred, or replaced by fresher inputs. The model runtime can expose these properties as a transport specification, which the sender combines with dynamic signals such as deadline slack and network conditions to prioritize chunks across modalities and channels. This allows barge-in, speech, control traffic, and visual context to take precedence over one another without requiring the transport to infer what any chunk means. The opportunity is therefore to make network transport execution-aware without making it semantics-aware. Our third observation is that full-duplex AI presents a new playback decision: when should each model response begin? Conventional real-time communication maintains an ongoing media timeline, with buffering primarily used to absorb network jitter. A full-duplex model, in contrast, alternates between listening and producing finite responses. Each time

Chunk 1

Chunk2 ...

Downlink

Queue and Batch Timeline

Audio:16kbps

Arrive at 1200ms Delay!

Frame:480p Inference Engine

Text

1000ms

Chunk size: 500ms t0 Last input captured

t1

t2

t3

Uplink chunk sent

Inference starts

First output arrives

t4 Playback Starts

TTFR

Figure 3: Full-duplex omni-model transport flow.

a response starts, the receiver must establish a new playback timeline from the first output chunk. Starting immediately minimizes time to first response, but leaves little protection if subsequent chunks are delayed; waiting creates playback slack, but directly slows the interaction. Moreover, future arrivals depend not only on downlink conditions, but also on model serving progress. AI-native playback must therefore decide when each response is ready to begin, absorbing variability across the entire uplink-model-downlink path. We are building C ONFLUX, an AI-native real-time communication system that instantiates these observations across transport, model serving, and playback stages. We describe its design principles and provide initial evidence of its potential.

2 Motivation 2.1 Transport Pattern in Full-Duplex AI Full-duplex omni-models treat chunks as the basic transport objects. As shown in Figure 3, each uplink chunk represents a bounded interaction interval, e.g., a 500 ms window, and contains multiple aligned modalities, such as speech samples at a given bitrate, selected video frames with configurable resolution and frame rate, text, and control events. These chunks are transmitted over the uplink using transport channels such as WebRTC or WebSocket. Then, each input chunk proceeds through the server-side model pipeline, where it is preprocessed, encoded, and batched. The autoregressive LLM subsequently samples an interaction action: silence, producing no media output; delegate, invoking a background task; or response. A response action generates a token sequence, which is rendered by modality-specific generators, such as TTS, and packaged into a sequence of output chunks for downlink delivery. This process establishes a logical correspondence between each input chunk and the resulting output chunks. The receiver maintains a playback timeline anchored by the first output chunk, with subsequent deadlines determined by chunk duration and playback buffer. Chunks arriving before their deadlines enable continuous playback, whereas late arrivals introduce gaps and user-perceived glitches. 2

The Model in the Middle: Toward AI-Native Real-Time Communication

First round

Table 2: Modality-specific transport characteristics.

Q: Describe large language model. A: Large language model is a type...

Modality Urgent? Reliable? Replaceable? Adaptive?

Q: Stop, what is on my hand? (barge-in)

Good

Control Audio Video Text

A: It is a white cup.(barge-in succeed) Q: Stop, what is on my hand? (barge-in)

Bad

A: of artificial intelligence...(continue last question)

Figure 4: A barge-in example. Table 1: Barge-in latency by network condition. Metric

Good Constrained Limited

Bandwidth (Mbps) 20 RTT (ms) 20 Jitter (ms) 0 Barge-in latency (ms) 542.7

2.2

10 40 5 686.6

5 80 10 1202.9

2.3

High High Medium Medium

High Medium Low High

No No Yes No

No Yes Yes No

Model-Defined Transport Requirements

Recent work [27, 40, 41] shows that model accuracy does not improve monotonically with input volume: insufficient fidelity can cause inference errors, while additional data provides little benefit beyond a sufficient quality level. However, identifying this boundary from the semantic context of each request requires task-specific online analysis and is costly and difficult to operationalize in transport; for example, CLIPbased ROI identification incurs 27.13% additional inference cost [32, 41]. Our key observation is that multimodal inputs already have distinct and well-defined roles in model interaction. As shown in Table 2, control events, audio, video, and text differ in urgency, reliability, replaceability, and adaptability. For example, control events and speech directly affect turn-taking, visual frames provide replaceable grounding evidence, and text often carries exact symbols that require reliable delivery. Omni-models should therefore expose explicit modalityspecific transport requirements, including minimum fidelity, tunable operating ranges, latency sensitivity, reliability, and cross-modal alignment constraints. The transport layer can then satisfy the minimum requirements needed for accurate inference while avoiding unnecessary transmission, and independently map each modality to suitable channels or network paths rather than applying a uniform policy to all inputs. The resulting objective is an accuracy-bounded performance optimization: deliver sufficient model-useful evidence within its validity window while minimizing transmission overhead and end-to-end delay.

Bad 2.5 160 20 5040.9

Network-Conditioned Model Behavior

In traditional media systems, a poor network condition mainly appears as receiver-side quality degradation, such as stalls and lower resolution [16, 22, 26, 28, 35]. In full-duplex AI media, however, it can affect what content is produced in the first place, because delayed input changes the model state from which subsequent outputs are generated. Take user barge-in as an example, the signal must enter the model’s inference loop so that the model can update its interaction state and decide to stop speaking and listen to the new input. As shown in Figure 4, after the model starts answering the first question, the user interrupts with a new visual question. Under a good network condition, the model receives the interruption in time and switches to the new responses, while under network degradation, the delayed barge-in causes it to continue the obsolete response. Table 1 shows this case quantitatively in MiniCPM-o-4.5 Demo [9] with voice activity detection (VAD [36]) setup: as network condition degrades, barge-in latency increases from 542.7 ms to 5040.9 ms. To ensure the timely processing of interaction-sensitive inputs, full-duplex model serving needs to become inherently network-aware. The inference scheduler should maintain lightweight per-session information, such as speaking state, input queuing delay, and, crucially, network conditions. Using these signals, the runtime can determine whether a chunk should be admitted immediately, preempt ongoing work, or tolerate a short delay to improve batching efficiency. In particular, network and session state reveal whether additional waiting would merely increase latency or further reduce the usefulness of ongoing generation. Therefore, full-duplex AI media should schedule model inference in a network- and session-aware manner rather than relying solely on request arrival order or GPU utilization [33, 38, 45].

Insight 2: Omni-models should expose modality-specific requirements to enable model-defined transport.

2.4

Interaction QoE and Playback Deadlines

Traditional real-time video transport optimizes continuous media streams, where receiver-side buffering primarily absorbs packet-arrival variation and maintains smooth playback [4, 6, 8, 46, 47]. AI-native real-time communication exposes a different optimization space: model outputs arrive as discrete response chunks whose timing depends on both network conditions and model execution. User-perceived quality therefore depends on how quickly a response begins and whether later chunks arrive before their playback deadlines. We characterize interaction QoE using two metrics: Time to First Response (TTFR), which measures the delay until the

Insight 1: Model inference should be per-session and network-aware to optimize end-to-end interaction quality. 3

Ziqian Liu, Minghao Li, and Yiming Qiu

Model-Centric Spec①

Sender-side Cross-layer Transport

Inference Latency ②

Model

Throughput

RTT

Bandwidth

Jitter

TTFR

1 2

①,②,③,④,⑤,⑥

3 4

Network

Playback Deadline Miss

Signal Gatherer

Network-aware Model Inference

5 6

②,③,④,⑦,⑧

7 8 Receiver-side Playback Schedule

9 10

④,⑦,⑧

11

Figure 5: C ONFLUX Overview.

12 13 14

first useful output reaches the user, as shown in Figure 3, and Deadline Miss (DM), which captures the accumulated delay time of chunks that miss their playback deadlines. After the first chunk arrives, the receiver waits for a buffer interval before starting playback. Let 𝑡 0 denote the arrival time of the first output chunk 0, 𝑏 0 the initial playback buffer, and 𝑑𝑖 the media duration of chunk 𝑖. Playback begins at time 𝑡 0 + 𝑏 0 , which anchors the response timeline. Accordingly, the playback deadline of chunk 1 is 𝑡 0 + 𝑏 0 + 𝑑 0 . A smaller 𝑏 0 reduces TTFR, whereas a larger 𝑏 0 provides more slack for subsequent arrivals and reduces DM. Unlike the jitter buffer of a continuous stream, 𝑏 0 can be selected independently for each model response based on network conditions, model backpressure, and recent interaction latency.

15 16 17 18 19

Figure 6: Example model-centric transport specification.

which defines the input requirements, as well as runtime signals from network and inference engine that indicate whether the current input path can deliver and process chunks timely. As shown in Figure 6, the specification exposes the transport configuration available to the sender, including media channels and protocols, tunable parameters, runtime feedback signals, and the adaptation policy. The channels field declares the logical media channels, their transport protocol, and direction. The params field defines configurable knobs, such as audio bitrate, video resolution, and chunk duration, together with their valid ranges or candidate values and default settings. It can also define an adaptation order, allowing the sender to reduce less critical dimensions first, such as lowering video resolution before audio bitrate when speech freshness and intelligibility are more important than visual fidelity. The signals field combines network measurements, including RTT, jitter, loss, and available bandwidth, with server-side model backpressure indicators such as chunk serving latency, processing throughput, and TTFT. At runtime, the sender applies the configured policy to translate the specification and cross-layer observations into a joint transmission decision across modality channels. The decision determines which chunks should be degraded, deferred, or discarded based on their latency sensitivity, deadline slack, and relevance to the current interaction, allowing speech, barge-in, and control events to take precedence over less urgent visual content. The policy can range from lightweight threshold-based or delay-gradient AIMD controllers to model predictive control or learned policies, depending on the available system model and runtime data.

Insight 3: Playback should be scheduled based on network delay and model backpressure to maximize QoE.

3

Research Agenda

As illustrated in Figure 5, C ONFLUX uses model, network, and playback signals to coordinate three runtime actions. Model signals expose modality requirements and inference backpressure; network signals such as RTT, bandwidth, and jitter estimate chunk delivery conditions; and playback signals such as TTFR and DM reflect interaction quality. Guided by these signals, the sender adapts chunk size, modality channels, fidelity, and reliability; the server prioritizes chunks for immediate inference, batching, or degradation based on urgency and staleness; and the receiver adjusts the first-chunk playback buffer using end-to-end latency. Together, these coordinated actions form an end-to-end control loop that adapts transport, inference, and playback to the evolving interaction state.

3.1

channels: audio: protocol: WebSocket direction: uplink downlink params: audio_bitrate: channel: audio range: 8 48 default 24 kbps video_resolution: channel: video values: 240 360 480 720 default 480 px chunk_size: values: 0.5 1 2 default 1 s signals: server_tbt_threshold: 40 ms server_ttft_threshold: 120 ms loop: policy: delay-gradient-aimd interval: 1000 ms

Sender-Side Cross-Layer Transport

The sender is the first control point in the full-duplex AI media loop. It observes cross-layer signals before deciding how to schedule outgoing chunks and coordinate multiple modality channels. These signals include the model specification, 4

The Model in the Middle: Toward AI-Native Real-Time Communication

3.2

Table 3: Experiment network settings.

Network-Aware Model Inference

Full-duplex AI media requires network-aware inference scheduling that accounts for each session’s interaction state and end-to-end latency. The key question is not only whether delaying a chunk improves serving efficiency, but whether it changes the chunk’s meaning or usefulness. The scheduler therefore maintains lightweight per-session context, including duplex phase, turn and interaction states, recent QoE, network quality, output progress, and inference backlog. Each chunk is scheduled based on these states before admission, allowing the runtime to prioritize state-changing events while batching less urgent inputs for better throughput.

L1

L2

L3

L4

Bandwidth (Mbps) RTT (ms) Jitter (ms)

10.0 40 0

5.0 80 10

2.5 160 20

1.5 200 40

buffer reduces time to first response, whereas a larger buffer gives subsequent chunks more time to arrive before their playback deadlines. Network conditions and model backpressure determine end-to-end latency, shaping subsequent chunk arrivals relative to their playback deadlines. Therefore, at the beginning of each response, the reciever derives a congestion signal from recent playback and delivery dynamics. Based on these observed signals, the receiver adapts the playback buffer using a configurable congestion-control-inspired policy. The receiver then updates the buffer time 𝑏𝑡 at time 𝑡 as:   min {𝛾𝑏𝑡 −1, 𝑏 max } 𝐿𝑡 > 𝜙 high,     𝑏𝑡 = max {𝑏𝑡 −1 − 𝛿, 𝑏 min } 𝐿𝑡 < 𝜙 low,    𝑏𝑡 −1 otherwise, 

Box 1: State-Aware Inference Scheduling Policy. (1) Update network, QoE, and inference states. (2) For each pending chunk 𝑐 in session 𝑠, 𝑃 (𝑠, 𝑐) = 𝑃𝑟𝑖𝑜𝑟𝑖𝑡𝑦𝐹𝑢𝑛𝑐 (𝐼 (𝑠, 𝑐) + 𝑄 (𝑠) + 𝑁 (𝑠)), 𝑃: Priority, 𝐼 : Interaction state, 𝑄: QoE, 𝑁 : Network. (3) Sum over priority of current batched chunks 𝐶 and select action 𝐴 based on threshold 𝜃 . ( Í A DMIT N OW, 𝑐 ∈𝐶 𝑃 (𝑠, 𝑐) ≥ 𝜃, 𝐴(𝑠) = Í WAIT BATCH, 𝑐 ∈𝐶 𝑃 (𝑠, 𝑐) < 𝜃,

where 𝐿𝑡 denotes the observed end-to-end latency, 𝛾 > 1 is the multiplicative increase factor, 𝛿 > 0 is the additive decrease step, and 𝜙 high > 𝜙 low define a hysteresis region in which the buffer remains unchanged. When 𝐿𝑡 exceeds 𝜙 high , the receiver rapidly increases the buffer to absorb larger inter-chunk arrival gaps. When 𝐿𝑡 falls below 𝜙 low , it gradually decreases the buffer to improve responsiveness. Values between 𝜙 low and 𝜙 high do not trigger an adjustment, preventing small latency fluctuations from causing oscillation. For different models and deployment environments, the input signal can incorporate additional metrics, such as network jitter, to enable faster and more sensitive adaptation to short-term delivery changes.

(4) Commit the batch when it is full, its wait limit expires, or any queued chunk becomes urgent. At each scheduling epoch, the scheduler refreshes active session states and assigns each pending chunk a priority. Box 1 summarizes how the scheduler converts this per-session state into priority-based inference actions. The priority score captures whether the chunk may change the model state, such as a turn transition, barge-in event, or control signal. QoE ensures sessions that have recently experienced poor interaction quality, such as high end-to-end latency, delayed responses, stalls, or failed interruptions, will be given high priority, while network degradation captures unstable transport conditions that make additional waiting more harmful. The scheduler aggregates the priorities of queued chunks within each candidate batch and uses the resulting batch priority to determine when inference should begin. If the aggregate priority exceeds threshold, the batch is admitted immediately, allowing urgent interaction events and sessions with poor QoE or degraded network conditions to shorten the batching delay. Otherwise, the batch waits to improve serving efficiency.

3.3

Metric

4

Preliminary Validation

Experiment settings. We conducted a preliminary experiment using MiniCPM-o-4.5 [9], an open-source full-duplex omni-modal model, deployed on an AWS instance with one NVIDIA L40S GPU (48 GB VRAM). The client runs on another cloud provider, while both the client and the server are located in the US East region and communicate over the public Internet. We use the MiniCPM demo repository as baseline and emulate the four network profiles in Table 3 using Linux Netem, ranging from high-bandwidth, low-latency conditions to severely constrained links. For all profiles with nonzero jitter, we set the jitter correlation to 25%. Preliminary results and analysis. Figure 7 compares C ON FLUX with the baseline across increasingly constrained network profiles. As shown in Figure 7a, TTFR remains between

Receiver-Side Playback Scheduling

The receiver controls the last stage of the full-duplex loop by setting the playback buffer duration after the first output chunk arrives and before playback begins. This decision balances responsiveness against playback smoothness: a smaller 5

Ziqian Liu, Minghao Li, and Yiming Qiu

7500

Baseline

Conflux

L2

L3

serving conditions, and perturb individual input modalities to evaluate their impact on model accuracy, grounding quality, turn-taking, and response latency. Such a platform would enable controlled and reproducible comparisons of alternative transport, scheduling, buffering, and modality adaptation policies before deploying them in real systems. It could also help identify effective configurations across different models, workloads, and operating conditions, thereby accelerating the design of model-aware end-to-end optimizations.

Latency (ms)

6000 4500 3000 1500 0

L1

L4

Avg DM (ms/chunk)

(a) Time to first response. 125

Baseline

Conflux

L2

L3

Protocols for AI-native RTC. C ONFLUX currently schedules channels over existing protocols according to their transport characteristics, such as latency, reliability, ordering, and congestion-control behavior. However, their heavyweight abstractions are difficult to adapt at the packet level for AIspecific behavior. In contrast, real-time AI communication is naturally chunk-based, with explicit timing, priority, modality, and validity semantics. This motivates a new lightweight UDP- or QUIC-based [24] protocol that exposes such AInative delivery semantics directly to the transport layer.

100 75 50 25 0

L1

L4

(b) Per-chunk playback DM.

Figure 7: End-to-end QoE across network profiles. OmniGym User Simulator Profile Task List

Chunk-level serving engine observability. As full-duplex model serving shifts from request-level execution to continuous chunk processing, inference engines need fine-grained chunk-level observability and traceability. In particular, future systems should not only preserve the correspondence between input and output chunks, but also trace their execution across batching, modality encoding, LLM inference, and output generation, with a latency breakdown for each stage. Such visibility would enable accurate diagnosis of model backpressure and support chunk-aware scheduling across the serving, transport, and playback stages.

Network Emulator Inference Backend Telemetry

Accuracy, Network, QoE

Optimization Policy

Figure 8: An envisioned OmniGym. 1.39 and 1.88 s across all profiles, compared with 1.77–5.41 s for the baseline. The improvement is modest under L1, at 21.3%, but becomes substantially larger as network conditions deteriorate, reaching 74.4% under L4. Notably, TTFR under L4 is lower than under L3, suggesting that the current scheduling policy responds differently across network regimes and leaves room for further exploration and tuning. Figure 7b reports the average per-chunk lag beyond its playback deadline. C ONFLUX keeps this lag below 10 ms across all profiles, while the baseline increases from 46 ms under L1 to 94 ms under L4, indicating that chunks generated and delivered by C ONFLUX are more closely aligned with the intended playback timeline even under L4 constrained network conditions. Together, these results show that the design improves both initial responsiveness and sustained deadline adherence under network degradation.

5

Generalizing to other real-time AI media. Although C ON FLUX targets full-duplex omni-model interaction, its abstraction can be extended to other real-time AI media, including live-stream recommendation, real-time translation, meeting assistants, and interactive avatars. These workloads consume live multimodal streams under task-specific latency, fidelity, and alignment requirements. By expressing such requirements as model-centric specifications, the framework can combine them with network state and inference backpressure to make delivery and scheduling decisions across diverse applications.

6

Concluding Remarks

Full-duplex AI-native interaction transforms sender-receiver real-time communication into a closed-loop system in which network transport, model inference, and user playback jointly determine interaction quality. This paper argues that modelfacing chunks provide a common abstraction for coordinating these stages and presents C ONFLUX, an AI-native realtime communication framework that combines model-defined transport requirements, network- and session-aware inference scheduling, and adaptive receiver-side playback. Preliminary

Discussion and Future Work

Exploring optimization policies with OmniGym. Since the networking community currently has limited understanding of full-duplex models and their impact on networks, a promising direction is to build an OmniGym for systematically exploring and evaluating optimization mechanisms for AI-native real-time communication. As shown in Figure 8, OmniGym could simulate diverse user interactions, emulate network and 6

The Model in the Middle: Toward AI-Native Real-Time Communication

References

results demonstrate substantial QoE gains under constrained networks, highlighting the potential of jointly designing transport, serving, and playback architectures for real-time AI.

[1] Introducing gpt-live. https://openai.com/index/introducing-gpt-live/. [2] WebRTC: Real-time communication in browsers. W3C recommendation, World Wide Web Consortium, 2025. [3] Webtransport. W3C working draft, World Wide Web Consortium, 2026. [4] N. Agarwal, R. Pan, F. Y. Yan, and R. Netravali. Mowgli: Passively Learned Rate Control for Real-Time Video. In USENIX Symposium on Networked Systems Design and Implementation (NSDI ’25), Apr. 2025. [5] R. Arakawa, H. Yakura, K. Akuzawa, and S. Kubo. Ai for meeting minutes: Promises and challenges in designing human-ai collaboration on a production saas platform. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–6, 2025. [6] G. Carlucci, L. De Cicco, S. Holmer, and S. Mascolo. Analysis and design of the google congestion control for web real-time communication (webrtc). In Proceedings of the 7th International Conference on Multimedia Systems, pages 1–12, 2016. [7] X. Chen, Z. Feng, W. Zhao, J. Ding, K. Sun, and X. Zhang. From bits to tokens:{Knowledge-Driven} generative communication of multimodal data. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), pages 191–207, 2026. [8] Y. Cheng, Z. Zhang, H. Li, A. Arapin, Y. Zhang, Q. Zhang, Y. Liu, K. Du, X. Zhang, F. Y. Yan, et al. {GRACE}:{Loss-Resilient} {RealTime} video through neural codecs. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), pages 509–531, 2024. [9] J. Cui, B. Xu, C. Wang, T. Yu, W. Sun, Y. Xu, T. Wang, Z. He, W. Ma, T. Cai, et al. Minicpm-o 4.5: Towards real-time full-duplex omni-modal interaction. arXiv preprint arXiv:2604.27393, 2026. [10] D-ID. Ai agents: Create an interactive visual agent. https://www.did.com/ai-agents/. [11] A. S. Dandage, V. R. Rupnar, T. A. Pise, and A. Mulani. Real-time language translation application using tkinter. International Journal of Digital Communication and Analog Signals, 11(01), 2025. [12] Doubao. Doubao. https://www.doubao.com/, 2026. [13] J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan. A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2):230–244, 2022. [14] Duolingo. What is duolingo max? https://www.duolingo.com/help/wh at-is-duolingo-max, 2026. [15] I. Fette and A. Melnikov. The websocket protocol. Technical report, Internet Engineering Task Force. [16] S. Fouladi, J. Emmons, E. Orbay, C. Wu, R. S. Wahby, and K. Winstein. Salsify:{Low-Latency} network video through tighter integration between a video codec and a transport protocol. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18), pages 267–282, 2018. [17] M. Geleta, H. Sodoma, and H. Gamper. Spatial audio rendering for real-time speech translation in virtual meetings. International Journal of Human–Computer Interaction, pages 1–21, 2026. [18] Google. Gemini live. https://gemini.google/overview/gemini-live/. [19] HeyGen. Liveavatar: Lifelike real-time avatar api for conversational ai. https://www.liveavatar.com/. [20] M. Houtti, M. Zhou, L. Terveen, and S. Chancellor. Observe, ask, intervene: Designing ai agents for more inclusive meetings. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1–18, 2025. [21] INMO. Inmo. https://inmolens.com/, 2026.

7

Ziqian Liu, Minghao Li, and Yiming Qiu

[22] S. Khairy, G. Mittag, V. Gopal, F. Y. Yan, Z. Niu, E. Ameri, S. Inglis, M. Golestaneh, and R. Cutler. ACM MMSys 2024 Bandwidth Estimation in Real Time Communications Challenge. In ACM Multimedia Systems Conference (MMSys ’24), Apr. 2024. [23] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611–626, 2023. [24] A. Langley, A. Riddoch, A. Wilk, A. Vicente, C. Krasic, D. Zhang, F. Yang, F. Kouranov, I. Swett, J. Iyengar, et al. The quic transport protocol: Design and internet-scale deployment. In Proceedings of the conference of the ACM special interest group on data communication, pages 183–196, 2017. [25] Y. Ma, Z. Song, Y. Zhuang, J. Hao, and I. King. A survey on vision– language–action models for embodied ai. IEEE Transactions on Neural Networks and Learning Systems, 2026. [26] H. Meng, Y. Zhuang, Y. Noushirvani, X. Huang, and Z. Meng. {MAE}: More adaptive video encoder for consistent low latency in {HighQuality} {Real-Time} communication. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), pages 1417–1430, 2026. [27] J. Meng and B. Li. Sema: Semantic transport for real-time multimodal agents. arXiv preprint arXiv:2604.20940, 2026. [28] Z. Meng and M. Xu. Latency Optimization in Interactive Multimedia Streaming. Springer, 2024. [29] Microsoft. Using copilot vision with microsoft copilot. https://suppor t.microsoft.com/en-us/microsoft-copilot/using-copilot-vision-withmicrosoft-copilot, 2025. [30] OpenAI. Chatgpt. https://chatgpt.com/, 2026. [31] Qwen Team. Qwen3.5-omni technical report. arXiv preprint arXiv:2604.15804, 2026. [32] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021. [33] C. Ruan, Y. Chen, D. Tian, Y. Shi, Y. Wu, J. Li, and C. Li. Libra: Flexible request partitioning and scheduling for serving unbalanced and dynamic {LLM} workloads. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), pages 1243– 1258, 2026. [34] R. Shahmerdanova. Artificial intelligence in translation: Challenges and opportunities. Acta Globalis Humanitatis et Linguarum, 2(1):62– 70, 2025. [35] Y. Shen, R. Chen, B. Wang, J. Chen, H. Zhang, M. Wang, M. Xu, and Z. Meng. Mortise: Auto-tuning congestion control to optimize {QoE} via {Network-Aware} parameter optimization. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), pages 2303–2321, 2026. [36] J. Sohn, N. S. Kim, and W. Sung. A statistical model-based voice activity detection. IEEE signal processing letters, 6(1):1–3, 1999. [37] SoulFun. Soulfun. https://www.soulfun.ai/, 2026. [38] B. Sun, Z. Huang, H. Zhao, W. Xiao, X. Zhang, Y. Li, and W. Lin. Llumnix: Dynamic scheduling for large language model serving. In 18th USENIX symposium on operating systems design and implementation (OSDI 24), pages 173–191, 2024. [39] Tavus. Conversational video interface. https://www.tavus.io/cvi. [40] J. Wu, Z. Ren, L. Liu, and X. Zhang. Chat with ai: The surprising turn of real-time video communication from human to ai. In Proceedings of the 24th ACM Workshop on Hot Topics in Networks, pages 402–410, 2025. [41] J. Wu, Z. Ren, J. Zhong, L. Liu, and X. Zhang. Artic: Ai-oriented real-time communication for mllm video assistant. arXiv preprint

arXiv:2602.12641, 2026. [42] xAI. Grok. https://grok.com/, 2026. [43] D. Yao, J. Zhou, C. Yang, C. Qin, H. Hou, Z. Liang, C. Wang, Y. Cao, S. Ye, S. Xie, et al. Joyai-vl-interaction: Real-time vision-language interaction intelligence. arXiv preprint arXiv:2606.14777, 2026. [44] P. Yin, J. Zhu, H. Gao, C. Zheng, Y. Huang, T. Zhou, R. Yang, W. Liu, W. Chen, C. Guo, et al. vllm-omni: Fully disaggregated serving for anyto-any multimodal models. arXiv preprint arXiv:2602.02204, 2026. [45] D. Zhang, H. Wang, Y. Liu, X. Wei, Y. Shan, R. Chen, and H. Chen. {BlitzScale}: Fast and live large model autoscaling with o (1) host caching. In 19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25), pages 275–293, 2025. [46] H. Zhang, A. Zhou, Y. Hu, C. Li, G. Wang, X. Zhang, H. Ma, L. Wu, A. Chen, and C. Wu. Loki: improving long tail performance of learningbased real-time video adaptation by fusing rule-based models. In Proceedings of the 27th Annual International Conference on Mobile Computing and Networking, pages 775–788, 2021. [47] H. Zhang, A. Zhou, J. Lu, R. Ma, Y. Hu, C. Li, X. Zhang, H. Ma, and X. Chen. Onrl: Improving mobile video telephony via online reinforcement learning. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking, pages 1–14, 2020. [48] L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, et al. Sglang: Efficient execution of structured language model programs. Advances in neural information processing systems, 37:62557–62583, 2024. [49] Zhipu AI. GLM-Realtime. https://zhipu-ef7018ed.mintlify.app/cn/gui de/models/sound-and-video/glm-realtime, 2026.

8

Record · ID 410991 · SHA-256 575d7149d1ed27f8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.