1
SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming
arXiv:2609.18075v1 [cs.MM] 16 Sep 2026
Shiqi Xu, Soung Chang Liew, Life Fellow, IEEE, Yuyang Du, Member, IEEE
Abstract—Conventional video metrics such as PSNR, SSIM, and VMAF measure visual distortion or perceptual quality, but they do not directly capture semantic preservation: whether compression retains a video’s objects, actions, and temporal narrative. This gap matters for adaptive bitrate (ABR) streaming because user experience depends not only on visual quality and playback smoothness but also on whether such semantic information is preserved. Existing Quality-of-Experience (QoE)driven bitrate-selection and resource-allocation methods primarily aim to minimize rebuffering and bitrate switching while maximizing perceptual video quality, without explicitly considering semantic preservation. To address this gap, we introduce video semantic fidelity (SF), a metric that quantifies how well a compressed video preserves the semantic content of its source. An offline multimodal large language model (MLLM) generates structured descriptions of the reference and compressed versions of the video, and a separate text-only large language model (LLM) evaluates their semantic correspondence. The resulting content-dependent SF–bitrate profiles are cached and queried by the online bitrate selector without invoking MLLMs at runtime. Evaluations on three subjective QoE benchmarks show a consistent positive association between SF and mean opinion scores (MOS). A separate human semantic-rating study directly evaluates semantic preservation and shows that SF correlates more strongly with human judgments than conventional video metrics. We then embed these profiles into a 5G MEC-assisted video-on-demand (VoD) resource-allocation framework at the base station. When wireless resources cannot support high bitrate levels for all users, the framework uses the SF–bitrate profiles to select bitrate levels jointly across users and reduce the semantic loss caused by the required bitrate reductions. NS-3 simulations with the 5G NR module show that the proposed framework achieves higher average and worst-user SF than the evaluated baselines, with a widening advantage as the wireless resources available to each user decrease. Index Terms—Video semantic fidelity (SF), adaptive bitrate (ABR) streaming, multimodal large language models (MLLMs), quality of experience (QoE), wireless resource allocation
I. I NTRODUCTION Video streaming has become a major and bandwidthintensive source of mobile data traffic in modern wireless networks. As radio resources remain limited, efficient resource allocation is critical to maintaining user viewing quality. When prioritizing and scheduling traffic, traditional wireless resource The authors are with the Department of Information Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China (e-mail: {xs024, soung, yuydu}@ie.cuhk.edu.hk). Corresponding author: Soung Chang Liew. The work was supported in part by the Hong Kong Innovation and Technology Fund (Project Number: ITS/362/24). The experimental work in this paper was conducted in the JC STEM Lab of Advanced Wireless Networks for Mission-Critical Automation and Intelligence funded by The Hong Kong Jockey Club Charities Trust.
Fig. 1. Same bitrate reduction, unequal semantic loss: high-motion content (top) can lose action and event-level cues, whereas static content (bottom) largely preserves high-level semantics.
allocators focus on optimizing Quality-of-Service (QoS) metrics, such as latency, throughput, and packet loss [1]–[3]. These metrics describe how efficiently and reliably packets traverse the network, but they do not directly capture user viewing experience. Recognizing this gap, a growing body of work has shifted toward Quality-of-Experience (QoE)-driven resource allocation. These approaches incorporate application-layer quality signals, such as resolution, rebuffering events, and perceptual quality, into mean opinion score (MOS) prediction or optimization models [4]–[8]. However, these signals focus on how a video looks, not what it conveys (i.e., the semantic). For example, Video Multi-Method Assessment Fusion (VMAF) [9] is a widely used perceptual video quality metric proposed by Netflix. VMAF and related perceptual metrics are sensitive to the level of compression. They capture visual quality rather than semantic preservation. A viewer’s understanding of the video also depends on whether objects, actions, and event context survive compression. In adaptive bitrate streaming (ABR), each video is encoded at multiple bitrate levels, and the radio resources available to a user determine which levels can be delivered. Fig. 1 shows that bitrate reductions can affect different videos in different ways. In an athletics race, a high-motion scene such as a sprint finish requires sufficient bitrate to preserve key semantic details. Under aggressive compression, subject roles, motion cues, action continuity, and event context may become difficult to recognize, making it harder for viewers to follow the competition and its key moments. In contrast, a static
2
scene, such as a landscape view with minimal visual changes, can often be delivered at lower bitrates while preserving its main objects and scene context. This asymmetry, where the same bitrate reduction can lead to very different losses of meaning across videos, calls for an allocation criterion that accounts for semantic preservation rather than visual quality alone. Multimodal large language models (MLLMs) offer a promising route for measuring such semantic effects. By jointly reasoning over visual and textual modalities, they can describe a video’s semantic content in natural language [10]. Recent work has begun using MLLMs for video quality scoring [11]–[13]. Related semantic-aware transmission studies have used semantic information to guide resource management and video streaming [14], [15]. However, in ABR, video semantics is still rarely treated as a measurable signal for bitrate selection and resource allocation. Few studies quantify how much meaning is preserved when the same video is encoded at different bitrate levels. This paper introduces semantic fidelity (SF), a metric for measuring how much meaning is retained under compression. To compute SF, an MLLM describes different bitrate versions of a video in natural language, and a separate text-only large language model (LLM) compares the resulting descriptions. The descriptions from the MLLM cover visible subjects, actions, scene context, on-screen text, and event flow. The SF score of a compressed video is computed offline by comparing its description with that of the source reference of the same video. The resulting SF scores of compressions at different levels are cached as SF–bitrate profiles, allowing online bitrate selection to use the SF scores for resource allocation without invoking MLLMs at runtime. The profiles reveal substantial content-dependent heterogeneity: some streams require higher bitrates to preserve meaning, while others remain semantically stable at lower bitrates. This heterogeneity creates an optimization opportunity that semantic-agnostic allocation cannot exploit. Building on the offline SF profiles, we study a resource allocation problem in a video-on-demand (VoD) delivery system assisted by mobile edge computing (MEC). The system model consists of two parts: offline profiling and online delivery, as shown in Fig. 2. For offline profiling, each video is encoded at multiple bitrate levels. The MLLM/LLM pipeline compares each bitrate version with the source reference and uses the resulting scores to construct an SF profile that relates semantic fidelity to bitrate. The MEC video server, co-located with the base station, caches these bitrate versions and their SF profiles. For online delivery, the base station allocates the available radio resources among competing video streams to optimize their semantic fidelity. The base station uses the cached SF profiles to select a bitrate level for each stream. The 5G radio scheduler then assigns radio resources to transmit each stream at its selected bitrate level. The main contributions of this paper are as follows: 1) We define semantic fidelity (SF), an MLLM-based metric for measuring video semantic preservation. An offline pipeline compares each compressed version with the source reference of the same video and caches the resulting scores as SF–bitrate
Offline Video Semantic Fidelity Profiling 1
2
3
4
5
Source video reference
Encoded bitrate ladder
MLLM/LLM-based frame & scene comparison
Segment-level SFbitrate profiles
Encoded ladders & SF profile cache
Online MEC-Assisted VoD Delivery Online control at the base station MEC video server & cache
1
UE clients
SF-aware multi-user bitrate selection
UE 1
UE 2
Encoded segment ladders
Cached SF profiles
2
5G RB delivery UE N
Playback buffer
Playback buffer
Playback buffer
5G NR
Fig. 2. System overview of offline semantic-fidelity profiling and online MEC-assisted VoD delivery.
profiles across bitrate levels. Validation on three independent QoE benchmarks and a human semantic-rating study indicates that SF correlates strongly with MOS in the evaluated datasets and aligns with direct semantic-preservation judgments. In essence, SF provides a valuable resource allocation signal complementary to pixel-domain quality metrics. 2) Using the cached SF–bitrate profiles, we cast multiuser video delivery as a semantic-aware resource allocation problem in a wireless downlink. Different videos exhibit markedly different SF–bitrate relations, so that the same bitrate increase can produce very different semantic gains depending on content. The allocation framework uses this content-dependent relation to prioritize additional bitrate for semantically sensitive streams. 3) We design a two-timescale 5G resource allocation framework at the base station. At the video bitrate-selection stage, a controller maps the user playback-buffer states to a resource budget. An α-fair multiple-choice knapsack (MCKP) formulation then selects one bitrate level for each user under this budget to reduce overall SF loss. At the wireless-scheduling stage, the base station assigns the available 5G time-frequency resource blocks to deliver the selected segments. 4) We validate the complete framework through NS-3/5G New Radio (NR) simulations across multiple system bandwidths and user loads. SF-aware allocation improves semantic fidelity over non-SF baselines, with a larger advantage when wireless resources are limited. The experiments further show that allocation without SF may not protect semantically sensitive streams sufficiently, a gap revealed by measuring semantic fidelity. II. MLLM-BASED S EMANTIC F IDELITY P ROFILING FOR C OMPRESSED V IDEO This section describes how we compute semantic-fidelity (SF) profiles for each source video. We first construct a set of bitrate levels for each video, with the source reference at the highest level. An LLM compares the MLLM-generated caption of each compressed version with that of the source reference. The pipeline combines these comparisons with
3
TABLE I M AIN NOTATION USED THROUGHOUT THE PAPER . Symbol v, n BRv , qv , ∆, rv,i v , ψ v,frame Cl,i l,i
wlv , Wkv CLk,i v , Tseg t, ti , Ts Di , D̂i , λ xn,j , ξ
Description
Symbol
Description
Video and user indices
i, j, k, l
Measured bitrate set, number of bitrate levels, bitrate step, and bitrate of level i MLLM-generated caption and frame-level SF component Profiling-frame and segment semantic-importance weights Segment payload and physical playback duration of one segment in seconds Slot index, starting slot of batch i, and slot duration in segment-display-time units Actual and target transmission times of batch i in segment-display-time units, and buffer-smoothing parameter Bitrate-level selection indicator and switch-penalty coefficient
v Lv , Fl,i
Section II: bitrate level, scene, segment, and frame; Section III: batch index (i) and bitrate level (j) Number of frames and frame l at bitrate level i for video v Scene-level temporal SF and final SF of profiling frame l Segment-level SF and α-fair semantic loss at bitrate level i User set, number of users, and number of resource blocks (RBs) per slot Common buffer level in segments at the start of batch i and its target Actual and estimated RB capacities for user n in batch i, measured in bits per RB
v , ψv τj,i l,i v , Lv,α Sk,i k,i
N , N, M Bi , Bth ηn,i , η̂n,i ∗ jn,i
Bitrate level selected for user n in batch i
Offline Video Semantic Fidelity Profiling 1
2
3
4
5
Multiple bitrate levels
MLLM-Based Semantic Fidelity Profiling
Final Semantic Fidelity
Per-segment aggregation
Segment-level cached profiles
A. Frame-level semantic fidelity
Fixed-duration segments All source and encoded frames
MLLMgenerated captions
Text-only LLM judge
Framelevel SF
Framelevel SF
Scenelevel temporal SF
Source reference
B. Scene segmentation, importance weighting, and temporal SF
MEAN
Geometric mean Reference scene summary
⋮
Encoded versions
Scene-level temporal SF
Content-dependent SF–bitrate curves Cached profile fields
Semantic importance weight
Shared scene partition Encoded scene summary
Frame-level SF Segment SF
SF across bitrates
Semantic Fidelity scores
MEAN
Frame weights Segment importance
Segment importance
α-fair semantic loss Segment payload
Fig. 3. Detailed offline semantic-fidelity profiling pipeline. The pipeline compares MLLM-generated descriptions across bitrate versions, adds scene-level temporal and importance information, and aggregates the resulting scores into per-segment SF–bitrate profiles. The profiles are computed once offline and cached for online bitrate selection.
scene-level analysis and aggregates the results into segmentlevel profiles. Each cached profile records the SF and bit requirements associated with a segment and its available bitrate levels. Fig. 3 summarizes the pipeline.
We formalize the frame-level SF computation pipeline as follows: For each video v, let rv,max be the average bitrate of the source reference. Let 0 < ∆ < rv,max be the bitrate quantization step. The quantized target bitrates for the compressed versions, together with the source-reference bitrate, are
A. Frame-Level Semantic Fidelity across Bitrate Levels We define the frame-level component of SF by comparing each compressed frame with its corresponding sourcereference frame through their MLLM-generated captions. Specifically, VideoLLaMA3-7B [16] generates a structured caption for each frame, covering visible subjects, actions, scene context, and on-screen text. Qwen3-8B [17], a separate text-only model, then serves as the semantic-similarity judge for the compressed and reference captions.
BRtv = { ∆, 2∆, . . . , (qv − 1)∆, rv,max },
lr
v,max
m
, (1) The final entry is the source reference. By construction, the final interval satisfies 0 < rv,max − (qv − 1)∆ ≤ ∆. The first qv − 1 target bitrates in BRtv are used by the H.264/AVC encoder [18] to generate lower-rate versions of video v. The encoder output may differ slightly from the target bitrate, so the actual bitrate levels are given by the measured qv =
∆
4
bitrates: BRv = { rv,1 , rv,2 , . . . , rv,qv −1 , rv,qv },
(2)
where rv,i is the measured bitrate of the i-th lower-rate encoded version for i = 1, . . . , qv − 1, and rv,qv = rv,max is the bitrate of the source reference. The source reference and the qv − 1 compressed encodings form qv bitrate-specific versions. For each video v of duration Tv seconds, all bitrate versions contain Lv profiling frames. Frames with the same index l correspond to the same temporal position across all versions of video v. For each version i, the frames form the following sequence: v v Fv,i = [F1,i , F2,i , . . . , FLvv ,i ],
∈ RLv ×Hv ×Wv ×3 ,
i = 1, . . . , qv .
(3)
PLv v,frame versus meaFig. 4. Average frame-level SF component L1 l=1 ψl,i v sured bitrate rv,i for the compressed encodings with rv,i ≤ 1 Mbps. Scores plateau beyond 1 Mbps, with diminishing returns toward the maximum.
Here, Hv and Wv are the frame height and width, respectively, v and Fl,i is profiling frame l of version i. The captioning function ϕcaption by the MLLM maps each frame to a text description1 : v v Cl,i = ϕcaption (Fl,i ). (4) The LLM similarity judge ϕsimilarity compares each v compressed-frame caption Cl,i with the source-reference capv tion Cl,qv . It returns an integer score on a 1–9 scale. We linearly rescale this score to [0, 1] to define the frame-level semantic-fidelity component: v v ϕsimilarity (Cl,q , Cl,i )−1 v , i = 1, . . . , qv − 1. (5) 8 The source reference is assigned perfect fidelity by definition: v,frame ψl,q ≡ 1. Higher values indicate greater semantic agreev ment between the compressed and source-reference captions. v,frame ψl,i =
B. Illustrative Frame-Level SF–Bitrate Behavior We evaluate three representative videos spanning a range of motion and visual complexity. Video 1 – Space Drift contains relatively stable scenes with limited motion. Video 2 – Running Race involves fast-moving athletes and rapid dynamic visuals. Video 3 – Flowing Nature features moderate motion in natural landscapes such as rivers and forests. Fig. 4 plots the average frame-level SF against the measured bitrate for the compressed encodings with rv,i ≤ 1 Mbps for all i. All three curves rise steeply at low bitrates and gradually saturate, but their rates of saturation differ. Video 3 (Flowing Nature) saturates earliest, since its natural-landscape semantics remain identifiable even under aggressive compression. Video 1 (Space Drift) follows closely: its static scenes preserve dominant objects and scene context at relatively low bitrates. Video 2 (Running Race), in contrast, requires substantially higher bitrates because its rapid dynamic visuals carry richer per-frame details that remain identifiable only at higher fidelity. Thus, the same bitrate increase does not raise SF by the same amount for different videos. Beyond these inter-video differences, frame-level semantic fidelity also fluctuates substantially within each video. We 1 The prompts and decoding settings for all MLLM and LLM functions are listed in Appendix A-A.
Fig. 5. Per-frame bitrate demand b̄vl (1-s moving average of the minimum bitrate required to reach SF ≥ 0.875) for the three representative videos.
capture this with b̄vl , defined as the average bitrate required to reach an SF score of 0.875 over the available profiling frames in a 1-second window centered at frame l: 1 X v,frame b̄vl = min {rv,i | ψm,i ≥ 0.875}. (6) rv,i ∈BRv |Bl | m∈Bl
Here, Bl contains the profiling frames in the 1-second window centered at frame l. Near the beginning or end of the video, only the available frames are used. For Fig. 5, the videos have a frame rate of 30 fps, so a full window contains 30 frames. Fig. 5 shows that each video alternates between localized high-demand bursts (most pronounced in Video 2’s running scenes) and low-demand stretches, indicating that the bitrate requirements to preserve semantics in each video are timevarying. C. Scene Segmentation and Semantic Importance Weighting The frame-level SF score defined above evaluates each frame in isolation and does not account for video semantics shared across related frames. Video semantics also unfold over time, as related frames jointly convey actions and events. We treat a scene as a contiguous temporal interval that may span multiple shots but contains semantically related content. Because scene boundaries reflect the temporal organization of the source content rather than a particular compression level, we detect them from the source reference and reuse them for
5
all compressed versions. Scenes may also differ in semantic importance, so we assign each scene a semantic-importance weight and propagate it to the scene’s frames for later segmentlevel aggregation. To obtain this scene partition, we use BaSSL [19], a selfsupervised scene-segmentation method that detects scene-level boundaries from the temporal context of neighboring shots. Applying BaSSL to the source reference produces a scene partition, Scenev = [Scenev1 , Scenev2 , . . . , ScenevJv ],
(7)
where each Scenevj is a contiguous set of frame indices in {1, . . . , Lv }. For each scene, we use VideoLLaMA3-7B as the summarization function ϕsummary to generate a scene-level caption from the scene’s source-reference frames. Given the scenelevel caption prompt (Appendix A-A, Prompt 3), the MLLM outputs a single-paragraph caption: v Gvj = ϕsummary Fl,q , (8) v l∈Scenev j
where Gvj is the scene-level caption for the j-th scene of video v. The semantic weighting function ϕweight uses Qwen3-8B to assign a positive score to each scene-level caption and normalizes the scores across scenes: Scene (W1,v , . . . , WJScene ) = ϕweight (Gv1 , . . . , GvJv ), v ,v Scene Wj,v > 0,
Jv X
Scene Wj,v = 1.
(9)
j=1 Scene Each normalized score Wj,v represents the share of the video’s overall semantic importance assigned to scene j. We distribute this scene-level weight uniformly among the scene’s profiling frames and apply a common scaling factor Lv :
wlv =
Lv W Scene , |Scenevj | j,v
∀l ∈ Scenevj .
(10)
This definition preserves the relative scene and PLv importance gives the frame weights unit mean, L1v l=1 wlv = 1, across videos of different durations. Finally, we collect the frame-level weights as w
v
v = [w1v , w2v , . . . , wL ]. v
(11)
D. Temporal Semantic Fidelity via Scene-Level Narrative Comparison v,frame The frame-level score ψl,i defined in Section II-A evaluates the semantic preservation of each frame independently. It does not capture whether the temporal progression of a scene remains recognizable after compression. Compression artifacts can make the visual cues for action progression and state transitions harder to recognize. We therefore define a complementary scene-level temporal SF and combine it with the frame-level score.
Section II-C has defined the source-reference scene-level caption Gvj . To evaluate temporal preservation under compression, we apply the same summarization function ϕsummary to the frames of scene j in each compressed version: v Gvj,i = ϕsummary Fl,i , i = 1, . . . , qv − 1, (12) l∈Scenev j
where the same scene index set Scenevj is reused across bitrate v versions while the pixel content Fl,i differs. The Qwen3-8B text-only judge ϕtemporal compares each compressed-version scene-level caption with the corresponding source-reference scene-level caption and returns a temporal-fidelity score in [0, 1]: v τj,i = ϕtemporal Gvj , Gvj,i ∈ [0, 1], i = 1, . . . , qv − 1. (13) The source reference is assigned perfect temporal fidelity by v ≡ 1. definition: τj,q v v,frame The frame-level score ψl,i measures semantic preservav tion in frame l, whereas the scene-level temporal score τj,i measures temporal preservation within scene j. These scores capture complementary aspects of semantic preservation, and the final SF should be high only when both aspects are preserved. For each frame index l ∈ Scenevj , we therefore define the final SF as their geometric mean: q v,frame v v ψl,i τj,i ∈ [0, 1], ∀l ∈ Scenevj , i = 1, . . . , qv . ψl,i ≜ (14) The geometric mean makes the final SF sensitive to a low value in either component: a high frame-level score cannot fully compensate for a low temporal score, and vice versa. It treats the two components symmetrically and returns their common value when the scores are equal. For numerical stability, we floor each component score at 10−6 before computing the geometric mean. E. Segment-Level Cached Profiles for Adaptive Streaming Adaptive streaming systems such as MPEG-DASH and HTTP Live Streaming (HLS) divide each encoded video version into segments of fixed duration Tseg [2]. We group the profiling frames by segment and aggregate each group into a cached SF profile. Let d be the number of frames in a fulllength segment. Video v contains Kv = ⌈Lv /d⌉ segments. For k = 1, . . . , Kv , the profiling-frame index set of segment k is Segvk = {l ∈ {1, . . . , Lv } | (k − 1)d < l ≤ kd} .
(15)
v v Per-segment aggregation. Let ψ vl = [ψl,1 , . . . , ψl,q ] be the v SF vector of frame l across bitrate levels. We aggregate these vectors using the importance weights wlv from Section II-C. The segment-level SF is the importance-weighted mean of these SF vectors, and the segment weight is the mean frame weight within the segment. The resulting segment-level values are P wlv ψ vl
Skv =
l∈Segv k
P l∈Segv k
wlv
∈ [0, 1]1×qv ,
X 1 Wkv = wlv > 0. v |Segk | v l∈Segk
(16)
6
Here, Skv is the importance-weighted SF vector of segment k, and Wkv is its mean semantic-importance weight. Because the segments the profiling frames, their weights satisfy PKv partition v 1 v k=1 |Segk |Wk = 1, keeping the average weight scale Lv independent of the video duration. α-fair semantic loss. We define the segment-level α-fair semantic loss [20] Lv,α k,i at bitrate level i based on the segment’s semantic importance and fidelity: v v 1−α Wk 1 − (Sk,i ) , α ̸= 1, Lv,α (17) 1−α k,i = v v −Wk log Sk,i , α = 1. v where Wkv and Sk,i are the segment’s semantic importance and SF at bitrate level i, respectively. The fairness parameter α ≥ 0 controls the emphasis on low SF: α = 0 yields Wkv (1− v Sk,i ), whereas larger values penalize low-SF segments more v strongly. For numerical stability when α ≥ 1, we replace Sk,i v −3 in (17) with max(Sk,i , ϵ), where ϵ = 10 throughout this paper. Segment payload. The cached profile also stores the payload CLk,i v , defined as the size in bits of the k-th encoded streaming segment at bitrate level i and measured offline from the corresponding video file. It is a content-side value and is independent of the wireless channel. Thus, each segment is associated with a payload vector across all bitrate levels: k,2 k,qv ]. CLkv = [CLk,1 v , CLv , . . . , CLv
(18)
III. S YSTEM M ODEL AND SF-AWARE R ESOURCE A LLOCATION A. System Architecture and Synchronized Batch Operation We consider a video-on-demand (VoD) service over a 5G network with mobile edge computing (MEC). Each encoded video is divided into segments of common playback duration Tseg seconds. A video server co-located with the base station (BS) caches each segment at multiple bitrate levels together with its pre-computed SF profile. Since the BS controls the shared downlink and can access these cached profiles, it can coordinate bitrate selection and radio-resource allocation across users. The downlink uses 5G orthogonal frequencydivision multiple access (OFDMA) and operates in transmission time slots indexed by t ∈ {0, 1, 2, . . .}. The segment playback duration Tseg is typically much longer than the duration of a slot. Similarly, the delivery of a segment may span multiple slots. We measure time in units of Tseg , so a segment has duration 1 and a slot has duration Ts . The BS serves the user set N = {1, . . . , N } with M resource blocks (RBs) per slot. Each user is associated with one user equipment (UE). The bitrate level of a segment remains the same during delivery, whereas RB assignments can be updated in every slot. Bitrate selection and RB scheduling therefore operate at different time scales. At the BS, we group the next segments of users into a batch containing N segments, one per user. This grouping allows the BS to coordinate segment-level bitrate selection across users under a shared RB budget. For i = 0, 1, 2, . . ., let ti be the slot index at which the delivery of batch i begins. At ti , the buffer controller sets
Fig. 6. Two-timescale SF-aware video delivery with batch-level bitrate selection and slot-level RB delivery.
a target RB budget for batch i based on the playback-buffer state. The bitrate-level selector then selects one bitrate level for each segment in the batch in accordance with the target RB budget. Batch i continues until all N user segments are delivered, and the next batch begins at ti+1 . Thus, each batch spans one or more slots, and batches do not overlap. Note that the actual RB resources consumed by the batch may not be exactly equal to the target RB budget because the target RB budget was computed based on an estimation of the channel conditions, which may be different from how the channel conditions unfold during the delivery. This will manifest in the buffer settling at a level different from a target buffer level. The target RB budget for the next batch will take this into account to ensure that the buffer fluctuates only minimally around a target level. In short, the system operates over two timescales. Fig. 6 summarizes the two-timescale framework. The slow time scale operation is related to setting a target RB budget (and therefore the bit rates of videos) to control the buffer level and the video quality. The fast time scale is related to assigning RBs during the actual delivery of the video. B. Slow-Timescale Buffer Control and Bitrate Selection Buffer-Based Batch Budget At t0 , we assume that all users are initialized with the same number of buffered segments. Each batch contains one segment from each user, and a completed batch delivers one segment to every user’s buffer. If playback does not stall, all users consume (playback) the same amount of buffered video during the delivery of a new segment because the elapsed transmission time is common to all users. Their buffer levels remain equal at each ti . We can therefore use a single model to describe the buffer dynamics of all users. Let Bi be the common buffer level at ti , measured in number of segments. We index the first segment delivered after initialization as segment 0. Specifically, segment i is delivered between ti and ti+1 . Thus, Bi is the buffer level immediately upon the complete delivery of segment i − 1. The buffer update should account for the video consumed while a batch is being transmitted, given as follows: Bi+1 = Bi + 1 − Di .
(19)
Between ti and ti+1 , a new segment, segment i, is added to the buffer and Di video segments are consumed (i.e., have
7
departed the buffer because of playback). The expression for the departure Di is Di = (ti+1 − ti )Ts .
(20)
where ti+1 − ti is the number of transmission slots used to deliver the new segment and Ts is the slot duration in units of a segment playback time. In other words, Di is the delivery time of segment i in unit of the duration of a segment playback. Let Bth be a target buffer level for Bi . A larger target provides a greater buffer margin against variations in transmission time. The buffer controller adjusts the target transmission time of each batch to keep the buffer level close to Bth . In general, the judicious setting of Bth is critical. If it is set too low, buffer underflow and freezing of video display may occur due to the random fluctuations of communication resources and video bit-rate requirements. Setting it unnecessarily high may cause the initial buffer build-up time to be longer (in general, when the user skips the video to another section, the buffer build-up time will be incurred). A model of the buffer dynamic is required to identify an optimal Bth . For the time being, let us assume that we already have a target Bth . Our goal is to maintain the buffer level Bi to be close to Bth by controlling the amount of time taken to deliver a segment to the buffer. The delivery time Di in (20) depends on the number of bits in the new user segments and hence the quality of the video being delivered. If the current buffer level is below Bth , we will need to reduce the video quality going forward, and if the current buffer level is above Bth , we can afford to increase the video quality. At ti , the delivery of segment i − 1 has just completed, and the buffer controller can observe the buffer level Bi . To control the buffer level at ti+1 , Bi+1 , we can focus on setting a target delivery time D̂i for segment i. Suppose that the controller sets D̂i = Bi + 1 − Bth and the actual transmission time for batch i, Di , turns out to be equal to this target. Then according to (19) and (20), Bi+1 would equal Bth , as given below: Bi+1 = Bi + 1 − Di = Bi + 1 − (Bi + 1 − Bth ) = Bth . In general, we need to bound the quantity Bi + 1 − Bth by max{Bi + 1 − Bth , 0} since the delivery time for segment i cannot be negative. Furthermore, even if Bi + 1 > Bth , trying to move the buffer level to Bth in one single epoch may be too aggressive and may cause large fluctuations in video quality. Before batch 0, no measured batch transmission time is available. We initialize D̂0 = 1, corresponding to one segment playback time. For i ≥ 1, we use the following smoothing scheme to smooth out the delivery time of successive segments to limit too drastic a change from one segment to the next: D̂i = λDi−1 + (1 − λ) max{Bi + 1 − Bth , 0},
i ≥ 1. (21)
The parameter λ ∈ (0, 1) controls the smoothing by setting the relative weights of the previous transmission time Di−1 and the buffer-based term max{Bi + 1 − Bth , 0}. A smaller λ gives more weight to the buffer-based term, so the target delivery time responds more strongly to the current buffer level and the buffer may return to Bth more quickly. However, the
RB budget and selected bitrate levels may also change more drastically between batches. A larger λ gives more weight to Di−1 and keeps the target delivery time closer to the previous transmission time, producing smoother bitrate decisions but a slower buffer response. For a given Bth , this slower response may increase the risk of buffer underflow.2 SF-Aware Bitrate Selection. At ti (i.e., at the beginning of batch i), the buffer controller first sets the target delivery time D̂i based on the current buffer level. The target delivery time also determines the number of scheduling slots and hence the target number of RBs for batch i. The bitrate-level selector then chooses one bitrate level for each user’s next segment while keeping the total estimated RB requirement within this target. A shorter target delivery time yields a smaller number of RBs and therefore favors lower bitrate levels, whereas a larger target delivery time permits higher bitrate levels. To select bitrate levels under the target RB budget, the bitrate-level selector estimates the RB requirement of each available bitrate level. The segment payload at each bitrate level is measured in bits, whereas the target RB budget is measured in number of RBs. Converting a segment payload into an RB requirement requires an estimate of the payload delivered per RB. Because channel conditions differ across users and change over time, the payload delivered per RB depends on both the user and batch. However, at ti , the bitratelevel selector must choose the bitrate levels before the payload delivered per RB for batch i is known. Before batch i is transmitted, the bitrate-level selector estimates the RB capacity for each user. The selector uses this capacity estimate to convert the segment payload at each bitrate level into an estimated RB requirement. We use one capacity value for each user in each batch and apply it to all RBs that may be assigned to that user. That is, we assume channel response for a particular user is flat across frequency as well as time during the batch delivery time, but the channel responses of different users may be different. Also the channel response of the same user may vary from batch to batch—i.e., slow block fading. Specifically, ηn,i , measured in bits per RB, is the number of payload bits that an RB can deliver to user n during batch i. Since we assume each user’s RB capacity varies slowly at the batch timescale, the capacity observed in batch i − 1 provides a useful estimate for batch i. For i ≥ 1, the selector estimates ηn,i using the capacity ηn,i−1 observed during the delivery of user n’s segment i − 1 in batch i − 1: η̂n,i = ηn,i−1 ,
n ∈ N.
For batch 0, the bitrate-level selector uses an initial estimate η̂n,0 > 0 for each user. For each user, the number of RBs required to deliver the next segment depends on the selected bitrate level. Because the target number of RBs is limited, some bitratelevel combinations require more RBs than this target permits. The bitrate-level selector assigns one bitrate level to each 2 If B + 1 ≥ B i th and Di = D̂i for every batch i, the buffer level converges to Bth for any λ ∈ (0, 1). Appendix B provides the closed-form analysis for this case.
8
user jointly according to the α-fair semantic-loss criterion, using the estimated RB capacities {η̂n,i }n∈N in the resource constraint. For batch i, let j = 1, . . . , qn index the qn bitrate levels available for user n’s next segment. At bitrate level j, Lα n,j is the cached α-fair semantic loss, and CLn,j is the corresponding segment payload in bits. Let xn,j be a binary variable that equals 1 when bitrate level j is selected for user n, and 0 otherwise. We omit the batch index i from these symbols because the problem is formulated separately for each batch. For i ≥ 1, the bitrate-level selector solves qn XX xn,j Lα min n,j {xn,j }
n∈N j=1 qn XX
+ξ
C. RB Delivery and Batch Completion ∗ xn,j j − jn,i−1
2
n∈N j=1
s.t.
qn X
each of the N users, the DP considers every integer state up to ⌊M D̂i /Ts ⌋ and at most qmax = maxn qn bitrate levels. The DP retains the minimum objective value for each state. Its complexity is O(N qmax ⌊M D̂i /Ts ⌋). If even the lowest-bitrate combination requires more than M D̂i /Ts RBs, (22) has no feasible solution. The batch must nevertheless deliver one complete segment to every user. We therefore select the lowest bitrate level for all users and continue scheduling until the batch is complete. This fallback gives the smallest estimated RB requirement among the available bitrate combinations and allows the system to proceed to the next batch.
xn,j = 1,
∀n ∈ N ,
(22)
j=1
xn,j ∈ {0, 1}, ∀n ∈ N , j = 1, . . . , qn , qn XX M D̂i CLn,j ≤ . xn,j η̂ Ts n,i j=1 n∈N
The first term in the objective minimizes the sum of the cached α-fair semantic losses Lα n,j defined in Section II-E. A larger loss indicates poorer preservation of video semantics. The second term discourages bitrate changes by imposing a 2 ∗ penalty ξ j − jn,i−1 based on the difference between the bitrate levels selected for consecutive segments. Here, ξ ≥ 0 ∗ is the switch-penalty coefficient, and jn,i−1 is the bitrate level selected for user n in batch i − 1. The squared difference 2 ∗ j − jn,i−1 equals 0 when candidate bitrate level j is the same as the previous selection, while a larger difference produces a larger penalty. This penalty is included because selecting different levels for a user in two consecutive batches changes the bitrate between consecutive segments and may cause noticeable quality fluctuations. The first constraint restricts that only one bitrate level is selected for each user. The second constraint defines xn,j as a binary variable. In the last constraint, CLn,j is measured in bits and η̂n,i in bits per RB, so ⌈CLn,j /η̂n,i ⌉ is the estimated number of RBs needed to deliver user n’s segment at level j. The target delivery time D̂i corresponds to D̂i /Ts slots. Since each slot contains M RBs, M D̂i /Ts is the target number of RBs available over the target delivery time. The last constraint therefore requires the total estimated number of RBs needed by the selected bitrate levels not to exceed this target number. The multiple-choice knapsack problem is NP-hard in general, and a direct search would have to examine every combination of users’ bitrate levels. A pseudo-polynomial dynamic programming (DP) algorithm can process the users one at a time while using the remaining number of RBs as its state. Because this state must be indexed by integers, we round each estimated RB requirement up and the target number of RBs M D̂i /Ts down. Rounding in these directions makes the resource constraint stricter: any solution that satisfies the rounded constraint also satisfies the original constraint. For
Once the bitrate levels are selected at ti , the BS begins delivering the segment payloads of batch i. For user n, the ∗ ∗ . selected bitrate level jn,i determines the payload CLn,jn,i Given M RBs per slot and the actual RB capacities ηn,i , the number of slots required to deliver all N segment payloads in batch i is: & ' ∗ 1 X CLn,jn,i . (23) ti+1 − ti = M ηn,i n∈N
The inner ceiling gives the integer number of RBs required by each user. The sum gives the total RB requirement of batch i, and the outer ceiling converts this requirement into an integer number of slots. Multiplying this slot count by the slot duration Ts gives the actual delivery time Di = (ti+1 − ti )Ts . The actual RB capacities ηn,i may differ from the estimates η̂n,i used for bitrate selection. Consequently, the actual delivery time Di may differ from its target D̂i . At ti+1 , the buffer controller uses Bi+1 and Di to set D̂i+1 , while the observed capacity ηn,i becomes the estimate η̂n,i+1 for the next bitrate decision. This completes one batch-level control cycle. IV. S EMANTIC F IDELITY E VALUATION ON Q O E B ENCHMARKS AND A H UMAN S TUDY This section evaluates how the proposed SF metric relates to perceived QoE and to human judgments of semantic preservation. Section IV-A uses three public ABR datasets to examine the relationship between SF and the overall MOS, and whether SF provides information beyond that captured by conventional video metrics. Section IV-B compares SF scores with direct human ratings of video semantic preservation. A. QoE Relevance and Metric Complementarity on Public ABR Datasets To evaluate SF, we use three public ABR datasets containing test videos and their mean opinion scores (MOS). Each test video is a complete ABR playback that may switch between different bitrate levels over time. Each MOS reflects viewers’ overall QoE for the corresponding test video. LIVE-NFLXII [21] contains 420 H.264 test videos generated from 15 original videos. Waterloo SQoE-III [22] contains 450 H.264 test videos generated from 20 original videos using 6 ABR algorithms. Waterloo SQoE-IV [23] includes 225 H.264 test
9
TABLE II SROCC AND PLCC OF SF AND CONVENTIONAL VIDEO METRICS AGAINST HUMAN MOS. Metric
SF VMAF ST-RRED MS-SSIM SSIM PSNR
LIVE-NFLX-II
SQoE-III
SQoE-IV
SROCC
PLCC
SROCC
PLCC
SROCC
PLCC
0.8096 0.8267 0.7171 0.6954 0.7036 0.6688
0.7573 0.8507 0.7188 0.7210 0.7302 0.6812
0.8104 0.5613 0.4706 0.5214 0.5246 0.4606
0.7799 0.6723 0.2553 0.5895 0.3527 0.4953
0.8389 0.7856 0.6324 0.5713 0.7149 0.7145
0.8705 0.8553 0.4485 0.4293 0.7478 0.7015
Note: Because ST-RRED [24] uses the opposite scoring direction, we reverse its correlation signs for comparison with the other metrics.
videos generated from 5 original videos using 45 ABR– network combinations. Each test video is associated with an overall MOS in the dataset, and we also compute a corresponding SF score for the same test video. Let v index a test video, and let Lv be the number of frames in test video v. Following the procedure defined in Section II, we compute the SF ψlv and semanticimportance weight wlv for each frame l in test video v. The SF score of test video v is PLv v v l=1 wl ψl , (24) SFv = P Lv v l=1 wl QoE Relevance. To compare SF with MOS, we use the Spearman rank-order correlation coefficient (SROCC) and the Pearson linear correlation coefficient (PLCC). SROCC indicates whether test videos with higher SF generally receive higher MOS, whereas PLCC indicates whether SF and MOS increase together at a consistent rate. Table II reports the resulting SROCC and PLCC values for SF and the conventional video metrics across the three datasets. Across all three datasets, SF maintains SROCC above 0.80 and PLCC above 0.75, indicating a consistently strong association with MOS. Unlike SF, VMAF, SSIM, and PSNR do not maintain consistently strong correlations with MOS across all three datasets. We also repeat the SF evaluation using two additional MLLMs: Qwen2-VL-7B and InternVL2-8B. Both MLLMs show a similar overall pattern as VideoLLaMA3-7B, with SF consistently associated with MOS across all three datasets (Appendix A-C). Comparison with Video Metrics. Fig. 7 shows the pairwise SROCC between SF, conventional video metrics, and MOS on LIVE-NFLX-II. Correlations among the metrics indicate how similarly they score the test videos, whereas correlations with MOS indicate how closely each metric follows viewers’ overall ratings. As shown in Fig. 7, the five conventional video metrics have pairwise SROCC values of at least 0.87, showing that, to a large extent, they score videos similarly. SF has lower SROCC values of 0.66–0.72 with these metrics but retains the second-highest SROCC with MOS. This pattern indicates that SF does not simply reproduce the scores of conventional video metrics. VMAF achieves a higher SROCC of 0.83 on this dataset, but the strong correlation between SF and MOS suggests that semantic preservation is also relevant to viewers’ overall QoE and warrants greater attention in video-streaming research.
Fig. 7. Pairwise SROCC matrix on LIVE-NFLX-II.
B. Human Evaluation of Semantic Preservation Section IV-A shows a positive association between SF and overall MOS. In the MOS datasets, viewers rated their overall viewing experience rather than semantic preservation specifically – i.e., they were not explicitly asked to focus on semantic preservation when the dataset was collected. Our main result was that SF, being a semantic preservation metric, is positively associated with quality of experience as measured by MOS. This implies that human viewers may have subconsciously and indirectly factored semantic preservation into their quality of experience. In this section, we examine whether SF agrees with direct human judgments of semantic preservation. Specifically, unlike the MOS dataset in Section IV-A, for the dataset here, participants were explicitly asked to rate how well each compressed video preserved the semantic content of its source reference video. We selected 30 source-reference videos from LIVE-NFLXII and SQoE-III and encoded each video at four bitrates (200 kbps, 500 kbps, 1.5 Mbps, and 4 Mbps), producing 120 compressed videos. Each compressed video was paired with its source reference for evaluation. Participants viewed each compressed video and its source reference side by side, with their left–right positions randomized. Participants then rated how well each compressed video preserved the semantic content of its source reference on a 1–9 scale, with higher scores indicating better semantic preservation. We averaged the ratings for each compressed video. Fig. 8 shows the SROCC and PLCC between each evaluated metric and the average human rating. SF has the highest SROCC and PLCC among all evaluated metrics. Among the four conventional video metrics, VMAF has the highest correlations. SF exceeds VMAF by 0.17 in SROCC and 0.11 in PLCC. The higher SROCC and PLCC values indicate that SF is more closely associated with human ratings of semantic preservation than the conventional video metrics. In the human study, SF correlates more strongly with human ratings of semantic preservation than any of the evaluated conventional video metrics. SF also maintains a consistent association with MOS across the three datasets. Taken together,
10
Correlation with human ratings
SROCC
PLCC
TABLE IV BASELINE COMPARISON USING SF AT N = 6 AND 10 MH Z . VALUES ARE
1.0 0.8
AVERAGES OVER TEN SIMULATION RUNS WITH DIFFERENT RANDOM SEEDS . 0.803 0.776
0.6
0.638
0.671 0.605
0.640
0.598
0.626 0.547
0.602
0.4 0.2 0.0
SF
VMAF
MS-SSIM
PSNR
SSIM
Fig. 8. SROCC and PLCC between each metric and human ratings of semantic preservation.
these findings indicate that semantic preservation is a relevant component of streaming QoE but is not fully captured by conventional video metrics. We therefore argue that semantic preservation should be considered a distinct component of streaming QoE, alongside visual quality and playback factors. V. S YSTEM -L EVEL E VALUATION A. Simulation Setup We implement the SF-aware framework described in Section III using the 5G NR module in NS-3. SF-Aware uses the buffer controller in Section III to adjust the target RB budget for each batch and keep the common buffer level close to Bth . The simulations consider a single-cell downlink in which one BS serves all UEs. We vary the number of UEs and system bandwidth across the experiments. Each UE receives a different video so that the evaluation covers different contentdependent SF–bitrate profiles. For each video, we construct the available bitrate levels and cached segment-level SF–bitrate profiles as described in Section II. Table III summarizes the simulation parameters. B. Comparison with Baselines We compare SF-Aware with six baselines in three groups: 1) conventional ABR algorithms, 2) equal resource sharing and 3) variants that keep the same bitrate-level selector but replace SF with other scoring signals. TABLE III K EY SIMULATION PARAMETERS . Parameter
Value
NS-3 version Carrier frequency Numerology Default scenario BS transmit power System bandwidths Segment duration Tseg Bitrate step ∆ Target buffer level Bth Buffer-smoothing parameter λ Fairness parameter α Switch-penalty coefficient ξ
3.39 + NR module 4 GHz 1 (30 kHz subcarrier spacing) 3GPP UMi Street Canyon 43 dBm 15 / 12 / 10 MHz 2.0 s 500 kbps 5 segments 0.5 1 0.4
Scheme
Avg. SF
Worst-User SF
Sw./min
Stalls/ UE
Jain’s Index
SF-Aware ES BOLA [25] RobustMPC [26] VMAF-Aware [9] P.1203-Aware [28] CLIP-Aware [29]
0.931 0.859 0.856 0.869 0.907 0.893 0.853
0.908 0.699 0.683 0.705 0.787 0.790 0.680
4.00 20.00 24.50 18.25 5.50 4.75 4.00
0.00 0.00 0.00 0.33 0.00 0.00 0.00
0.9997 0.9895 0.9881 0.9905 0.9963 0.9945 0.9830
Note: Avg. SF is the mean of S̄n across UEs, whereas WorstUser minimum S̄n among UEs. Jain’s index is J = P SF is2 the P N 2 ( N n=1 S̄n ) /(N n=1 S̄n ). Switching frequency is the average number of bitrate-level changes per UE per minute of playback. Stalls/UE is the number of playback stalls caused by buffer depletion, averaged over UEs.
The two conventional ABR baselines, BOLA [25] and RobustMPC [26], do not use the buffer controller proposed in Section III. Instead, each video stream has a separate playback buffer located at its UE, and the corresponding ABR algorithm selects the bitrate of the next segment. For both baselines, the BS uses the QoS-aware scheduler in the NS-3 5G NR module [27] to allocate RBs for the requested segments. BOLA independently selects the bitrate of each requested segment at each UE based on the local buffer level and the utility and payload of each available bitrate level. RobustMPC selects the next bitrate by balancing video quality, bitrate changes, and rebuffering risk over a future horizon. Together, BOLA and RobustMPC serve as representative baselines for two common classes of ABR algorithms: buffer-based methods and model predictive control (MPC)-based ABR methods, respectively. Unlike BOLA and RobustMPC, the four remaining baselines use the same BS-side buffer controller as SF-Aware to determine the total target RB budget for each batch. EqualShare Budget (ES) divides this total RB budget equally among the UEs. For each UE, ES selects the highest bitrate level whose estimated RB requirement does not exceed its assigned share. The other three baselines use the same bitratelevel selector as SF-Aware but replace the SF scores with scores from other metrics. VMAF-Aware uses perceptual video-quality scores from VMAF [9]. P.1203-Aware uses scores from P.1203 [28], a standardized QoE model that combines video quality with playback events such as stalling and quality changes to produce an overall QoE score. CLIPAware uses CLIP similarity scores [29]. These baselines compare SF with perceptual video quality, QoE, and CLIP similarity under the same bitrate-level selector. For each UE n, let Sk be the cached SF of its delivered segment k at the selected bitrate level. Let Wk be the segment’s semantic-importance weight defined in Section II-E. We compute the overall SF S̄n of the stream delivered to UE n over the same segment range for all schemes: P k Wk Sk S̄n = P . (25) k Wk
11
SF-Aware VMAF-Aware
Cached SF for resource allocation Evaluated SF using different MLLM and LLM
(a) SF
(b) VMAF
ES CLIP-Aware
(a) Average
1.00
(c) P.1203
RobustMPC P.1203-Aware
BOLA
(b) Worst user
1.0
Cached SF
0.95
SF-Aware ES BOLA
0.8
0.85 0.7
0.80
RobustMPC
4
6
8
10
12 4
6
8
10
12
0.6
Number of UEs
VMAF-Aware P.1203-Aware
Fig. 10. SF versus the number of UEs at 15 MHz: (a) average and (b) worstuser values.
CLIP-Aware 0.80
0.9
0.90
0.85
0.90
Semantic fidelity
0.95 92
94
VMAF
95
4.30 4.36 4.42
P.1203
Fig. 9. SF, VMAF, and P.1203 for SF-Aware and six baselines at N = 6 and 10 MHz. In panel (a), the blue markers use cached SF from VideoLLaMA37B and Qwen3-8B, and the orange markers show SF evaluated using InternVideo2.5-8B and Prometheus 2 7B.
Table IV shows that SF-Aware attains the highest average and worst-user SF. Compared with the strongest non-SF result for each metric, SF-Aware increases average SF by 0.024 and worst-user SF by 0.118. All schemes using the BS-side buffer controller complete without stalls, whereas RobustMPC averages 0.33 stalls per UE. BOLA also avoids stalls, but this comes at the cost of the highest switching frequency, averaging 24.50 switches per minute. Overall, SF-Aware improves semantic fidelity, particularly for the worst user, while maintaining a low switching frequency. C. SF Evaluation of Delivered Videos Using Different Models The comparison in Table IV uses the cached SF profiles that also guide bitrate selection in SF-Aware. Since SF-Aware is designed to improve the SF scores in these profiles, evaluating its performance with the same scores may favor SF-Aware. We therefore also evaluate the delivered videos using a different MLLM and a different LLM judge. That is, the cached SF profiles used for resource allocation are obtained based on an MLLM and an LLM judge. But for evaluation, we use a different MLLM and a different LLM judge. Specifically, we use the SF pipeline in Section II to generate a new set of SF–bitrate profiles using InternVideo2.5-8B [30] and Prometheus 2 7B [31] in place of VideoLLaMA3-7B and Qwen3-8B, respectively. For each scheme, we then evaluate its SF using the new profiles according to (25), keeping its bitrate selections unchanged. We also report VMAF and P.1203 for the delivered videos to compare their perceptual-quality and QoE results. Fig. 9 reports SF from both sets of models, together with VMAF and P.1203, for all seven schemes. SF-Aware achieves the highest SF with both sets of models, showing that its advantage over the evaluated baselines is preserved when a different MLLM and LLM judge are used for evaluation. Future work could also integrate SF into a complete QoE
model or maximize SF subject to explicit constraints based on perceptual-quality metrics and a standardized streaming-QoE model. D. Load and Bandwidth Effects When the available wireless resources cannot support high bitrate levels for all UEs, the bitrate-selection rule determines which streams receive higher or lower bitrate levels. The selected bitrate levels affect how much semantic content is preserved in each UE’s video. We examine tighter resource constraints in two ways: by increasing the number of UEs at a fixed bandwidth and by reducing the system bandwidth for a fixed number of UEs. Fig. 10 examines the effect of user load by varying user N from 4 to 12 at a fixed bandwidth of 15 MHz. At N = 4, both average and worst-user SF reach the maximum value of 1.0 for every scheme. The schemes cannot be distinguished in this low-load setting because the resource constraint does not force a reduction in SF. As N increases, both average and worst-user SF decrease. SF-Aware achieves the highest values for N = 6–12, and its advantage is most pronounced for worst-user SF at larger N . The SF–bitrate profiles capture how the semantic fidelity of each video changes across bitrate levels, while the α-fair semantic loss places greater emphasis on avoiding low-SF outcomes. Together, these components help protect users whose semantic fidelity would otherwise degrade most as resource competition increases. Fig. 11 shows the effect of system bandwidth when the number of UEs is fixed at N = 6. As the bandwidth decreases from 15 to 10 MHz, both average and worst-user SF decrease overall. SF-Aware achieves the highest values at every bandwidth, and its margin over the strongest baseline widens as the bandwidth decreases. The larger gains at lower bandwidths indicate that SF provides more useful guidance for bitrate selection when wireless resources are limited. VI. D ISCUSSION AND C ONCLUSION This paper investigated how video semantic preservation can be measured and used when limited wireless resources force bitrate tradeoffs among video streams. We introduced video semantic fidelity (SF), a metric for quantifying the semantic content retained by a compressed video relative to
12
SF-Aware VMAF-Aware
RobustMPC P.1203-Aware
ES CLIP-Aware
(a) Average
BOLA
(b) Worst user
0.950
0.9
Cached SF
0.925 0.900
0.8
0.875 0.7
0.850 15
12
10
15
12
10
System bandwidth (MHz)
Fig. 11. SF versus system bandwidth at N = 6: (a) average and (b) worstuser values.
its source reference. We developed an offline MLLM/LLM pipeline to construct content-specific SF–bitrate profiles for adaptive video delivery. The resulting profiles show that the same bitrate reduction can produce different semantic losses across videos. We evaluated whether SF is positively associated with overall QoE and aligns with human judgments of semantic preservation. Across three public ABR datasets, SF shows a consistent positive association with MOS. In the human study, SF corresponds more closely to direct judgments of semantic preservation than the evaluated conventional video metrics. We therefore argue that video semantic fidelity should be considered a distinct component of streaming QoE, alongside visual quality and playback factors. We further evaluated SF as an input signal for resource allocation. We incorporated the SF–bitrate profiles into a 5G MEC-assisted VoD framework, allowing bitrate selection to account for differences in semantic loss across videos. SFAware matches the baselines in the saturated low-load setting and achieves higher average and worst-user SF as the number of users increases or the system bandwidth decreases. SFAware also retains the highest SF when the delivered videos are evaluated using a different MLLM and a different LLM judge. These results indicate that SF provides useful guidance for bitrate selection when limited wireless resources require tradeoffs among video streams. The current study has several limitations. First, the SF pipeline is based on MLLM-generated descriptions and LLM judgments, which may not capture all fine-grained visual features. Second, the evaluation focuses on the association between SF and overall QoE rather than its integration into a complete QoE model. Finally, the resource-allocation experiments consider synchronized batch operation; extending the framework to asynchronous delivery remains future work. Future research directions. The results motivate three directions for future work: 1) Larger semantic-fidelity benchmarks. A public benchmark with direct human ratings would provide a common basis for developing and comparing semantic-fidelity metrics. Such a benchmark should cover diverse content, compression levels, codecs, languages, and temporal structures.
Ratings of retained semantic content should be collected separately from visual-quality and overall-QoE ratings. 2) More accurate and efficient SF computation. Future work should improve the recognition of identities and specialized content while reducing the computational cost of offline profiling. Smaller MLLMs and selective frame processing are possible directions for reducing this cost. Image features or visual-token similarity could also be incorporated into SF computation to capture fine-grained visual details. 3) Broader integration into adaptive delivery. Future systems could combine SF with visual quality and playback factors in a complete QoE model. Evaluating the resulting framework under asynchronous delivery, rebuffering, and a wider range of network conditions would extend the current system-level evaluation. A PPENDIX A P ROMPT T EMPLATES , P ROFILING C OST, AND A DDITIONAL SF VALIDATION Sections II-A, II-C, and II-D define five model-based functions (ϕcaption , ϕsimilarity , ϕsummary , ϕweight , ϕtemporal ). Together, these functions produce the Semantic Fidelity score and the scene-level importance weights. Section A-A provides representative prompt templates for each function. Section A-C provides additional validation of the SF formulation. All prompts are invoked with greedy decoding (temperature = 0, top_p = 1). A. Prompt templates Prompt 1 — Frame caption (ϕcaption ). Input: a single decoded video frame. Output: a structured description with four labeled semantic fields followed by a one-sentence summary that condenses all present fields. You are an objective video-frame captioner. Analyze the provided frame and produce exactly five labeled lines in the order below. Omit content for any field that is genuinely absent from the frame; do not invent details. Do not mention image quality, compression, or production aspects. SUBJECT: Identify the primary subject (the most salient person, animal, or object). Include count, gender or species, distinguishing clothing or colors, and body pose. ACTION: Describe, in one sentence, the action or physical state the subject is currently performing (e.g., sprinting toward the net, standing at a podium, lying on a sofa). CONTEXT: Describe the scene setting in one sentence, covering location or environment, time of day, weather conditions, and the most salient background elements. ON-SCREEN TEXT: Transcribe verbatim any clearly readable text visible in the frame (signs, captions, overlays, scoreboards, HUD elements, jersey numbers). Write NONE if no text is present. SUMMARY: Write one self-contained sentence that integrates all the above present fields into a coherent description of the frame. Output exactly these five labeled lines and nothing else.
Prompt 2 — Caption-pair similarity (ϕsimilarity ). Input: a reference structured caption and a candidate structured caption,
13
both produced by Prompt 1 for the same frame. Output: an integer score in {1, 2, . . . , 9}. You are an objective evaluator of semantic equivalence between two structured video-frame captions. Each caption contains the labeled fields SUBJECT, ACTION, CONTEXT, ON-SCREEN TEXT, and SUMMARY, produced for the same frame under different compression levels. Score how much of the REFERENCE caption’s semantic content is preserved in the CANDIDATE on the following 1-to-9 integer scale. Evaluate each field independently, then derive a holistic score: 9 -- Identical across all fields. Same subject identity, action, context, and on-screen text (if any). 7 -- Substantially equivalent. Minor descriptive details differ in at most one field; the main subject, action, and context agree. 5 -- Partially equivalent. Subject matches but at least one of ACTION, CONTEXT, or ON-SCREEN TEXT is wrong, vague, or missing. 3 -- Largely different. Only superficial overlap across fields; a reader of the CANDIDATE would misidentify the frame’s content. 1 -- Completely unrelated. If multiple fields degrade simultaneously, choose the lower candidate score (e.g., SUBJECT correct but both ACTION and CONTEXT wrong ⇒ score 3). ON-SCREEN TEXT mismatches are weighted equally to ACTION mismatches. Output the integer score and nothing else. REFERENCE: {C_ref} CANDIDATE: {C_cand}
Prompt 3 — Scene summarization (ϕsummary ). Input: the ordered sequence of frames belonging to a single scene cluster Scenevj . Output: a one-paragraph scene-level caption. You are an objective video scene summarizer. The N frames provided are temporally ordered profiling frames from a single scene of a video. Describe the scene in one paragraph of three to five sentences. The summary should cover (i) the main subjects and their identifying attributes, (ii) the events or activities that occur within the scene, (iii) the environment or setting, and (iv) how the events or states progress within the scene. Do not mention frame numbers, image quality, compression, or production aspects. Output the summary paragraph and nothing else.
Prompt 4 — Scene-importance weighting (ϕweight ). Input: the ordered list of Jv scene-level captions (Gv1 , . . . , GvJv ). Output: Jv real numbers in (0, 1), each reflecting the corresponding scene’s relative contribution to the overall video narrative. These scores are normalized across scenes before use. You are an objective evaluator of scene-level semantic importance. You will receive J textual summaries describing the consecutive scenes of a single video, in playback order. Read all J summaries together to understand the video’s overall narrative arc, then assign each scene a strictly positive importance score in (0,1) reflecting its relative contribution to the viewer’s understanding of the video as a whole. Higher scores indicate scenes that carry distinctive subjects, pivotal events, or content that the viewer would be unlikely to recover from neighboring scenes. Lower scores indicate scenes that are repetitive, transitional, or whose content is largely redundant with adjacent scenes. Do not assign zero to any scene. The J scores need not sum to a fixed value. Output the J scores as a comma-separated list in playback order and nothing else. SCENE 1: {summary_1} SCENE 2: {summary_2} ... SCENE J: {summary_J}
Prompt 5 — Temporal-narrative judge (ϕtemporal ). Input: a
pair of scene narratives produced by ϕsummary on the same scene, the first from the source reference and the second from a compressed version. Output: a JSON object containing a temporal-fidelity score in [0, 1]. You are an objective evaluator of how well a compressed video preserves the temporal narrative of a reference video at the scene level. You see only two textual scene descriptions, not the videos themselves. Internally weigh, in the following order, (i) subject and identity persistence, (ii) action progression, (iii) event order, (iv) state transitions of subjects and objects, (v) interactions between subjects and objects, (vi) missing short but pivotal events, and (vii) hallucinated or contradictory events. If multiple aspects degrade simultaneously, choose the lower of the candidate scores. Do not output sub-scores, reasoning, or natural-language explanations. Output exactly one JSON line and nothing else, of the form {"temporal_semantic_fidelity": x}, where x is a real number in [0,1]. REFERENCE: {C_ref} CANDIDATE: {C_cand}
B. Offline profiling cost The SF pipeline runs once per source video, and the cached profiles are reused during online allocation. Table A.I reports representative per-unit computation times measured on a single NVIDIA RTX 3080 GPU (20 GB). The total profiling cost depends on the number of decoded frames, bitrate levels, and detected scenes in each source video. TABLE A.I R EPRESENTATIVE PER - UNIT COMPUTATION TIMES OF THE OFFLINE SF- PROFILING STAGES ON A SINGLE NVIDIA RTX 3080 GPU (20 GB). Stage Scene segmentation (BaSSL) Frame captioning Scene-level captioning Scene weighting Frame similarity Temporal judge
Time per unit 120 ms / video second 1 s / frame 3 s / scene and version 300 ms / video 100 ms / frame pair 160 ms / scene pair
C. Cross-MLLM stability of the Semantic Fidelity formulation To examine whether the MOS-correlation results reported in Section IV-A depend on the choice of MLLM backbone, we repeat the SF evaluation using two additional opensource MLLMs of comparable size: Qwen2-VL-7B [32] and InternVL2-8B [33]. The two alternative MLLMs use different visual encoders and language backbones from VideoLLaMA37B. In all three configurations, we keep the prompt templates and the Qwen3-8B text-only model unchanged; only the MLLM used for frame captioning and scene summarization is replaced. Table A.II reports per-dataset SROCC and PLCC against MOS on LIVE-NFLX-II, SQoE-III, and SQoE-IV under the three backbones. Across the three datasets, SF shows similar correlations with MOS under all three MLLM backbones. This consistency indicates that the MOS-correlation results are not highly sensitive to the MLLM used for frame captioning and scene summarization.
14
TABLE A.II C ROSS -MLLM SROCC AND PLCC OF S EMANTIC F IDELITY ON THREE STREAMING -Q O E DATASETS . MLLM Backbone
LIVE-NFLX-II
SQoE-III
SQoE-IV
SROCC
PLCC
SROCC
PLCC
SROCC
PLCC
0.810 0.775 0.846
0.757 0.780 0.844
0.810 0.706 0.723
0.780 0.682 0.706
0.839 0.814 0.804
0.871 0.774 0.777
VideoLLaMA3-7B† Qwen2-VL-7B InternVL2-8B †
Main backbone used in Section IV.
A PPENDIX B C LOSED -F ORM A NALYSIS OF THE B UFFER C ONTROLLER This appendix derives the buffer trajectory and settling-time estimate for the controller in (21) when Bi + 1 − Bth ≥ 0 and the playback buffer remains nonempty. When Bi +1−Bth ≥ 0, the max term equals Bi + 1 − Bth , and the controller follows the linear recurrence analyzed below. We first assume that the actual delivery time equals its target and then analyze how the buffer changes when they differ for one batch. A. Buffer Trajectory without Delivery-Time Mismatch When Bi + 1 − Bth ≥ 0, the controller in (21) becomes D̂i = λDi−1 + (1 − λ)(Bi + 1 − Bth ),
i ≥ 1.
(26)
From the buffer recurrence (19), Di = −Bi+1 + Bi + 1,
bound to 5% of the initial buffer deviation gives an estimate of the number of batches required for the buffer deviation to fall below this level: ln 0.05 . (34) i0.05 ≈ 1 + 0.5 ln λ B. Effect of Delivery-Time Mismatch When Bi + 1 − Bth ≥ 0, the actual delivery time may differ from its target because RB capacities are estimated and transmission occurs in discrete slots. Define the delivery-time mismatch as wi ≜ D̂i − Di , so that Di = D̂i − wi . The buffer-deviation recurrence becomes B̃i+1 − 2λB̃i + λB̃i−1 = wi .
Consider a unit impulse wi = δi , where δ0 = 1 and δi = 0 for i ̸= 0, under the zero initial conditions B̃0 = B̃−1 = 0. Let hi be the resulting buffer-deviation response. Solving (35) gives √ λ(i−1)/2 sin(iθ) ui−1 , θ = sin−1 1 − λ, (36) hi = √ 1−λ where ui = 1 for i ≥ 0 and ui = 0 otherwise. Since Di = 1 when the buffer remains at Bth , the delivery-time deviation is D̃i ≜ Di − 1 = −hi+1 + hi . Using √ √ sin((i + 1)θ) = λ sin(iθ) + 1 − λ cos(iθ) (37) gives
Di−1 = −Bi + Bi−1 + 1. (27) D̃i = λ
−Bi+1 + Bi + 1 = λ(−Bi + Bi−1 + 1) (28)
Rearranging yields Bi+1 − 2λBi + λBi−1 = (1 − λ)Bth .
(29)
Let B̃i = Bi − Bth be the buffer deviation. The recurrence becomes B̃i+1 − 2λB̃i + λB̃i−1 = 0. (30) Its characteristic equation is z 2 − 2λz + λ = 0, with roots p √ (31) z1,2 = λ ± j λ(1 − λ) = λ e±jθ , √ √ where j2 = −1, cos θ = λ, and θ = sin−1 1 − λ. Hence, B̃i = λi/2 [c1 cos(iθ) + c2 sin(iθ)] .
"r i/2
When Di = D̂i , substituting (27) into (26) gives + (1 − λ)(Bi + 1 − Bth ).
(35)
(32)
With the initialization D̂0 = 1 and D0 = D̂0 , we have B1 = B0 and B̃1 = B̃0 . Applyingp these initial conditions to (32) gives c1 = B0 −Bth and c2 = (1 − λ)/λ (B0 −Bth ). For i = 0, 1, 2, . . ., the buffer trajectory is therefore Bi = Bth + (B0 − Bth )λi/2 " # r (33) 1−λ × cos(iθ) + sin(iθ) . λ √ Because both characteristic roots have magnitude λ < 1, Bi converges to Bth for every λ ∈ (0, 1). The buffer deviation is bounded by |B0 − Bth |λ(i−1)/2 . For B0 ̸= Bth , setting this
# 1−λ sin(iθ) − cos(iθ) , λ
i ≥ 0, λ ∈ (0, 1).
(38) Because the recurrence is linear, replacing wi = δi with wi = aδi multiplies both √ the buffer and delivery-time deviations by a. The factor λ determines how quickly these deviations decay, so the decay becomes slower as λ approaches one. R EFERENCES [1] 3GPP, “Technical specification group services and system aspects; release 15,” 3rd Generation Partnership Project (3GPP), Tech. Rep., 2018. [2] T. Stockhammer, “Dynamic adaptive streaming over HTTP—standards and design principles,” in Proc. 2nd Annu. ACM Conf. Multimedia Syst., 2011. [3] R. Prasad and A. Sunny, “QoS-aware scheduling in 5G wireless base stations,” IEEE/ACM Transactions on Networking, vol. 32, no. 3, pp. 1999–2011, 2024. [4] H. D. Moura et al., “Improved video QoE in wireless networks using deep reinforcement learning,” in Proc. 19th Int. Conf. Netw. Service Manage. (CNSM). IEEE, 2023. [5] A. Ahmad et al., “Supervised-learning-based QoE prediction of video streaming in future networks: A tutorial with comparative study,” IEEE Communications Magazine, vol. 59, no. 11, pp. 88–94, 2021. [6] J. Feng et al., “QoE fairness resource allocation in digital twin-enabled wireless virtual reality systems,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 11, pp. 3355–3368, 2023. [7] H. Du et al., “Attention-aware resource allocation and QoE analysis for metaverse xURLLC services,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 7, pp. 2158–2175, 2023. [8] S. Liu, J. Xie, and X. Wang, “QoE enhancement of the industrial metaverse based on mixed reality application optimization,” Displays, vol. 79, p. 102463, 2023. [9] Z. Li, A. Aaron, I. Katsavounidis, A. Moorthy, and M. Manohara, “Toward a practical perceptual video quality metric,” Netflix Technology Blog, Jun. 2016.
15
[10] J. Wang et al., “A comprehensive review of multimodal large language models: Performance and challenges across different tasks,” arXiv preprint arXiv:2408.01319, 2024. [11] Z. Zhang, Z. Jia, H. Wu, C. Li, Z. Chen, Y. Zhou, W. Sun, X. Liu, X. Min, W. Lin, and G. Zhai, “Q-Bench-Video: Benchmark the video quality understanding of LMMs,” in Proc. IEEE/CVF Conf. Computer Vision Pattern Recognition (CVPR), 2025. [12] L. Cao, W. Sun, W. Zhang, X. Zhu, J. Jia, K. Zhang, D. Zhu, G. Zhai, and X. Min, “VQAThinker: Exploring generalizable and explainable video quality assessment via reinforcement learning,” in Proc. AAAI Conf. Artificial Intelligence (AAAI), 2026. [13] Y. Mi, Y. Shu, Y. Li, C. Hui, P. Zhou, and S. Liu, “CLiF-VQA: Enhancing video quality assessment by incorporating high-level semantic information related to human feelings,” in Proc. ACM Int. Conf. Multimedia (ACM MM), 2024, pp. 9989–9998. [14] L. Yan, Z. Qin, C. Li, R. Zhang, Y. Li, and X. Tao, “QoE-based semanticaware resource allocation for multi-task networks,” IEEE Transactions on Wireless Communications, vol. 23, no. 9, pp. 11 958–11 971, 2024. [15] Z. Yan, J. Pei, H. Wu, H. Tabassum, and P. Wang, “Semantic-aware adaptive video streaming using latent diffusion models for wireless networks,” IEEE Wireless Communications, vol. 32, no. 5, pp. 30–38, 2025. [16] B. Zhang et al., “VideoLLaMA 3: Frontier multimodal foundation models for image and video understanding,” arXiv preprint arXiv:2501.13106, 2025. [17] Qwen Team, “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [18] Joint Video Team, “Advanced video coding for generic audiovisual services,” ITU-T Recommendation H.264 & ISO/IEC 14496-10, Tech. Rep., 2005. [19] J. Mun, M. Shin, G. Han, S. Lee, S. Ha, J. Lee, and E.-S. Kim, “Boundary-aware self-supervised learning for video scene segmentation,” in Proc. Asian Conf. Computer Vision (ACCV), 2022, pp. 485–501. [20] J. Mo and J. Walrand, “Fair end-to-end window-based congestion control,” IEEE/ACM Trans. Netw., vol. 8, no. 5, pp. 556–567, Oct. 2000. [21] C. G. Bampis, Z. Li, I. Katsavounidis, T.-Y. Huang, C. Ekanadham, and A. C. Bovik, “Towards perceptually optimized adaptive video streaming—a realistic quality of experience database,” IEEE Trans. Image Processing, vol. 30, pp. 5182–5197, 2021. [22] Z. Duanmu, A. Rehman, and Z. Wang, “A quality-of-experience database for adaptive video streaming,” IEEE Trans. Broadcast., vol. 64, no. 2, pp. 474–487, 2018. [23] Z. Duanmu, W. Liu, Z. Li, D. Chen, Z. Wang, Y. Wang, and W. Gao, “Assessing the quality-of-experience of adaptive bitrate video streaming,” arXiv preprint arXiv:2008.08804, 2020. [24] R. Soundararajan and A. C. Bovik, “Video quality assessment by reduced reference spatio-temporal entropic differencing,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 23, no. 4, pp. 684–694, Apr. 2013. [25] K. Spiteri, R. Urgaonkar, and R. K. Sitaraman, “BOLA: Near-optimal bitrate adaptation for online videos,” IEEE/ACM Transactions on Networking, vol. 28, no. 4, pp. 1698–1711, 2020. [26] X. Yin, A. Jindal, V. Sekar, and B. Sinopoli, “A control-theoretic approach for dynamic adaptive video streaming over HTTP,” in Proc. ACM SIGCOMM, 2015, pp. 325–338. [27] K. Koutlia, S. Lagén, and B. Bojović, “Enabling QoS provisioning support for delay-critical traffic and multi-flow handling in ns-3 5GLENA,” in Proc. Workshop ns-3, 2023, pp. 45–51. [28] W. Robitza, M. N. Garcia, and A. Raake, “A modular HTTP adaptive streaming QoE model—candidate for ITU-T P.1203 (‘P.NATS’),” in Proc. QoMEX, 2017, pp. 1–6. [29] A. Radford et al., “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. (ICML). PMLR, 2021. [30] Y. Wang, X. Li, Z. Yan, Y. He, J. Yu, X. Zeng, C. Wang, C. Ma, H. Huang, J. Gao, M. Dou, K. Chen, W. Wang, Y. Qiao, Y. Wang, and L. Wang, “InternVideo2.5: Empowering video MLLMs with long and rich context modeling,” arXiv preprint arXiv:2501.12386, 2025. [31] S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo, “Prometheus 2: An open source language model specialized in evaluating other language models,” arXiv preprint arXiv:2405.01535, 2024. [32] P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y. Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin, “Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution,” arXiv preprint arXiv:2409.12191, 2024.
[33] Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, J. Ma, J. Wang, X. Dong, H. Yan, H. Guo, C. He, B. Shi, Z. Jin, C. Xu, B. Wang, X. Wei, W. Li, W. Zhang, B. Zhang, P. Cai, L. Wen, X. Yan, M. Dou, L. Lu, X. Zhu, T. Lu, D. Lin, Y. Qiao, J. Dai, and W. Wang, “How far are we to GPT-4V? closing the gap to commercial multimodal models with open-source suites,” Sci. China Inf. Sci., vol. 67, no. 12, p. 220101, Dec. 2024.