ConceptioArchivearXiv CS
arXiv CSopen access

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

K AIROS : A DATASET FOR F INE -G RAINED V IDEO L ANGUAGE M ODELING OVER S PACE , T IME , AND DYNAMICS Ruibo Ming1 , Lei Sun1 , Deheng Zhang1 , He Zhang2 , Jialu Li2 , Jian Wang3 , Zhendong Li1 , Mengshun Hu1 , Danda Pani Paudel1 , Luc van Gool1 , Jinjin Gu1 1 INSAIT, Sofia University “St. Kliment Ohridski” 2 Adobe Research 3 Snap Research The reference frame description of Shot 0:

A wide, high-angle shot captures a bustling athletics stadium during an event, with a large crowd ... The blue running track encircles a green infield, where athletes and officials are ... A large digital screen above the stands displays a blue graphic with the hashtag "#The Moment". The atmosphere is electric, underscored by the roar of the crowd and distant music, with ... Text on the track's inner edge reads "BERLIN 2018" and ... The camera remains static, offering a sweeping view of the venue's scale and the anticipation of the competition.

...

The reference frame description of Shot 52:

...

arXiv:2609.08755v1 [cs.CV] 8 Sep 2026

Long-form video sources:

Sports

Gaming

Ego-centric

Vlogs

shot 0 Arts

Media

Public safety

Embodied AI

Drones

AIGC

Computer use

Formal comms

shot 52

shot 90

shot 204

ENTITY BANK 👤 Piotr Lisek

👤 Timur Morgunov

👤 Armand Duplantis

shot 205

📦 Large digital screen

📦 Blue running track

📦 Scoreboard

19,004 videos → 5,420 hours

10 - 30 min duration (average 17 min)

00:00

05:00

10:00

15:00

20:00

25:00

BENCHMARK …

shot 91, 13:13

transition between shot 90 and shot 90, 13:00
 shot 90, 12:59

The camera offers a low-angle perspective dominated by a long, Piotr Lisek’s head is slightly more A male pole vaulter, identified as shot 91

turned toward the camera ... Piotr Lisek, stands in the A hard cut transitions from the yellow pole ..., positioned on the blue running track with white lane foreground with both fists athlete’s celebration to a low-angle markings. The background reveals the stadium's tiered stands densely shot 90, 13:01

packed with a vibrant crowd of spectators, creating an energetic clenched and raised in a triumphant shot of the pole, shifting Piotr Lisek has lowered his arms atmosphere ... On the infield, near the track, a cameraman operates a gesture, his mouth open in a yell of perspective to emphasize the large camera rig ... Ambient crowd murmuring and background celebration. He wears a white tank from a celebratory pose and turned equipment and the scale of the his head ... His mouth is now closed. venue, reinforcing the athleticism music contribute to the lively atmosphere, while the commentator's top with 'POLSKA' ... The The camera has slightly panned to voiceover states, "Great clearance there from Lisek" ... background is filled with a dense and setting of the event. the right ... crowd of spectators in the tiered shot 91, 13:15

stadium stands, blurred to The athlete's body has rotated further, with their legs now extended downward and their arms reaching … emphasize the athlete ... The forward ... The text on the blue banner in the background has changed to read "GLASGOW 2019 camera is positioned at a medium ATHLETICS CHAMPIONSHIPS" ... distance, capturing his full upper shot 90, 13:12

body and the immediate Piotr Lisek is now sitting upright on shot 91, 13:17

surroundings, conveying the raw the landing mat ... A man in a blue The athlete has completed the vault and is now descending ... The commentator’s voiceover notes, “And emotion of his successful shirt is visible to the right, holding still a bit of space between him and the bar,” indicating the clearance was successful. The camera has performance. up a white flag ...
 slightly panned to follow the athlete’s descent ...

Space

19,004

videos

What is where? 5,420

hours

Time

When does it happen?

10-30 min

duration

1 FPS

annotations

Dynamics

12

domains

How does it change?

35

categories

Question:

Just after the celebratory shot of Piotr Lisek, when the camera shows a low-angle view of the yellow pole vaulting pole on the blue track, what changes occur to the scene?

Answers:

An athlete in a white and red outfit becomes visible in mid-air above the pole, having just cleared the bar, with their arms extended upwards.

The athlete raises both arms overhead in a celebratory gesture, and the standings board is updated with new performance records for other athletes.

The camera pans slightly to the left, revealing more infield equipment and personnel, and the scoreboard updates to show new results for Piotr Lisek.

A lower-third graphic appears, displaying the athlete's name and performance statistics along with the European Championships logo.

Reasoning:

The correct answer is from [00:13:14] because a blurred figure of an athlete in a white and red outfit is now visible in mid-air, having just cleared the pole vault bar, with their arms extended upward. Distractor 1 is from [00:02:25] where a lower-third graphic appears. Distractor 2 is from [00:15:30] where the camera pans left and the scoreboard updates. Distractor 3 is from [00:11:01] where the athlete raises arms in celebration and the standings board updates.

202

scenarios

820

benchmark
 videos

2,870

human-

verified MCQs

Figure 1: K AIROS represents long-form videos as structured, time-resolved annotation streams. Each video is decomposed into shots, with each shot annotated by a reference-frame description, an in-shot differential chain, and a video-level entity bank that links recurring identities across shots. At 1 FPS, the annotations capture three axes of fine-grained video understanding: Space, Time, and Dynamics. K AIROS-Bench is derived from K AIROS by converting these structured annotations into temporally grounded questions, whose answers are tied to explicit evidence spans.

A BSTRACT Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce K AIROS, a video dataset for video-language modeling with time-resolved annotations. K AIROS consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. K AIROS provides a general-purpose foundation for modeling visual experiences over time.

1

1

I NTRODUCTION

Video-language modeling (Venugopalan et al., 2015; Sun et al., 2019; Zhu & Yang, 2020; Li et al., 2020; Luo et al., 2020; Fu et al., 2021; Li et al., 2022; Xu et al., 2021; Alayrac et al., 2022; Li et al., 2023; Zhang et al., 2023; Li et al., 2024a) seeks to represent and reason over video content through language as it unfolds over both space and time. In a video, entities appear, interact, and evolve continuously, forming structured patterns across both spatial layouts and temporal progressions. Such spatiotemporal evolution gives rise to rich dynamics, including event progression, state transitions, and causal dependencies, and provides the foundation for grounded reasoning about what happens, why it happens, and what may follow. Therefore, fine-grained video-language modeling requires jointly modeling Space, Time, and Dynamics: Space specifies what is present and how it is spatially arranged, Time captures when events occur across the video timeline, and Dynamics characterizes how entities, actions, and contexts evolve and interact to form coherent video-level understanding. A fundamental barrier to achieving this goal lies in the data. Significant efforts have been made to advance video understanding, including construction of video caption datasets (Chen et al., 2024b; Farré et al., 2024), video instruction tuning and conversational video datasets (Luo et al., 2023; Zhang et al., 2024; Chen et al., 2024b; Ren et al., 2024), and recent long video benchmarks (Chen et al., 2024a; Qin et al., 2025; Mangalam et al., 2023; Wu et al., 2024; Cheng et al., 2025; Fu et al., 2025; Chandrasegaran et al., 2024; Li et al., 2024b; Wang et al., 2025a). Despite these advances, almost all prior works still fall short on all three fronts outlined above. 1. Space: Existing datasets still lack spatial granularity. Most rely on holistic video captions and place far less emphasis on fine-grained visual annotation than image datasets. As a result, their annotations typically foreground only the most salient objects, leaving many entities, attributes, and relations essential for a comprehensive account of video content unspecified. 2. Time: Existing datasets offer limited temporal resolution. Most provide annotations at the clip level, such as a single caption or a set of questions, without fine-grained temporal alignment. While such annotations capture high-level semantics, they fail to reflect how visual content, events, and scene states evolve over time. 3. Dynamics: Existing datasets, especially benchmarks, largely fail to probe fine-grained spatiotemporal grounding. This includes aspects such as pinpointing precisely when an event occurs, tracing how states change over time, and examining how observations from different moments jointly support reasoning about temporal order, interactions, and causality. As a result, they provide only limited supervision and evaluation of such dynamics, leaving models inadequately assessed in their ability to perform grounded reasoning over dynamic video structure. These limitations are especially pronounced in long-form videos (e.g., longer than 10 minutes), where richer temporal dependencies accumulate over extended durations but remain sparsely annotated. We propose K AIROS, a new dataset for fine-grained video-language modeling over Space, Time, and Dynamics. The core of K AIROS lies an automated annotation pipeline that produces fine-grained spatiotemporal descriptions at multiple levels of granularity: within each shot, it records detailed visual content; across frames within a shot, it captures temporal variation and scene evolution; and over the full video, it integrates these observations into spatially detailed and temporally resolved descriptions. A key component of this framework is a unified, text-centric representation that maintains consistency across long video contexts. K AIROS integrates entity consistency directly into frame-level annotations, allowing recurring entities and semantic attributes to be resolved over time. This design links entities across distant segments through language, while incorporating multimodal signals such as speech and environmental audio into a coherent narrative. As demonstrated in Figure 1, built on this pipeline, K AIROS provides rich, time-resolved annotations for constructing both training data and benchmarks for video-language models. Since evidence may be distributed across distant temporal segments, K AIROS supports tasks such as cross-temporal reference resolution, causal reasoning, and compositional reasoning. Its annotations can be converted into dense dynamic captions, question-answer pairs, and explicit reasoning paths grounded in temporally localized evidence. This enables questions about what happened, when an entity appeared, how entities interacted, and how a scene evolved over time, together with rationales that connect answers to the corresponding video evidence. 2

Our final dataset contains 19,004 videos with an average duration of 17 minutes, totaling 5,420 hours of annotated video across 202 diverse scenario types. We focus on videos ranging from 10 to 30 minutes in length to better reflect real-world temporal horizons while maintaining high annotation quality. This regime is particularly important because the limitations of existing datasets become even more severe in long-form videos, where richer temporal dependencies accumulate over extended durations while annotations remain sparse. From a subset of 820 videos, we further construct a benchmark of 2,870 questions, all of which are passed through sanity checks and human verification. These questions are temporally grounded to specific moments in the video and include explicit reasoning rationales. As illustrated in Figure 1, K AIROS exhibits a substantial leap compared to caption-based datasets in annotation density and information richness. Our extensive experiments on state-of-the-art video language models (OpenAI, 2026c; The Gemini Team, 2026; Intelligence, 2024; ByteDance Seed, 2026; Chen et al., 2024c; Bai et al., 2025; Xiaomi, 2025) reveal that performance degrades when the evidence spans several shots or the whole video. Besides, we highlight that fine-tuning on training data derived from K AIROS can improve the performance of video-language models on other long video benchmarks. Our contributions are threefold. 1. We introduce K AIROS, a new long-form video-language dataset designed for fine-grained modeling of Space, Time, and Dynamics. 2. We develop an automated annotation pipeline that produces multi-level spatiotemporal representations for long videos, which captures detailed and structured visual content and enabling coherent tracking of recurring entities and semantic attributes over time. 3. We construct a temporally grounded benchmark and demonstrate the utility of K AIROS for evaluating and improving video-language models.

2

T HE K AIROS DATASET

Annotation format. K AIROS represents each video as a dense, temporally unfolding annotation stream. Rather than assigning a single coarse clip-level description, it refreshes annotations at 1 FPS to capture fine-grained visual, auditory, and semantic changes as they occur. The format also preserves video structure: frames are organized into locally coherent shots, while a video-level entity matching mechanism tracks identities, attributes, and relationships as they persist, disappear, reappear, or are disambiguated across shots. Formally, each extracted frame is stored as a structured record with four fields: (1) an absolute timestamp, (2) the corresponding video frame, (3) a shot index, and (4) an audio-aware description. Together, these records provide a temporally grounded interface for fine-grained video understanding, retrieval, temporal reasoning, entity tracking, and video generation. 2.1

K AIROS A NNOTATION P IPELINE

The dense and structured nature of K AIROS annotations makes manual annotation prohibitively expensive. We therefore develop an automated pipeline that first structures the video globally and then performs fine-grained annotation. As shown in Figure 2, this design preserves both local details and long-range video dynamics under a tractable computational budget. Video structure parsing. We model cross-shot entity consistency to handle entities that reappear across distant shots, undergo visual changes, or are later identified through speech and context, thereby capturing continuity and long-range relationships across the video. We first employ TransNetV2 (Soucek & Lokoc, 2024) to segment each video into shots. After shot segmentation, we identify entities appearing in the reference frame of each shot and construct a video-level entity bank. We define eight types of entities: person, animal, object, vehicle, text, location, food, and clothing. Each detected entity is associated with a canonical label and a short visual-detail string, describing distinguishing attributes. We then link each entity mention emitted for a new reference frame to the video-level entity bank. The bank is injected into every reference-frame prompt, which requires that an entity already in the bank be referred to by its canonical name verbatim; a mention whose normalized label exactly matches a bank entry is recorded as a new appearance of that entity. This lets later shots reuse established identities instead of introducing 3

VLM Annotation

Shot Detection

(1 FPS) Qwen3-VL-8B-Instruct + vLLM

TransNetV2

Ref Description

Audio Processing

(Shot Initial Frame)

Speech

Differential Frames Descriptions

faster-whisper-large-v3

Video Input

(Subsquent Frames)

Environmental Sound

Inline Entity Matching

Qwen2-Audio-7B-Instruct

Information Sampling For 10 question types

Question Correct Answers

Distractors Reasoning

Organize into three formats: MCQ

Leakage Control Text-only Audit

GPT-5.4

Multiple Choice Question

Claude-Opus-4.7

OpenQA Open-ended Question Answering

Gemini-3.1-Pro

SFT Supervised Fine-Tuning Data

Gemma-4-31B-it

GPT-4o-mini

Remove Questions that at least

3 of the 5 models answer correctly.

Final Benchmark

Gemini-2.5-Flash

Data Organization

Human Check

Question Generation

Figure 2: K AIROS transforms raw long-form videos into structured, time-resolved annotations that integrate visual, temporal, entity-level, and audio information, supporting dense captioning, reasoning supervision, and K AIROS-Bench construction. duplicates. In this way, the entity bank serves both as a tracking mechanism and as global context for spatio-temporal descriptions. In-shot fine-grained annotation. Given the parsed video structure, K AIROS annotates each shot at 1 FPS with an initial-and-subsequent scheme. The first sampled frame in each shot is treated as the reference frame and receives a full moment-level description, including spatial layout, visible entities, attributes, relations, on-screen text, and relevant audio cues. Subsequent 1 FPS samples are treated as differential frames, which record only fine-grained changes relative to the previous second, such as actions, state transitions, motion, interactions, visibility changes, and newly appearing or disappearing entities. This design reduces redundancy while preserving fine-grained temporal evolution. The core visual annotation process is driven by Qwen3-VL-8B-Instruct (Bai et al., 2025) and accelerated with vLLM (Kwon et al., 2023). For each initial or differential frame, the model receives the sampled frame, the current entity bank, and the relevant context history and audio information. The generated description is structured along three axes: Space, describing what is present at a moment; Time, anchoring the description to an absolute point on the video timeline; and Dynamics, describing how the scene evolves. Audio and transition annotation. We combine speech transcription and non-speech audio summarization. faster-whisper-large-v3 (SYSTRAN, 2024; Radford et al., 2023) transcribes speech into timestamped sentence-level segments, while Qwen2-Audio-7B-Instruct (Chu et al., 2024) summarizes ambient and event-level sounds in 30-second windows. Both streams are aligned with the annotation timeline: speech is attached to the corresponding frame, while nonspeech summaries are attached only to reference frames to preserve context without redundancy. K AIROS also explicitly annotates shot-boundary transitions to preserve cross-shot continuity. For each boundary, the pipeline records the editing technique (hard cut, fade, dissolve, wipe, etc.), the narrative purpose of the cut, and the visual contrast across it. These descriptions link adjacent shots and prevent the video annotation from becoming a set of disconnected shot-level records. Annotation computation analysis. Because K AIROS annotations are produced by a fully automated pipeline, the process scales naturally with available GPU parallelism. In our production setting, one H200 annotates about 4 hours of video per hour; with tensor parallelism over two H200s, the pipeline reaches about 8× real time. This corresponds to roughly 1,355 H200-hours for the K AIROS dataset of 5,420 hours, with preprocessing and audio analysis running at over 20× real time and contributing little to the overall cost. 2.2

V IDEO C URATION

To provide a suitable setting for studying dynamic and fine-grained video understanding, we explicitly prioritize long-form narratives with extended durations and a high density of shot transitions. Such videos naturally encapsulate rich event progressions, entity interactions, and causal dependencies over time. We source our raw data from the two most prominent long-form video platforms: 4

Figure 3: Per-category video count of K AIROS dataset, grouped by the 12 parent domains (bar color). Each bar stacks the curated benchmark subset (820 videos, dark bottom), the YouTube nonbenchmark remainder (middle), and the Bilibili (top hatched). YouTube and Bilibili. Using an agent-based search strategy, we constructed an initial pool of candidate videos ranging from 10 minutes to 4 hours in length. This yielded a total of 28,282 URLs, comprising 15,054 from YouTube and 13,228 from Bilibili. To ensure semantic diversity, each candidate URL is pre-annotated and filtered using a stringent three-level taxonomy encompassing 12 high-level domains, 35 categories, and 202 leaf-level scenarios (e.g., A. Sports → A.1. Ball Games → A.1.1. Basketball). During processing, excessively long videos are segmented into consecutive clips of no more than 30 minutes to maintain annotation fidelity while preserving long-range context. Our final annotated dataset consists of 19,004 high-quality videos, of which 9,434 are from YouTube, and 9,570 are from Bilibili. The total is 5,420 hours. As illustrated in Figure 3, K AIROS exhibits substantial diversity.

3

T HE K AIROS -B ENCH

The dense, time-resolved annotations of K AIROS provide structured evidence for downstream videolanguage tasks. We use them to construct K AIROS-Bench, a benchmark for fine-grained video understanding, by guiding LLM-based question generation with a taxonomy of capabilities and temporal tiers. Each item is derived from the relevant evidence in the annotation stream and produced as a complete multiple-choice instance, including the question, answer, distractors, and rationale. The same mechanism can also be adapted to generate instruction-response pairs for downstream training. Because benchmark construction requires a higher quality standard than raw generation, we further apply a strict verification process before including questions in K AIROS-Bench. 3.1

B ENCHMARK D ESIGN

We design the K AIROS-Bench around three orthogonal axes that are often entangled in long-form video understanding: Space, Time, and Dynamics. The Space axis measures what is present at a single moment, including scenes, entities, spatial relations, on-screen text, and audio cues. The Time axis measures where the required evidence resides and how long the evidence span is, ranging from moment-level cues to whole-video context. The Dynamics axis measures how the video evolves, including within-shot changes, cross-shot continuity, temporal ordering, causal relations, counterfactual reasoning, counting, and holistic integration. To systematically probe these dimensions, we partition questions into 17 capabilities across four cognitive levels: second-level perception, intrashot evolution, cross-shot reasoning, and whole-video understanding. Each question is therefore associated with both a capability label and a temporal tier, allowing model performance to be analyzed at a finer granularity. Given a video, the generator samples evidence only from the timestamped K AIROS annotation stream, and asks Gemini-2.5-Flash (Comanici et al., 2025) to produce a complete multiplechoice question (MCQ) in a single call. This design makes question construction scalable while keeping every question grounded in the same fine-grained evidence used by the dataset. We design ten source-type samplers, each targeting a different temporal scale and semantic structure in the annotation stream: 5

• Reference-frame samplers. The ref samplers target local evidence from reference frames. ref perception samples scene, entity, and spatial information, ref ocr samples on-screen text, and ref audio samples speech or environmental sound. • Within-shot samplers. The diff samplers target short-range dynamics within a shot. diff change samples a reference frame with several subsequent differential descriptions, while diff sequence samples the full differential chain of a shot. • Cross-shot samplers. entity tracking, transition, cross shot, and long range sample entities, events, and transitions across multiple shots, from adjacent-shot changes to multiminute evidence windows. • Full-video sampler. full video samples evidence over an entire video or a long segment, supporting holistic questions that require extended temporal integration. Each sampler returns the materials needed for question generation, including the relevant annotation descriptions, their timestamps, the evidence span, neighboring reference-frame context, and a pool of candidate distractors. The evidence span determines the temporal tier of the question: T1 for moment-level evidence, T2 for evidence within 60 seconds, T3 for evidence within 300 seconds, T4 for evidence within 900 seconds, and T5 for evidence beyond 900 seconds. At the same time, the source type restricts the possible capability labels from the Space and Dynamics taxonomy. Thus, the temporal tier and capability label are not assigned post hoc after question generation; they are determined by the evidence sampling process itself. In the second step, we provide Gemini-2.5-Flash (Comanici et al., 2025) with the sampled evidence, the neighboring reference-frame context before and after the target evidence, and a coarse temporal hint indicating the approximate position of the evidence in the video. The model is required to return a structured response containing four fields: question, answer, distractors, and reasoning. The question, correct answer, three distractors, and rationale are generated atomically in the same call. This encourages internal consistency between the answer and the reasoning, and avoids a separate rewriting stage that could make the distractors superficially different from the correct answer. The distractors design. A key design choice is to construct distractors from real annotations rather than hallucinating from scratch. Specifically, the distractor pool is drawn from non-overlapping time windows of the same video. Thus, wrong options remain linguistically and semantically plausible because they describe content that actually appears in the video, but they refer to the wrong time point. This reduces text leakage and prevents models from relying on commonsense priors or eliminating obviously implausible choices. To answer correctly, a model must locate the relevant temporal evidence and understand the corresponding visual, auditory, or dynamic content. 3.2

C HECKS OF THE B ENCHMARK

Leakage control and sanity checks. A common failure mode in video benchmarks is text leakage, where a model can infer the correct answer from the question wording or option priors alone, without actually understanding the video. We address this issue at both the generation and filtering stages. At the generation stage, distractors are designed to be plausible rather than fabricated. This prevents models from answering by simply eliminating obviously implausible choices, instead forces them to locate and understand the relevant visual, auditory, and dynamic evidence. With this design, some questions may still be solvable from textual priors. We therefore apply a strict text-only audit using a committee of five strong language models: GPT-5.4 (OpenAI, 2026b), Claude-Opus-4.7 (Anthropic, 2026), Gemini-3.1-Pro (The Gemini Team, 2026), GPT-4o-mini (OpenAI, 2026a), and Gemma-4-31B-it (Google DeepMind, 2026). Each model receives only the question and the four shuffled answer options, without access to the video frames, audio, or textual video annotations. If at least three of the five models answer a question correctly, the question is marked as text-solvable and removed. Since each model is evaluated with a single shuffled option order, the random majority floor is approximately 10.4%. Across 25,707 raw questions, this text-only audit removes 19,430 items and leaves 6,277 video-dependent candidates. Human review. After the text-only audit, we further introduce a human verification stage to construct the final curated evaluation set. Each surviving MCQ is independently reviewed by at least two human annotators, who watch the corresponding video and evaluate the item according to four criteria. First, question validity checks whether the question is clear, unambiguous, and answerable given the video evidence. Second, accuracy and grounding verifies that the correct answer 6

Table 1: Using a text-only Gemini-3.1-Pro solver with identical prompts and option shuffling, K AIROS-Bench shows the lowest leakage among all benchmarks, measured by absolute lift over the random baseline, despite having the second-longest question stems. Benchmark CG-Bench (Chen et al., 2024a) DeVE-QA (Qin et al., 2025) EgoSchema (Mangalam et al., 2023) LongVideoBench (Wu et al., 2024) Video-Holmes (Cheng et al., 2025) Video-MME (Fu et al., 2025) HourVideo (Chandrasegaran et al., 2024) MVBench (Li et al., 2024b) LVBench (Wang et al., 2025a) K AIROS (Ours)

# Options

Random

Accuracy

Leakage ↓

# Avg words ↑

var 2–8, N̄ =6.86 fixed 5 fixed 5 fixed 4 fixed 6 fixed 4 fixed 5 var 2–5, N̄ =3.51 fixed 4

14.57% 20.00% 20.00% 25.00% 16.67% 25.00% 20.00% 28.45% 25.00%

50.20% 53.60% 53.60% 56.00% 47.60% 54.80% 39.00% 41.00% 37.40%

+35.63 pp +33.60 pp +33.60 pp +31.00 pp +30.93 pp +29.80 pp +19.00 pp +12.55 pp +12.40 pp

48 37 154 87 54 40 97 32 39

fixed 4

25.00%

34.60%

+9.60 pp

144

Benchmark Capability Distribution (17 capabilities, 4 axes)

Figure 4: Statistics of K AIROS-Bench. Distribution of the curated 2,870 questions across three independent labeling axes: evidence span, source type, and evaluated capability, showing the benchmark coverage over space, time, and dynamics. matches the visual or auditory facts and that the associated temporal evidence is accurate. Third, answer uniqueness ensures that none of the distractors can reasonably be interpreted as another correct answer. Fourth, video dependence confirms that the question cannot be answered easily without watching the video. Any item that fails the review criteria is discarded. For the final leaderboard set, we retain only questions with unanimous pass verdicts, yielding 2,870 human-curated MCQs over 820 videos. As shown in Table 1, we keep the text-only accuracy of K AIROS-Bench substantially lower than that of common video benchmarks, indicating that the retained questions require video-grounded evidence rather than language-only reasoning. 3.3

B ENCHMARK S TATISTICS

Following the generation, leakage-control, and human-review steps above, the released K AIROSBench contains 2,870 human-verified MCQs (and matched Open-ended QA pairs) over 820 videos drawn from all 35 categories. Figure 4 summarizes the three labeling axes: the Time axis is deliberately weighted toward the long-context tail (T3–T5 jointly account for ≈30% of items), the Source axis is dominated by reference-frame and within-shot samplers but retains meaningful mass on cross-shot and full-video samplers, and the Capability axis spreads over all 17 rows so that peraxis evaluation surfaces specific failure modes rather than a single overall score. Representative questions across these axes are shown in Figure 5. 7

...

Diff: Sequence

...

Ref: Audio

B4: Camera Movement

A5: Audio Compre

hension

Ref: OCR

A4: Te

xt Reading

shown against the sky, following the close-up of the snowboarder's boot and preceding

Just after the snowboarder mentions taking the board out in the early season, when the snowboarder in a camouflage jacket is carving down the slope, what do they say about

Just after the snowboarder finishes discussing the board's carving capabilities, in the scene where a building is visible near a ski lift, what text appears on the sign of that

the shot of them on the groomed slope, what is the camera's behavior?

the snowboard?

building?

The camera remains static, maintaining a low-angle perspective.

The snowboarder states that the board possesses an incredible snap.

The sign on the building near the ski lift displays the word

Regarding the camera movement during the sequence where the ski lift chairs are

...

Ref: Perception

...

...

...

Diff: C

A1: Scene Recognition

hange

B1: C

hange Detection

'Ruby'.

Entity Tracking

C1: Entity Continuity

Considering the opening segment of the video, which of the following best describes

flare, what is the most prominent feature of the stage background?

In the segment where George Michael is singing passionately into the microphone with warm stage lighting, what changes are visible in his head position and facial expression?

The background is shrouded in darkness and haze, revealing only faint red and blue

His head slightly lowers and tilts more toward his right shoulder, and his mouth opens

The audience is visible as blurred figures in the background, illuminated by a warm

wider to emphasize a vocal note, while the microphone is held a bit lower and closer to

orange glow from below, while a hazy atmosphere pervades the venue.

During the opening segment, shortly after George Michael is seen walking across the stage, in the scene where he is illuminated by a powerful blue spotlight with a lens

stage lights in the distance.

'

the audience s appearance in a later scene where George Michael is performing under a spotlight?

his lips.

...

Cro

ss Shot

s l Reasoning

...

Long Range

C4: Cau a

What is the primary reason the game transitions from the live soccer match where Nikki is ahead 1-0 against Guest?

C

3: Temporal Order

is the sequence of events leading up to the second match against '

Nihad' where Nikki is

coins at stake for the upcoming match, preparing for the ne

'

xt phase of gameplay.

In the scene where Maran Zucker, D.D.S. is explaining skincare routines to the camera, just after the text 'LINKS BELOW' is no longer visible on screen, what detail is prominent on her attire?

kicking off?

'

The transition occurs to display the player s chosen opponent and the amount of gold

A2: Entity Identification

Ref: Perception

Considering the initial match selection screen featuring 'Nikki' and 'Mohammed', what

The initial match selection shows Nikki against Mohammed, then there s a match with

Her black V-neck scrub top features her embroidered name,

Dagmar that ends in a tie, followed by a match selection screen before Nikki starts the

clearly visible on the left chest.

'Maran Zucker, D.D.S.',

game against Nihad.

Full Video

D2: Tempora

s

l Localization

C2:

Tran ition

When does the player 'Nikki' first encounter an opponent whose name is clearly displayed as 'Mohammed' before a match begins?

The correct answer is supported by [00:10:42] because at this

Narrative Transition

And then continues speaking about her 'first go-to after a retinoid,' what is the editing technique employed and its purpose?

timestamp, 'Nikki' (seen at 00:06:41 and 00:06:50) and before occurs in the second quarter of the video, just after a match 'Guest_212349804' has concluded and Nikki is selecting a new opponent. This

any subsequent matches. Distractor 1. Distractor 2 ... [00:12:16].

against

...

Ref: Perception

A

3: Spatial Reasoning

...

Diff: Sequence

B2: Action Sequencing

Diff: C

When the woman with auburn hair is weaving in a loose end of beige yarn, just before she begins sewing together two sections of a knitted piece, what is visible in the softly blurred background?

In the sequence where the woman is shown threading a loose end of yarn with a darning needle, immediately following the shot where she is inspecting stitches with the darning needle, what is the correct order of actions?

A large ball of beige yarn rests nearby, accompanied by a dark gray knitting needle, j

the fabric; then, her right hand holds the yarn taut as it s threaded; finally, her left hand

indicating the ongoing pro ect.

A hard cut is used to maintain a continuous flow of her skincare advice, highlighting her xplains application instructions. hand gestures as she e

Distractor 3 ...[00:13:20].

First, the woman switches the darning needle to her left hand while her right hand holds

'

hange

B

3: State Tracking

In the close-up segment of Tayo Rockson, following the wide shot of him on stage, what is the sequence of his facial expressions and head movements?

His mouth first opens mid-speech, then widens further as he emphasizes a point, and finally closes as he turns his head away from the camera.

lifts the darning needle higher.

Full Video

D

3: Counting

How many distinct times does the video show a close-up of hands weaving in loose ends of yarn into the knitted piece?

Cro

The correct answer is supported by [00:08:02, 00:10:32] because these are the only two instances where the hands are specifically

ss Shot

C5: Counter Factua

l

If Tayo Rockson had not mentioned learning a lesson at age 20 about being lost in Greece, what would have been less likely to occur immediately after his statement about gathering information?

shown weaving in loose ends. Distractor 1 is incorrect as it The video shows close-ups of hands weaving in loose ends of yarn two distinct times.

happens twice. Distractor 3 is incorrect as it happens only twice. Distractor 4 is incorrect as it happens only twice.

The camera would have been less likely to transition from a wide shot of him speaking

'

'

to a close-up of the ONE title card with the global skyline.

Figure 5: Representative questions from K AIROS-Bench across the three axes. Each block shows three keyframes, the question and correct answer, and the labels assigned to the question. The examples cover the full evidence span: from single-frame perception, through within-shot evolution and cross-shot continuity, up to full-video. 3.4

R ESULTS OF THE K AIROS -B ENCH

We evaluate 7 closed-source models and 14 open-source models. Three models receive the video natively: the two Gemini models and LLaVA-Video-7B. Every other model receives uniformly sampled frames at the per-model budget listed in Table 8. The results are shown in Table 2. Accuracy drops once the evidence spans several shots or the whole video. To further analyze text leakage, we evaluate public video-MCQ benchmarks under an identical protocol: we draw a stratified sample of 500 questions from each benchmark and answer them with Gemini-3.1-Pro (The Gemini Team, 2026) using text only, without any visual input. We define leakage as the solver accuracy minus the random-answering baseline for each benchmark. K AIROSBench exhibits the lowest leakage despite having the second-highest word count. This is notable because K AIROS-Bench intentionally avoids exposing numeric timestamps that would allow a model to directly retrieve or attend to the referenced segment. Instead, each question contains a linguistic anchor that pins the query to a specific moment in the video while still requiring the model to localize the relevant event from visual content. This design prevents models from taking a timestamp-based shortcut, yet it also makes text-only leakage a more serious concern because the questions must include richer natural-language grounding. To mitigate this risk, K AIROS-Bench undergoes a strict auditing process involving five models followed by human verification. 8

Table 2: MCQ leaderboard on the K AIROS-Bench with full per-capability accuracy. The 17 capabilities are grouped by their temporal scope. #Q is the question count per capability. Single Frame Model

Overall

Within Shot

Cross Shots

Full Video

A1

A2

A3

A4

A5

B1

B2

B3

B4

C1

C2

C3

C4

C5

D1

D2

D3

2,870

388

152

74

177

184

408

201

31

354

366

14

230

53

97

5

97

39

Closed-source Gemini-3.1-Pro (The Gemini Team, 2026) GPT-5.5 (OpenAI, 2026c) Gemini-2.5-Flash (Comanici et al., 2025) Nova-2-Lite (Intelligence, 2024) Seed-2.0-Lite (ByteDance Seed, 2026) GPT-4o (Achiam et al., 2023) GPT-4o-mini (OpenAI, 2026a)

59.0 57.4 54.1 46.8 42.6 37.9 31.0

61.1 57.1 53.4 47.7 50.8 42.8 32.7

54.6 57.4 60.5 47.4 48.0 43.4 27.6

67.6 78.3 74.3 59.5 58.1 50.0 32.4

72.3 64.5 66.1 48.6 47.5 44.1 32.2

66.8 39.9 53.8 37.0 31.5 31.0 23.9

53.7 55.6 56.4 55.9 41.7 32.4 35.0

42.8 46.9 47.3 40.3 35.8 32.8 29.9

51.6 60.9 45.2 45.2 48.4 25.8 41.9

54.8 69.5 54.2 47.2 40.4 39.5 24.9

64.5 58.2 57.1 42.9 41.8 42.4 35.2

14.3 18.2 14.3 35.7 7.1 21.4 21.4

61.3 58.8 47.0 43.9 37.4 33.5 31.3

69.8 57.1 56.6 60.4 39.6 34.0 41.5

46.4 37.3 32.0 48.5 35.0 28.9 27.8

80.0 100.0 100.0 100.0 100.0 100.0 80.0

81.4 74.0 52.6 39.2 48.5 42.3 19.6

30.8 35.3 41.0 33.3 51.3 28.2 41.0

Open-weight InternVL3-78B (Chen et al., 2024c) InternVL3.5-38B (Wang et al., 2025b) GLM-4.5V (Hong et al., 2025) Qwen3-VL-30B-A3B (Bai et al., 2025) MiMo-VL-7B (Xiaomi, 2025) Qwen3-VL-8B (Bai et al., 2025) InternVL3-8B (Zhu et al., 2025) Qwen2.5-VL-7B (Bai et al., 2023) InternVL3.5-8B (Wang et al., 2025b) Step3-VL-10B (Team, 2025) GLM-4V-9B (GLM et al., 2024) Gemma-4-31B (Google DeepMind, 2026) CogVLM2-Video-13B (Hong et al., 2024) LLaVA-Video-7B (Li et al., 2024a)

52.0 47.0 46.8 45.4 42.3 42.2 41.5 41.0 40.8 39.7 38.6 38.2 36.3 26.8

47.9 46.9 44.8 43.3 41.5 42.0 46.1 37.6 42.0 41.5 39.7 42.3 33.8 33.0

46.7 46.7 44.1 42.8 41.4 43.4 36.2 40.8 36.8 40.1 36.2 46.7 36.8 16.4

64.9 62.2 54.1 58.1 54.1 54.0 55.4 54.0 43.2 50.0 37.8 58.1 36.5 33.8

53.7 50.3 53.7 54.2 47.5 52.5 45.8 44.1 47.5 53.7 35.0 42.9 31.1 24.3

37.0 29.9 34.8 28.8 24.5 26.1 30.4 35.3 29.3 23.9 27.7 26.1 29.4 28.8

61.3 53.4 50.0 46.8 42.2 42.6 51.7 42.4 44.6 38.7 41.2 33.1 40.2 22.1

45.3 35.3 39.8 37.3 37.8 38.3 38.8 39.3 37.8 29.9 33.3 26.9 33.3 22.4

58.1 38.7 51.6 41.9 32.3 41.9 38.7 41.9 38.7 35.5 54.8 38.7 51.6 25.8

62.4 54.5 65.8 63.8 55.1 50.6 35.3 48.6 44.1 45.2 46.6 42.1 49.4 28.0

50.3 42.1 41.3 42.4 41.8 42.1 37.4 38.0 41.3 41.5 39.9 39.3 36.9 31.1

50.0 50.0 35.7 28.6 35.7 42.9 28.6 35.7 50.0 35.7 42.9 35.7 7.1 21.4

46.1 41.3 37.4 38.7 40.4 34.4 39.1 37.8 39.6 31.7 36.5 34.8 30.9 27.8

64.2 62.3 45.3 54.7 54.7 52.8 54.7 45.3 52.8 49.1 47.2 41.5 39.6 32.1

44.3 52.6 43.3 41.2 37.1 37.1 40.2 43.3 25.8 38.1 40.2 38.1 32.0 17.5

100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 80.0 100.0 100.0 100.0 100.0 80.0

47.4 48.5 47.4 41.2 33.0 39.2 36.1 36.1 33.0 39.2 25.8 41.2 22.7 21.6

51.3 51.3 30.8 30.8 38.5 35.9 38.5 33.3 43.6 43.6 28.2 30.8 28.2 33.3

#Q

Table 3: OpenQA leaderboard on the K AIROSBench. Gemini-2.5-Flash (Comanici et al., 2025) judge scores each answer against the ground truth on a 0-3 scale, and we report per-tier binarized accuracy with scores no less than 2 counted as correct. Model #Q GLM-4.5V (Hong et al., 2025) MiMo-VL-7B (Xiaomi, 2025) Qwen3-VL-30B-A3B (Bai et al., 2025) InternVL3-78B (Chen et al., 2024c) Step3-VL-10B (Team, 2025) Qwen2.5-VL-7B (Bai et al., 2023) Qwen3-VL-8B (Bai et al., 2025) InternVL3.5-38B (Wang et al., 2025b) InternVL3.5-8B (Wang et al., 2025b) InternVL3-8B (Zhu et al., 2025) GLM-4V-9B (GLM et al., 2024) LLaVA-Video-7B (Li et al., 2024a)

Overall 2,870

T1 975

T2 1,050

T3 227

T4 397

T5 221

22.8% 22.6% 22.4% 19.9% 19.7% 19.2% 19.2% 16.9% 14.8% 14.7% 10.8% 8.3%

27.2% 27.3% 27.2% 26.4% 25.9% 22.8% 25.9% 22.8% 20.8% 21.0% 16.3% 9.6%

18.8% 19.7% 19.0% 14.1% 15.7% 15.5% 12.4% 11.5% 9.6% 9.4% 6.6% 3.2%

20.3% 22.0% 16.3% 17.6% 14.2% 15.4% 14.1% 15.9% 13.2% 13.2% 8.4% 6.2%

23.7% 19.1% 21.2% 20.2% 18.6% 18.9% 21.7% 18.6% 15.1% 14.1% 8.6% 14.4%

23.5% 22.2% 26.7% 20.8% 19.0% 24.9% 22.2% 14.0% 14.5% 14.9% 13.1% 17.6%

Table 4: K AIROS training data improves Qwen2.5-VL-7B-Instruct (Bai et al., 2023) on K AIROS-Bench and three external long-video benchmarks via LoRA (Hu et al., 2022) fine-tuning. At evaluation stage, we ablate the frame budget over {16, 32, 64}. Without external training data, the fine-tuned model improves over the base across all benchmarks. Benchmark

Model base (32-f) finetuned (16-f) finetuned (32-f) finetuned (64-f)

K AIROS 40.42 47.94 47.49 45.12

LongVideoBench 56.29 58.00 60.08 59.50

LVBench 38.69 39.18 40.81 42.15

Video-MME 57.11 57.63 59.56 59.30

We additionally run an OpenQA pass on open-source models on K AIROS-Bench: the model must generate the answer rather than pick a letter. We score with a text-only LLM judge Gemini-2.5-Flash, which gets only question, reference answer, and candidate. It returns an integer 0–3 (no less than 2 counts as correct). The results are shown in Table 3. 3.5

E MPOWERING V IDEO -L ANGUAGE M ODELS WITH K AIROS

We fine-tune Qwen2.5-VL-7B-Instruct with LoRA for 1 epoch, on 232,101 SFT data (including the question, answer and reasoning) derived from the K AIROS training split. The model is trained using uniformly sampled 32 frames per video. We evaluate on K AIROS-Bench and three public long-video benchmarks (LongVideoBench, LVBench, and Video-MME). Without external training data, the fine-tuned model improves over the base across all benchmarks. The results are shown in Table 4, proving that K AIROS serves as a valid supervision target.

4

C ONCLUSION

We present K AIROS, a dataset of 19,004 long-form videos with 1 FPS temporally grounded annotations, and K AIROS-Bench, a strictly audited benchmark of 2,870 MCQ and OpenQA questions organized along three orthogonal axes (Space, Time, Dynamics). K AIROS-Bench is the cleanest of nine public video-MCQ benchmarks under an identical text-only-leakage probe. Fine-tuning a 7B open-weight VLM on 232,101 SFT data corpus derived from K AIROS dataset improves accuracy on K AIROS and on three external long-video benchmarks despite using no training data from them. We release the dataset, benchmark, and pipeline to enable evaluation and supervision of long-form video understanding at the granularity at which it actually unfolds. 9

R EFERENCES Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in neural information processing systems, 34:24206–24221, 2021. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716– 23736, 2022. Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with natural language. In Proceedings of the IEEE international conference on computer vision, pp. 5803–5812, 2017. Anthropic. Introducing Claude Opus 4.7. https://www.anthropic.com/news/ claude-opus-4-7, 2026. Accessed: 2026-05-07. Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6836–6846, 2021. Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 1728–1738, 2021. Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Icml, volume 2, pp. 4, 2021. ByteDance Seed. Seed2.0. https://seed.bytedance.com/en/seed2, 2026. Accessed: 2026-05-07. Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299–6308, 2017. Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. Hourvideo: 1-hour videolanguage understanding, 2024. URL https://arxiv.org/abs/2411.04998. Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, and Limin Wang. Cg-bench: Clue-grounded question answering benchmark for long video understanding, 2024a. URL https://arxiv.org/abs/2412.12075. Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al. Sharegpt4video: Improving video understanding and generation with better captions. Advances in Neural Information Processing Systems, 37:19472–19495, 2024b. Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In CVPR, 2024c. 10

Junhao Cheng, Yuying Ge, Teng Wang, Yixiao Ge, Jing Liao, and Ying Shan. Video-holmes: Can mllm think like holmes for complex video reasoning?, 2025. URL https://arxiv.org/ abs/2505.21374. Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. Qwen2-audio technical report. arXiv preprint arXiv:2407.10759, 2024. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. Clip2video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097, 2021. Miquel Farré, Andi Marafioti, Lewis Tunstall, Leandro Von Werra, and Thomas Wolf. Finevideo. https://huggingface.co/datasets/HuggingFaceFV/finevideo, 2024. Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6202–6211, 2019. Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 24108–24118, 2025. Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv preprint arXiv:2111.12681, 2021. Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on computer vision, pp. 5267–5275, 2017. Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng Zhang, Qinkai Zheng, Rui Lu, Shuaiqi Duan, Shudan Zhang, Shulin Cao, Shuxun Yang, Weng Lam Tam, Wenyi Zhao, Xiao Liu, Xiao Xia, Xiaohan Zhang, Xiaotao Gu, Xin Lv, Xinghan Liu, Xinyi Liu, Xinyue Yang, Xixuan Song, Xunkai Zhang, Yifan An, Yifan Xu, Yilin Niu, Yuantao Yang, Yueyan Li, Yushi Bai, Yuxiao Dong, Zehan Qi, Zhaoyu Wang, Zhen Yang, Zhengxiao Du, Zhenyu Hou, and Zihan Wang. Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024. Google DeepMind. Gemma 4 Model Card. https://ai.google.dev/gemma/docs/ core/model_card_4, 2026. Accessed: 2026-05-07. Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 18995–19012, 2022. Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, et al. Cogvlm2: Visual language models for image and video understanding. arXiv preprint arXiv:2408.16500, 2024. Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, et al. Glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning. pp. arXiv–2507, 2025. 11

Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum? id=nZeVKeeFYf9. Amazon Artificial General Intelligence. The amazon nova family of models: Technical report and model card. Amazon Technical Reports, 2024. URL https://www.amazon.science/publications/ the-amazon-nova-family-of-models-technical-report-and-model-card. Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Video question answering with spatio-temporal reasoning. International Journal of Computer Vision, 127(10):1385–1412, 2019. Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang. Egotaskqa: Understanding human tasks in egocentric videos. In The 36th Conference on Neural Information Processing Systems (NeurIPS 2022) Track on Datasets and Benchmarks, 2022. Alexander Klaser, Marcin Marszałek, and Cordelia Schmid. A spatio-temporal descriptor based on 3d-gradients. In BMVC 2008-19th British machine vision conference, pp. 275–1. British Machine Vision Association, 2008. Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision, pp. 706–715, 2017. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pp. 611–626, 2023. Ivan Laptev. On space-time interest points. International journal of computer vision, 64(2):107–123, 2005. Jie Lei, Tamara L Berg, and Mohit Bansal. Detecting moments and highlights in videos via natural language queries. Advances in Neural Information Processing Systems, 34:11846–11858, 2021. Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024a. Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi. Align and prompt: Video-and-language pre-training with entity prompts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4953–4963, 2022. Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp. 19730–19742. PMLR, 2023. Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22195–22206, 2024b. KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. Science China Information Sciences, 68(10):200102, 2025. Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. Hero: Hierarchical encoder for video+ language omni-representation pre-training. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp. 2046–2065, 2020. Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision, pp. 323–340. Springer, 2024c. 12

Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3202–3211, 2022. Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou. Univl: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020. Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. Clip4clip: An empirical study of clip for end to end video clip retrieval. arXiv preprint arXiv:2104.08860, 2021. Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning. Neurocomputing, 508: 293–304, 2022. Ruipu Luo, Ziwang Zhao, Min Yang, Zheming Yang, Minghui Qiu, Zhongyu Wei, Yanhao Wang, and Cen Chen. Valley: Video assistant with large language model enhanced ability. ACM Transactions on Multimedia Computing, Communications and Applications, 2023. Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji. X-clip: End-toend multi-grained contrastive learning for video-text retrieval. In Proceedings of the 30th ACM international conference on multimedia, pp. 638–647, 2022. Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12585–12602, 2024. Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. NeurIPS, 2023. Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 2630–2640, 2019. Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. arXiv preprint arXiv:2407.02371, 2024. OpenAI. GPT-4o mini: advancing cost-efficient intelligence. https://openai.com/index/ gpt-4o-mini-advancing-cost-efficient-intelligence/, 2026a. Accessed: 2026-05-07. OpenAI. Introducing GPT-5.4. https://openai.com/index/ introducing-gpt-5-4/, 2026b. Accessed: 2026-05-07. OpenAI. GPT-5.5 system card. https://openai.com/index/ gpt-5-5-system-card/, 2026c. Accessed: 2026-05-07. Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine Miech, Alexandre Frechette, Hanna Klimczak, Raphael Koster, Junlin Zhang, Stephanie Winkler, Yusuf Aytar, Simon Osindero, Dima Damen, Andrew Zisserman, and Joao Carreira. Perception test: A diagnostic benchmark for multimodal video models. In Advances in Neural Information Processing Systems, 2023. Hangyu Qin, Junbin Xiao, and Angela Yao. Question-answering dense video events. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, pp. 884–894. ACM, 2025. doi: 10.1145/3726302.3729945. URL http: //dx.doi.org/10.1145/3726302.3729945. 13

Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PmLR, 2021. Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pp. 28492–28518. PMLR, 2023. Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. Timechat: A time-sensitive multimodal large language model for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14313–14323, 2024. Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba, Chen Zhao, Silvio Giancola, and Bernard Ghanem. Mad: A scalable dataset for language grounding in videos from movie audio descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5026–5035, 2022. Tomás Soucek and Jakub Lokoc. Transnet v2: An effective deep network architecture for fast shot transition detection. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 11218–11221, 2024. Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 7464–7473, 2019. SYSTRAN. faster-whisper. https://github.com/SYSTRAN/faster-whisper, 2024. StepFun Team. Step-3 is large yet affordable: Model-system co-design for cost-effective decoding, 2025. URL https://arxiv.org/abs/2507.19427. The Gemini Team. Gemini 3.1 Pro: A smarter model for your most complex tasks. https://blog.google/innovation-and-ai/models-and-research/ gemini-models/gemini-3-1-pro/, February 2026. Accessed: 2026-05-07. Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are dataefficient learners for self-supervised video pre-training. Advances in neural information processing systems, 35:10078–10093, 2022. Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, pp. 4489–4497, 2015. Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. Sequence to sequence-video to text. In Proceedings of the IEEE international conference on computer vision, pp. 4534–4542, 2015. Heng Wang and Cordelia Schmid. Action recognition with improved trajectories. In Proceedings of the IEEE international conference on computer vision, pp. 3551–3558, 2013. Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu. Dense trajectories and motion boundary descriptors for action recognition. International journal of computer vision, 103(1):60– 79, 2013. Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision, pp. 20–36. Springer, 2016. Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Ming Ding, Xiaotao Gu, Shiyu Huang, Bin Xu, et al. Lvbench: An extreme long video understanding benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 22958–22967, 2025a. 14

Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. Internvl3. 5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265, 2025b. Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7794–7803, 2018. Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191, 2022. Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023. Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14668–14678, 2022. Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. STAR: A benchmark for situated reasoning in real-world videos. In Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS), 2021. Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long-context interleaved video-language understanding. Advances in Neural Information Processing Systems, 37:28828–28857, 2024. Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of questionanswering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9777–9786, 2021. LLM-Core-Team Xiaomi. Mimo: Unlocking the reasoning potential of language model – from pretraining to posttraining, 2025. URL https://arxiv.org/abs/2505.07608. Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 conference on empirical methods in natural language processing, pp. 6787–6800, 2021. Antoine Yang, Arsha Nagrani, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vidchapters-7m: Video chapters at scale. Advances in Neural Information Processing Systems, 36:49428–49444, 2023. Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. CLEVRER: collision events for video representation and reasoning. In ICLR, 2020. Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynetqa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp. 9127–9134, 2019. Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. In Proceedings of the 2023 conference on empirical methods in natural language processing: system demonstrations, pp. 543–553, 2023. Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Llava-video: Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision (ECCV), pp. 803–818, 2018. 15

Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, et al. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479, 2025. Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8746–8755, 2020.

A PPENDIX A

R ELATED W ORK

Video Understanding and Video-Language Models. Early video understanding relied on handcrafted spatiotemporal descriptors and trajectory-based representations (Laptev, 2005; Klaser et al., 2008; Wang & Schmid, 2013; Wang et al., 2013). These methods established the fundamentally spatiotemporal nature of video analysis, but they typically produced clip-level decisions rather than persistent representations of entities and states over time. Deep architectures later shifted the field toward learnable spatiotemporal features (Tran et al., 2015; Carreira & Zisserman, 2017; Feichtenhofer et al., 2019; Wang et al., 2016; Zhou et al., 2018; Wang et al., 2018). While these models substantially improved video representations, their supervision was still largely aligned with clips or full videos rather than explicitly with second-level state transitions or temporally consistent entity attributes. Video Transformers further improved long-range temporal modeling through native crossframe attention (Bertasius et al., 2021; Arnab et al., 2021; Liu et al., 2022), while self-supervised pretraining methods demonstrated strong transfer without dense annotation (Tong et al., 2022; Wei et al., 2022; Akbari et al., 2021; Wang et al., 2022). Nevertheless, much of this literature still optimizes coarse temporal objectives and often relies on sparse temporal sampling or compressed visual tokens for efficiency, which can weaken sensitivity to subtle state changes in long videos. Video-language research evolved from sequence-to-sequence captioning models such as S2VT (Venugopalan et al., 2015) to large-scale video-text pretraining frameworks (Sun et al., 2019; Zhu & Yang, 2020; Li et al., 2020; Luo et al., 2020; Fu et al., 2021; Li et al., 2022; Xu et al., 2021). Large web-scale corpora such as HowTo100M (Miech et al., 2019) and WebVid (Bain et al., 2021), together with CLIP-style contrastive learning (Radford et al., 2021), enabled more scalable crossmodal alignment and inspired retrieval-oriented extensions (Luo et al., 2021; 2022; Ma et al., 2022; Fang et al., 2021; Bain et al., 2021). However, these corpora are often only weakly aligned in time, and most training objectives emphasize global or clip-level matching rather than second-level, temporally anchored scene grounding. Recent Video-LLMs connect visual encoders to large language models through adapters, query modules, or projection layers. Connector-based architectures such as Flamingo (Alayrac et al., 2022) and BLIP-2 (Li et al., 2023) established scalable multimodal interfaces, and subsequent systems, including Video-LLaMA (Zhang et al., 2023), Video-ChatGPT (Maaz et al., 2024), LLaVA-Video (Li et al., 2024a), LLaMA-VID (Li et al., 2024c), and TimeChat (Ren et al., 2024), extended this paradigm to video dialogue, instruction following, and long-context reasoning. Even so, a central tension remains between temporal coverage and spatiotemporal fidelity: sparse frame sampling, frame pooling, and aggressive token reduction improve throughput, but they can hinder fine-grained temporal localization, entity tracking, and temporally consistent reasoning. Video-Language Datasets and Benchmarks. Large-scale video-text pretraining datasets provide the foundation for many modern video-language models. Web-scale corpora such as WebVid (Bain et al., 2021), InternVid (Wang et al., 2023), and OpenVid-1M (Nan et al., 2024) collect millions to hundreds of millions of video-text pairs, enabling scalable video-text representation learning. These datasets improve coverage and diversity, but their annotations are weakly aligned and lack dense temporal supervision. More recent high-quality captioning datasets, such as ShareGPT4Video (Chen et al., 2024b), and FineVideo (Farré et al., 2024), move toward richer and more structured supervision. Another important line of work focuses on video instruction tuning and conversational video data. Datasets such as VideoChat-11K (Li et al., 2025), Video-ChatGPT-100K (Maaz et al., 2024), Valley-Instruct-65K (Luo et al., 2023), LLaVA-Video-178K (Zhang et al., 2024), ShareGPTVideo (Chen et al., 2024b), and TimeIT (Ren et al., 2024) transform existing captions, QA annotations, or video understanding tasks into instruction-following formats. While these datasets provide richer 16

language supervision, much of their annotations are still coarse in spatial and temporal granularity. As a result, they lack fine-grained grounding in space and time and provide limited supervision over video dynamics. Temporal grounding and dense video understanding benchmarks provide more explicit temporal supervision. Classic and widely used datasets such as ActivityNet (Krishna et al., 2017), CharadesSTA (Gao et al., 2017), DiDeMo (Anne Hendricks et al., 2017), and QVHighlights (Lei et al., 2021) annotate the relationship between natural language queries and temporal segments in videos. These benchmarks have played an important role in moving video-language evaluation beyond fullvideo classification and toward timestamp-aware grounding. More recent long-form or egocentric grounding resources, such as MAD (Soldan et al., 2022) and Ego4D (Grauman et al., 2022), extend this setting to movies or first-person videos. VidChapters-7M (Yang et al., 2023) further introduces large-scale chapter-level supervision for long videos, with chapter titles and timestamps collected from web videos. These datasets are highly relevant to temporal localization, but they still do not fully solve the problem of fine-grained state and attribute tracking. Moment boundaries are often coarse, query-level, or segment-level, and the annotations usually describe events or steps rather than maintain a persistent inventory of entities, attributes, relations, and state changes across time. Video question-answering benchmarks have also evolved from short-clip QA toward long-form and diagnostic evaluation. Earlier benchmarks evaluate video understanding through manually annotated or carefully constructed questions, including both open-ended and multiple-choice formats (Yu et al., 2019; Xiao et al., 2021; Grauman et al., 2022). More recent comprehensive benchmarks, including MVBench (Li et al., 2024b), LongVideoBench (Wu et al., 2024), Video-MME (Fu et al., 2025), and LVBench (Wang et al., 2025a), expand evaluation toward multi-task reasoning, long-video understanding, and full-spectrum video modeling. These benchmarks have been crucial for revealing the limitations of Video LLMs under long-context settings, but most of them still reduce evaluation to discrete multiple-choice accuracy. This makes leaderboards easy to compare, but it can also obscure whether a model truly tracks temporal evidence, maintains entity consistency, or merely exploits language priors and coarse scene summaries. Several diagnostic and multi-task benchmarks aim to evaluate more detailed perceptual and temporal capabilities. The Perception Test includes object tracks, point tracks, action segments, sound segments, multiple-choice video QA, and grounded video QA, providing a more diverse testbed for perception, grounding, and multimodal reasoning (Patraucean et al., 2023). Datasets such as STAR, CLEVRER, EgoTaskQA, and TGIF-QA focus on situated reasoning, physical reasoning, egocentric reasoning, state transitions, and repetition counting (Wu et al., 2021; Yi et al., 2020; Jia et al., 2022; Jang et al., 2019). These datasets isolate specific reasoning factors but are typically built on short or synthetic clips and thus complement rather than replace long-video benchmarks that require sustained temporal reasoning. Overall, existing datasets and benchmarks have substantially advanced video-language learning, but they leave an important gap for fine-grained, temporally faithful video understanding. Large-scale web corpora provide breadth but only weak temporal and spatial alignment. Instruction datasets improve the conversational interface of VideoLLMs but often rely on synthetic or model-generated supervision. Temporal grounding datasets provide timestamp supervision but usually focus on queryto-moment localization rather than persistent entity and state modeling. Long-video QA benchmarks expose the difficulty of reasoning over extended contexts, but their multiple-choice format can underdiagnose perceptual failures, temporal hallucinations, and entity-state inconsistencies. These limitations suggest the need for benchmarks and training resources that combine long temporal coverage with dense temporal anchoring, explicit spatial and entity-level grounding, and evaluation protocols that test not only whether a model gives the correct answer, but also whether it can maintain temporally consistent, evidence-grounded representations throughout the video.

B

A NNOTATION P IPELINE H YPERPARAMETERS

This section enumerates the hyperparameters used by every pipeline stage so that the corpus and the benchmark can be reproduced exactly. The four artifacts emitted per video are described in Table 5 of §C; the present section concerns how those artifacts are produced. 17

B.1

S HOT D ETECTION AND F RAME S AMPLING

Shot detector. Shot boundaries are produced by T RANS N ET V2 (Soucek & Lokoc, 2024) with a softmax threshold of 0.5 and a minimum scene length of 15 source frames. The threshold is intentionally slightly recall-biased: a missed boundary merges two semantically distinct shots and breaks every cross-shot question generated for the resulting merged span, whereas a spurious boundary at most introduces a redundant reference frame inside a real shot. Both hard cuts and gradual transitions are accepted as boundaries; the gradual-transition probability head is read from the same TransNetV2 forward pass. Frame sampling. After shot detection, the pipeline materialises a per-frame structural record and extracts a 1 FPS JPEG subset for downstream annotation. Frames are resized so that the longer image dimension is 1024 pixels, and saved at JPEG quality 90 . The resulting 1 FPS timeline is the spine of the annotation stream and is what every subsequent description, audio segment, and benchmark evidence span is keyed to. Reference and differential frames. Within each shot, the first extracted frame is treated as the initial frame and receives a full spatial description plus, when entity tracking is enabled, a structured entity list. All later 1 FPS samples in the same shot are treated as differential frames and are described only in terms of what changed relative to the previous state. For long shots, the pipeline reanchors every 300 differential frames by promoting the current frame to a new initial frame, which keeps prompt context bounded and reduces drift in long description chains. B.2

AUDIO P IPELINE

Two complementary audio models. Audio is processed by two complementary models running on the same physical device. Speech is transcribed with FASTER - WHISPER - LARGE - V 3 (SYSTRAN, 2024; Radford et al., 2023) at sentence resolution. Non-speech audio is summarized by Q WEN 2-AUDIO -7B-I NSTRUCT (Chu et al., 2024) over fixed 30 -second windows; this captures environmental and event-level sounds such as crowd noise, engine sounds, applause, footsteps, and music. Alignment. The two audio streams are aligned to the same 1 FPS timeline used by the visual annotation stage. Each speech sentence is attached to the first processed frame whose timestamp falls inside the sentence’s time span, and never attached twice. Environmental audio summaries are attached only to initial frames, not to every differential frame, so that nearby records are not burdened with repeated ambient descriptions. B.3

V ISUAL A NNOTATION P ROMPTS

Server. The visual annotation stage is driven by Q WEN 3-VL-8B-I NSTRUCT served through vLLM (Kwon et al., 2023) with tensor parallel size 2 , bfloat16 weights, sampling temperature 0.2 , and gpu memory utilization 0.85 . Four prompt templates. The pipeline emits descriptions through four prompt templates: • NARRATIVE REF WITH ENTITIES (initial frames). Returns structured JSON of the form {"description": ..., "entities": [...]}: a dense paragraph plus an entity list with canonical mention, entity type, and a visual details string for re-identification. Conditioned on the current frame, the current entity bank, and a narrative context window of the previous 20 shot descriptions. • DIFFERENTIAL (within-shot differential frames). Receives the previous and current frame plus a shot-local description chain capped at the previous 50 descriptions. Forbids restating alreadydescribed static content; allows an empty string when no meaningful change is observed. • TRANSITION (shot boundaries). Receives the two boundary frames of adjacent shots together with their shot descriptions, and asks the model to describe the cut in terms of editing technique, narrative purpose, and framing change without restating scene content already covered. • REFERENCE (fallback). A plain single-frame captioning path used when entity-structured output is not needed. 18

Audio injection. Aligned ASR and environmental audio are appended to every prompt as a shared suffix that contains only the not-yet-used segments anchored to the current timestamp; this brings spoken and ambient content into the same frame-level annotation stream without a separate fusion stage. B.4

C ROSS - SHOT E NTITY M ATCHING

Entity vocabulary. For each initial frame the model emits a structured set of entities drawn from the fixed vocabulary of 8 categories: person, animal, object, vehicle, text, location, food, clothing. Every instance carries a canonical phrase and a concise visual details string intended to preserve identifying appearance cues across shots. Label-exact matching. Cross-shot entity assignment relies on the prompt contract rather than on a similarity threshold. The current bank is rendered into each initial-frame prompt as one record per entity (canonical name, type, and visual details), limited to the 400 most recently seen entities, and the model is instructed to reuse a bank entity’s canonical name verbatim. An emitted mention is matched against the bank by normalized label equality: lower-casing, removal of a leading article and of edge punctuation, and whitespace collapsing. A hit appends an appearance (shot, frame, surface phrase) to the existing entity; a miss registers a new entity ID. B.5

M EGA - BATCH I NFERENCE

Single batch across the workload. Each inference round collects the initial-frame prompts of at most one shot per video (a long shot re-anchored every 300 frames contributes one prompt per segment) and then fills the remaining token budget with differential prompts drawn from segments whose initial frame has already been completed. Prompt lengths are estimated analytically before batching and any prompt that would overflow the round budget is deferred. The resulting workload is submitted through one vllm.generate batch call. Token budget and engine cadence. The hard cap is 1.2M prompt tokens per round, large enough to saturate both inference GPUs while remaining stable against out-of-memory spikes from heavy multi-image prompts. The vLLM engine is rebuilt every 150 rounds so that shared-memory artifacts that accumulate over long tensor-parallel sessions do not destabilise multi-hour annotation jobs. B.6

H ARDWARE AND T HROUGHPUT

The pipeline runs on a homogeneous H200 cluster. The prep and audio stages each exceed 20× real time on a single H200, while the infer stage reaches roughly 4× real time per H200, or 8× aggregate throughput under tensor-parallel-2 serving.

C

P ER - VIDEO S CHEMA AND D ISTRIBUTION D IAGNOSTICS

C.1

P ER - VIDEO S CHEMA

Every video directory holds the six artifacts listed in Table 5. The central supervision channel is descriptions.jsonl: one JSON record per 1 FPS extracted frame, containing the absolute timestamp, the shot identifier and shot start/end frames and times, a flag is reference distinguishing reference frames from differential frames, the description text, and, on reference frames that sit at a shot boundary, a populated transition field carrying the cross-shot dynamics description. This single record type is sufficient to support every axis of the derived benchmark and every granularity of the fine-tuning corpus — a downstream sampler need only filter by is reference and the presence of transition to obtain the artifact it wants. C.2

S OURCE C URATION

Two-platform pool, three-level taxonomy. Videos are drawn from two public long-form video platforms, YouTube and Bilibili, seeded by two curated download lists that together pool 28,282 candidate URLs (15,054 YouTube + 13,228 Bilibili). Each row is pre-tagged with a three-level content taxonomy: 12 parent domains (A. Sports, B. Gaming, C. 19

Table 5: Per-video output schema. descriptions.jsonl is the central supervision artifact; the other five files carry structural metadata that supports downstream sampling. File

Producer Role

frames/{frame id}.jpg prep prep cache.json prep audio segments.json

audio

descriptions.jsonl

infer

entities final.json

infer

resume state.json

infer

1 FPS JPEG samples, ≤1024 px, JPEG quality 90 . shot list and frame manifest, fingerprinted by the shotdetector and sampler config. FASTER - WHISPER - LARGE - V 3 speech segments and Q WEN 2-AUDIO -7B-I NSTRUCT environment summaries. per-frame reference and differential descriptions; transition field on shot boundaries. per-video entity bank: canonical label, type, visual details, first appearance. entity counter and ASR deduplication state; checkpointed at every shot boundary.

Ego-Centric & Daily Life, D. Vlogs & Ceremonies, E. Arts & Crafts, F. Media & Entertainment, G. Public Safety, H. Embodied AI, I. Drones & Remote Sensing, J. AIGC-related Content, K. Formal Communication, L. Computer Use); 35 categories nested within domains; and 202 scenarios at the leaf level (e.g. Basketball within A. Sports / I. Ball Games). The full taxonomy — every leaf scenario under its parent (domain, category) — is enumerated in Table 6. Table 6: Full 12 -domain / 35 -category / 202 -scenario content taxonomy of K AIROS. Each candidate URL in the 28,282 -URL pool was pre-tagged with one (domain, category, scenario) triple at curation time. The same taxonomy carries through to the annotation corpus and to the benchmark sampler, so any per-axis evaluation cut can be re-grouped by content type. Domain

Category

A. Sports

I. Ball Games

B. Gaming

I. Video Games

Scenarios

Basketball, Soccer/Football, Volleyball, Baseball, American Football, Golf, Snooker, Bowling II. Racket Sports Tennis, Badminton, Table Tennis III. Water & Ice Swimming, Diving, Water Polo, Ice Hockey, Curling, Figure Sports Skating, Speed Skating, Short Track IV. Athletics Racing, Relay Race, Marathon, Long Jump, High Jump, Hurdles, Pole Vault, Shot Put, Discus Throw, Javelin Throw, Hammer Throw V. Gymnastics Artistic Gymnastics, Trampoline VI. Combat Sports Boxing, Wrestling, Judo, Taekwondo, Karate, MMA, WWE, Kickboxing, Wushu, Fencing, Sumo VII. Racing Sports & Formula 1, Rally Racing, Off-road Racing, Motorcycle Racing, Equestrian Horse Racing, Equestrian VIII. Extreme Sports Skateboarding, BMX, Parkour, Surfing, Alpine Skiing, Snowboarding, Rock Climbing, Bouldering, Skydiving, Bungee Jumping, Scuba Diving First-Person Shooter, MOBA, VR Games, Soulslike, Roguelike, Sandbox Games, Racing Simulation, Sports Simulation, Horror Games, Puzzle Games II. Board & Strategy Go, Chess, Chinese Chess, Gomoku, Mahjong, Poker, Bridge Games

C. Ego-Centric & I. Household Activi- Cooking, Washing Dishes, Folding Laundry, Ironing Clothes, Daily Life ties Vacuuming, Assembling Furniture, Fixing Furniture, Cleaning, Child Care II. Daily Skills Typing, Handwriting, Tool Using, Equipment Operation, Playing Instruments, Conversation, Studying, Teaching, First Aid, Car Repair, Tire Change III. Personal Care Hand Washing, Makeup, Skincare, Taking Medicine, Rehabilitation (continued on next page...)

20

( ...continued from previous page) Domain

Category

Scenarios

IV. Outdoor Activi- Gardening, Agricultural Labor, Tourism ties D. Vlogs & Cere- I. Vlogs monies II. Ceremonies

Daily Vlogs, Travel Vlogs, Walking Vlogs, Running Vlogs, Riding Vlogs, Shopping Vlogs, Gym Vlogs Wedding Ceremony, Birthday Party

E. Arts & Crafts

Pencil Sketching, Oil Painting, Watercolor Painting, Calligraphy Pottery, Origami, Paper Cutting, Embroidery, Knitting, Woodworking, Sculpting, Jewelry Making, Glass Blowing, Leather Crafting, 3D Printing

I. Visual Arts II. Handcraft

F. Media & En- I. Drama Genres tertainment II. Shows & Performance III. News & Documentary IV. Musical & Opera V. Animation

Medical, Legal, Crime, Domestic, School, Xianxia, Spy, Office Talk Show, Stand-up Comedy, Magic, Street Performance, Circus, Acrobatics, Concert News Broadcast, Documentary Musical, Opera Animated Films, Animated Series

G. Public Safety

I. Driving II. Surveillance

H. Embodied AI

I. Embodied Interac- Manipulation, Tool Use, Navigation, Human-Robot Interaction, tion Dexterous Manipulation

I. Drones & Re- I. Aerial Video mote Sensing II. Satellite Video

Urban Driving, Highway Driving Traffic Surveillance, Public Space Monitoring, Home Security

Aerial Footage, Survey Flights Satellite Video, Time-lapse Observation

J. AIGC-related I. Generated Content Text-to-Video Samples, AI Stylized Animation, Deepfake Content II. Artifacts & Con- Compression Artifacts, Moiré, AI Motion Glitch, Frame Intersistency polation Artifacts, Temporal Consistency Issues K. Formal Com- I. Academic munication II. Business III. Politics

L. Computer Use

I. Software Tutorial II. Daily Activities

Conference Presentation, Plenary Speech, Poster Presentation, Job Talk, Seminar, Group Meeting, Dissertation Defense, TEDstyle Talk, Lecture Investor Pitch, Product Launch, Board Meeting, Contract Negotiation Stump Speech, Acceptance Speech, State of the Union, Press Conference, Legislative Debate, Diplomatic Negotiation, Parliamentary Session Word, Excel, PowerPoint, PhotoShop, LightRoom, Premiere, Blender, VS Code Browse, E-mail

Duration filtering. Candidate URLs are filtered to the 10 –30 minutes duration band: the length regime where clip-scale benchmarks stop scaling and where the Time axis becomes non-trivial. Sources longer than 30 min are split into consecutive fixed-length 30-minute parts (the partNN suffixes visible in the corpus) rather than truncated, so no footage is discarded. After duration filtering and download-side failures (privacy takedowns, geo-restrictions, channel deletion), the final annotation corpus contains 19,004 videos totalling 5,420 hours. C.3

L ENGTH AND S TRUCTURE

The corpus’s content taxonomy and length distribution are summarized in Figs. 6–9, with the deepest per-scenario decomposition in Fig. 10. Together these substantiate the design claim that K AIROS lives in the multi-shot, long-context regime that single-clip benchmarks do not exercise. 21

Software Tutorial Daily Activities Politics Academic Business

Racing Sports & Equestrian Extreme Sports Combat Sports

Artifacts & Consistency Satellite Video Aerial Video

Ball Games Water & Ice Sports

Embodied Interaction Driving

GH

News & Documentary

I

Athletics

L JK

A

Racket Sports

Domain (inner ring)

A. Sports B. Gaming C. Ego-Centric & Daily Life D. Vlogs & Ceremonies E. Arts & Crafts F. Media & Entertainment G. Public Safety H. Embodied AI I. Drones & Remote Sensing J. AIGC-related Content K. Formal Communication L. Computer Use

Gymnastics Video Games

B C

F Shows & Performance

Household Activities

D

E

Daily Skills Personal Care Outdoor Activities

Drama Genres Ceremonies Musical & Opera

Visual Handcraft Arts

Vlogs

Figure 6: Two-ring content taxonomy of the K AIROS corpus. Inner ring: the 12 parent domains (codes A–L). Outer ring: the 35 categories nested within domains, with leader-line labels. Wedge sizes are video counts. The corpus is non-uniform but well-spread: A. Sports and F. Media & Entertainment together account for the largest share, while small-tail domains (G. Public Safety, H. Embodied AI, I. Drones & Remote Sensing) are kept at single-digit shares so that downstream evaluation can isolate domain-specific failures.

Video Length Distribution (dataset)

1743

1750

Number of videos

1500 1250

1147

1219

1000 806

750

845

889 529

500 250 0

792 475

405

390

193 1

<10 10 12 12 14 14 16 16 18 18 20 20 22 22 24 24 26 26 28 28 30 30 32 32 35

Video duration (minutes)

Figure 7: Per-video duration distribution on the YouTube partition (9,434 videos, 2,857 hours). Median 17.0 min, mean 18.2 min, max 32.7 min. The bulk of the corpus sits in the 10 –30-min target band; the small <10 bar reflects partNN tail fragments left after splitting longer sources rather than truncation.

D

B ENCHMARK C ONSTRUCTION D ETAILS

D.1

S OURCE - TYPE S AMPLERS AND T IER M APPING

Each of the 10 source types corresponds to a single sampler over the timestamped annotation stream. A sampler returns the materials needed for question generation — the relevant annotations, their timestamps, the evidence span, neighbouring reference-frame context, and a pool of candidate distractors — without ever accessing the raw frames. The temporal tier of a generated question is determined by the realized evidence span itself, not assigned post hoc by the LLM. Each sampler also restricts the capability labels that the generator may emit (e.g. ref audio 22

Shot Count Distribution (dataset)

2207

Number of videos

2000 1500

1334 1041

1000

911

1232 972

855

882

500 0

25

25 50

50 75

75 100

100 150 150 200 200 300

Number of shots per video

>300

Figure 8: Shots-per-video distribution on the YouTube partition. Median 91 shots per video, mean 126.4 , max 1,235 , and 1,192,552 shots in total. The right tail is dominated by fast-cut sports broadcasts and clip compilations and motivates the differential-frame chain in the annotation pipeline.

Shot Duration Distribution (dataset) 253,714

250000

249,280

219,774 195,987

Number of shots

200000 150000

121,873

100000

76,338

59,750

50000

15,836

0

1

12

23

35

5 10

Shot duration (seconds)

10 20

20 60

>60

Figure 9: Per-shot duration distribution on the YouTube partition. Median 3.7 s, mean 8.6 s. The right tail (held shots, >60 s) is what makes Time-axis questions non-trivial in the derived benchmark, and is also where the 300 -frame re-anchor schedule is exercised.

can only be tagged with A5 audio comprehension, while transition is tagged with C2 narrative transition). Table 7 reports the per-source-type composition of the released benchmark, alongside the tier range that each sampler’s evidence span can fall into. Atomic generation. Given the sampled evidence, the neighbouring reference-frame context before and after the target evidence, and a coarse temporal hint, G EMINI -2.5-F LASH returns a structured JSON object containing four fields: question, answer, distractors, and reasoning. All four are produced atomically in the same call. This encourages internal consistency between the answer and the reasoning, and avoids a separate rewriting stage that could make the distractors superficially different from the correct answer. D.2

C ROSS - BENCHMARK T EXT- ONLY L EAKAGE — P ROTOCOL

The cross-benchmark leakage probe in Table 1 of the main text fixes the solver, the prompt, the option-shuffle protocol, and the sample size identically across all benchmarks: 23

Dataset Videos per Scenario

Opera News Broadcast E-mail Documentary Birthday Party Wedding Ceremony Musical Travel Vlogs Tool Use Daily Vlogs Table Tennis School Domestic Office Medical Concert Stand-up Comedy Crime Talk Show Acrobatics Street Performance Legal Magic Circus Badminton Browse Off-road Racing Horse Racing Trampoline Agricultural Labor Temporal Consistency Issues Rally Racing Motorcycle Racing Walking Vlogs Spy Navigation Folding Laundry Human-Robot Interaction Cleaning Fixing Furniture Vacuuming Assembling Furniture Formula 1 Time-lapse Observation Volleyball Washing Dishes Diving Soccer/Football Urban Driving

169 151 132 127 110 107 107 106 106 106 104 102 102 102 101 100 93 90 88 84 83 78 77 75 75 73 73 70 67 67 66 65 64 64 64 62 62 61 60 59 58 57 56

0

231

200

293 283 278

353

400

579

600

A. Sports B. Gaming

Swimming Taking Medicine Child Care Basketball PowerPoint Ironing Clothes Handwriting Board Meeting LightRoom Word Playing Instruments Excel Baseball Equipment Operation Aerial Footage Typing Moiré Contract Negotiation Gym Vlogs Conversation PhotoShop VR Games Investor Pitch MOBA Survey Flights Water Polo Premiere Manipulation Soulslike First-Person Shooter Artistic Gymnastics Oil Painting Dexterous Manipulation Parkour Rock Climbing Surfing Pencil Sketching Puzzle Games Calligraphy Frame Interpolation Artifacts Snowboarding Studying Watercolor Painting Bungee Jumping scuba diving Skydiving BMX Cooking Bouldering

Equestrian Product Launch Sandbox Games Compression Artifacts Taekwondo Sumo Kickboxing Ice Hockey Alpine skiing Judo Skateboarding Wrestling MMA WWE Wushu Highway Driving Karate High Jump Discus Throw Boxing Javelin Throw Fencing Hammer Throw Shot Put Xianxia Rehabilitation Long Jump Hurdles Pole Vault Marathon Gardening Satellite Video Skincare Makeup Curling Tourism Woodworking Leather Crafting Embroidery Knitting Short Track Racing Simulation Glass Blowing 3D Printing American Football Sculpting Jewelry Making Paper Cutting Relay Race

55 55 54 53 53 53 51 51 49 49 49 49 48 48 48 47 47 46 46 45 44 43 43 42 41 41 40 40 39 39 38 38 37 36 36 36 36 36 36 36 35 35 35 34 34 34 34 33 33

0

C. Ego-Centric & Daily Life D. Vlogs & Ceremonies

200

400

600

0

Domain (bar colour)

E. Arts & Crafts F. Media & Entertainment

G. Public Safety H. Embodied AI

Animated Series Lecture Diplomatic Negotiation Origami Sports Simulation AI Motion Glitch Racing Deepfake Poster Presentation Figure Skating Group Meeting Home Security Plenary Speech Dissertation Defense Parliamentary session Hand Washing Animated Films Golf Stump Speech Job Talk Pottery Seminar Traffic Surveillance Speed Skating Snooker Chinese Chess Conference Presentation Acceptance Speech Bowling TED-style Talk State of the Union Shopping Vlogs Running Vlogs Riding Vlogs Chess Teaching Public Space Monitoring Legislative Debate Press Conference Text-to-Video Samples Tennis Horror Games Mahjong VS Code Bridge Go First Aid Blender

33 33 31 31 30 30 30 30 30 29 29 28 27 27 27 27 26 25 25 25 24 24 24 23 23 23 21 21 21 21 20 20 19 19 19 17 16 16 16 15 15 15 15 14 14 14 14 14 13

200

400

Number of videos

600

I. Drones & Remote Sensing J. AIGC-related Content

13 13 13 12 12 12 11 11 10 10 10 10 10 10 10 10 10 9 8 8 7 7 7 7 6 6 6 6 5 5 5 4 3 3 3 3 3 3 3 3 2 2 2 1 1 1 1 1

0

200

400

600

K. Formal Communication L. Computer Use

Figure 10: Per-scenario video count across all 202 scenarios of the corpus, grouped by parent category and parent domain (bar color). Bars are sorted within each category by descending count; the resulting distribution is long-tailed but non-degenerate, with no single scenario dominating and every scenario receiving a meaningful share of the corpus. Table 7: Per-source-type composition of K AIROS-Bench after audit and human review. Allowed tiers is the range of evidence spans the sampler can realise; the tier of a question is set by the realized span, not by the source type. Source type

Evidence shape

ref perception ref ocr ref audio diff change diff sequence entity tracking transition cross shot long range full video

single reference frame, scene + entities + relations single reference frame, on-screen text single reference frame, aligned ASR / environment ref + several differential frames inside one shot full differential chain of one shot cross-shot entity reappearance two boundary frames + adjacent shot descriptions several adjacent or near-adjacent shots evidence span > 300 s whole-video integration

Allowed tiers

#Q

T1 T1 T1 T2 T2 T3–T5 T2–T3 T3–T4 T4–T5 T5

614 177 184 439 555 366 14 236 144 141

• Solver. GEMINI -3.1- PRO , no video and no audio, K=1 shuffle, seed 42, greedy decoding, with the smallest thinking budget the model accepts (128 tokens). • Sample size. 500 questions per benchmark, stratified over capability/tier/source where applicable, otherwise uniform. • Prompt. A fixed system prompt states that no video or images are available and asks for the answer letter on the first line followed by one sentence of rationale on the second; the user message contains the question stem and the shuffled options. No in-context examples and no reasoning before the answer. • Random baseline. Computed per benchmark from the realized option counts (1/N for fixed-N benchmarks, and the per-question mean of 1/Ni for variable-option benchmarks). • Leakage. Solver accuracy minus random baseline, reported in percentage points. 24

9 public video-MCQ benchmarks public-benchmark envelope Kairos (ours)

text-only leakage (pp above random baseline)

CG-Bench Linear fit over 9 public benchmarks (Gemini 3.1 Pro, text-only, 500 stratified each)

cg_bench 35

egoschema

text-only leakage lift (pp above random baseline)

deve_qa 30

video_holmes

longvideobench

video_mme

25 hourvideo

20 15

mvbench

Kairos (ours)

lvbench

10

35

Video-Holmes Video-MME

30

0 200

300

400

500

avg (question + options) length

characters

600

700

20

HourVideo

15

Politics

Daily Activities Software Tutorial Business

Kairos (ours)

MVBench LVBench

10 20

800

Figure 11: Leakage vs. average MCQ length. Linear fit of text-only-leakage (pp) on average total characters per question across the ten benchmarks of Table 1 of the main text. K AIROSBench is to the right of the cloud (long stems) and below the line (lower leakage than length predicts).

LongVideoBench

25

5 9 public video-MCQ benchmarks Kairos (ours) excluded from fit fit over 9 public benchmarks (R²=0.077, p=0.445)

EgoSchema

DeVE-QA

40

60

80

100

avg (question + options) length

120

words

140

160

180

Figure 12: Leakage–length scatter. Perbenchmark scatter underlying the fit in Fig. 11. Each dot is one of the ten benchmarks; axes are average total MCQ characters and the absolute lift over the random baseline. K AIROS is the rightmost low-leakage point.

Combat Sports

Academic

Extreme Sports

Generated Content Artifacts & Consistency Satellite Video Aerial Video Embodied Interaction Surveillance News & Documentary Musical & Opera Animation

I GH

Shows & Performance

F

J

E

Drama Genres

Athletics

K L

Domain (inner ring)

A

Ball Games

Water & Ice Sports

D

C

B

A. Sports B. Gaming C. Ego-Centric & Daily Life D. Vlogs & Ceremonies E. Arts & Crafts F. Media & Entertainment G. Public Safety H. Embodied AI I. Drones & Remote Sensing J. AIGC-related Content K. Formal Communication L. Computer Use

Racing Sports & Equestrian

Visual Arts

Racket Sports Gymnastics

Handcraft Ceremonies Vlogs Outdoor Activities Personal Care Daily Skills

Video Games Board & Strategy Games Household Activities

Figure 13: Two-ring view of the capability axis of K AIROS-Bench. Inner ring: the four cognitive levels A (Perception, Space), B (Events, within-shot Dynamics), C (Temporal, cross-shot Dynamics), D (Localisation, holistic). Outer ring: the 17 capability cells nested within each level.

Figure 11 fits a linear regression of leakage on average total characters across the ten benchmarks; K AIROS-Bench has the longest stems among the four-option benchmarks but lies below the regression line, i.e. its leakage is lower than its length would predict. Figure 12 shows the underlying scatter, broken down by benchmark. D.3

B ENCHMARK D ISTRIBUTION D IAGNOSTICS

Figures 13–18 expand the three-axis summary in Fig. 4 of the main text to the content and structural axes, none of which are exposed in the main paper. Figure 13 reframes the capability axis as a tworing pie (axis → capability) so that the relative weight of holistic Axis-D cells is legible. Figure 14 reports the benchmark video count by content category, and Fig. 15 the per-scenario decomposition; both confirm that the 820 benchmark videos preserve the corpus-level taxonomy distribution rather than over-sampling a small subset. Figures 16–18 report per-video duration, shots per video, and per-shot duration on the benchmark subset. 25

Number of questions

200

206 202

A. Sports B. Gaming C. Ego-Centric & Daily Life

191 164

150

158 155

D. Vlogs & Ceremonies E. Arts & Crafts F. Media & Entertainment

125

118

85

84

84 62

61

50

mb

Co

s

rts po

S S at eme tr Ex

J. AIGC-related Content K. Formal Communication L. Computer Use

109 106 105 101 86

rt po

G. Public Safety H. Embodied AI I. Drones & Remote Sensing

142

100

0

Domain

57

56

47

46

39

37

34

32

31

27

25

24

22

21

17

11

ics mes craft mes nres ports trian ities kills ance orial emic ency litics logs ction Care ports ness Arts ities ideo ance stics mes nies ation pera tary ideo tent ities iving i S t l let n o n v v l a a V e t s d iv d Po Dr t S Bus isua Acti rial V rveil ymna gy Ga rem Anim l & O ume llite V d Co Acti era nal Ath eo G Han all G ma G Ice S Eque d Act Daily rform re Tu Aca onsis V oor a c Int erso acke y e e a B e l G Ce Ae Su Vid ate sic Do Sat erat Dail P R &C Dra ater & rts & seho ied & P Softw td Str Mu s & od en cts Ou ws b & po Hou W a w G o f S m d e i E N g Sh ar Art cin Bo Ra

Figure 14: Per-category video count of K AIROS-Bench (820 videos), grouped by parent domain (bar color). The benchmark preserves the long-tail shape of the corpus (Fig. 3 of the main text); no domain is dropped, and the small-tail domains G, H, I are over-sampled relative to their share of the corpus to keep per-domain question counts non-trivial. Benchmark Questions per Scenario

Basketball Sports Simulation Rock Climbing Wrestling Legal Baseball Equestrian Snowboarding School Tool Use Racing Simulation First-Person Shooter Fencing Motorcycle Racing Jewelry Making Discus Throw Premiere Compression Artifacts Snooker Horse Racing Playing Instruments Pottery Soulslike Speed Skating Badminton Bowling Marathon Survey Flights Concert Relay Race Running Vlogs American Football Embroidery Surfing Office Karate Kickboxing Pole Vault Figure Skating Table Tennis PhotoShop Bungee Jumping Off-road Racing Agricultural Labor Sandbox Games Traffic Surveillance WWE MMA Wedding Ceremony

28 28 28 27 27 26 26 25 25 24 24 24 24 24 23 23 23 23 22 22 22 22 22 22 22 22 21 21 21 21 21 21 21 21 21 21 20 20 20 20 19 19 19 19 19 19 19

0

10

20

31

30

A. Sports B. Gaming

35

Water Polo AI Motion Glitch scuba diving Xianxia Legislative Debate 3D Printing Skydiving Acceptance Speech Sumo Cooking Leather Crafting Trampoline Daily Vlogs Alpine skiing Frame Interpolation Artifacts MOBA Aerial Footage Racing Formula 1 Hammer Throw Artistic Gymnastics Parkour Boxing Product Launch Talk Show Rehabilitation Long Jump Skateboarding Curling Fixing Furniture Gardening Animated Films Judo Equipment Operation Hurdles Moiré LightRoom Stand-up Comedy Wushu Assembling Furniture Swimming High Jump PowerPoint Tennis Acrobatics Home Security Excel Travel Vlogs Investor Pitch

Knitting Spy Oil Painting Manipulation Ice Hockey Makeup Shot Put Magic Medical Dexterous Manipulation Opera Conversation Deepfake TED-style Talk Gym Vlogs Human-Robot Interaction Domestic Contract Negotiation Ironing Clothes Skincare Calligraphy Board Meeting Parliamentary session Circus Taking Medicine Puzzle Games Chinese Chess Seminar Volleyball Documentary Diplomatic Negotiation News Broadcast Poster Presentation Street Performance Time-lapse Observation Birthday Party Tourism Studying Short Track Watercolor Painting State of the Union Musical Animated Series Golf Stump Speech Folding Laundry Temporal Consistency Issues Chess Group Meeting

19 19 18 18 18 18 18 18 18 18 18 17 17 17 17 17 17 17 17 17 17 16 16 16 16 16 16 16 16 16 16 16 16 16 15 15 15 15 15 15 15 15 15 15 15 15 15 15 14

0

10

C. Ego-Centric & Daily Life D. Vlogs & Ceremonies

20

30

14 14 14 14 14 14 14 14 14 14 14 13 13 13 13 13 13 13 13 13 13 13 12 12 12 12 12 12 12 12 12 12 12 12 12 12 11 11 11 11 11 11 11 11 11 11 11 11 11

0

Domain (bar colour)

E. Arts & Crafts F. Media & Entertainment

G. Public Safety H. Embodied AI

10

BMX Teaching Word Bouldering Vacuuming Washing Dishes Cleaning Satellite Video Plenary Speech VR Games Taekwondo Sculpting Handwriting Glass Blowing Javelin Throw Conference Presentation Woodworking Pencil Sketching Shopping Vlogs E-mail Rally Racing Text-to-Video Samples Navigation Crime Job Talk Typing Browse Diving Origami Paper Cutting Urban Driving Hand Washing Mahjong Lecture Child Care Riding Vlogs Highway Driving Dissertation Defense Soccer/Football Horror Games First Aid Public Space Monitoring Walking Vlogs Press Conference Blender Go Bridge

20

1

30

0

I. Drones & Remote Sensing J. AIGC-related Content

K. Formal Communication L. Computer Use

Number of questions

2 2 2

3 3

4 4 4

5 5

6 6 6 6 6 6

7 7

8 8 8 8 8 8 8 8

9 9 9 9 9 9 9 9 9 9

11 11 11 11 10 10 10 10 10 10

10

20

30

Figure 15: Per-scenario video count of K AIROS-Bench across the 202 scenarios of the content taxonomy. Bars are colored by parent domain. The benchmark video pool is intentionally spread thin across scenarios rather than concentrated in a few high-volume cells, so that per-scenario evaluation is well-defined for as many cells as possible at the cost of small absolute counts in the long tail.

E

E VALUATION P ROTOCOL

We evaluate 21 models in total: 7 closed-source proprietary and 14 open-weight. A unified runner groups questions by video id, prepares the video once per video for the relevant backend, then dispatches every question for that video in parallel (API: ThreadPoolExecutor) or batched (vLLM: generate batch). 26

Video Length Distribution (benchmark)

189

175

175

121

100 71

75 50

0

Shot Duration Distribution (benchmark)

72

67 43

38

35

31

123

120

125

100

100 77

75

84

77

67

50

0

<10 10 12 12 14 14 16 16 18 18 20 20 22 22 24 24 26 26 28 28 30 30 32 32 35

Video duration (minutes)

0

20,915

18,026

17500

15,846

15000 12500

10,042

10000 7500

6,782 5,116

5000

25

15

11

20,372

20000

Number of shots

127

125

Number of videos

Number of videos

150

25

Shot Count Distribution (benchmark)

172

150

2500

25

25 50

50 75

75 100

100 150

150 200

Number of shots per video

200 300

>300

0

1,370

1

12

23

35

5 10

Shot duration (seconds)

10 20

20 60

>60

Figure 16: Per-video duration Figure 17: Shots per bench- Figure 18: Per-shot duration on the 820 benchmark videos. mark video. on the benchmark videos.

E.1

BACKENDS AND V IDEO - INPUT N EGOTIATION

Five backends. The runner instantiates one of five backends per model based on its registry entry: • vLLM (open-weight): frame mode for every family except LLAVA - VIDEO -7 B, which is fed the video directly. • HuggingFace Transformers (open-weight): used for the few open-weight models without a vLLM video implementation, in frame mode only. • OpenAI direct (closed-source + OpenRouter-proxied open-weight): probes data:video/mp4 native video once per model and caches the result; falls back to N uniformly subsampled frames otherwise. • Gemini direct: File API upload + 2 s processing poll, cached per video path. Native video is the only input mode. • Anthropic direct: frame mode only (no native video support in the public API). Native-video models. Only GEMINI -2.5- FLASH, GEMINI -3.1- PRO, LLAVA - VIDEO -7 B run with native video input in our evaluation; every other model receives uniformly subsampled frames at the per-model budget reported in Table 8. E.2

P ER - MODEL D ECODING AND F RAME B UDGET

Table 8 enumerates the eval-time configuration used for every model in the leaderboard of Table 2 of the main text. The frame budget is the maximum number of frames the runner sends per question; the actual count for a given question equals min(budget, ⌊video dur · 1 ⌋) so that very short videos are never up-sampled past their native 1 FPS rate. Decoding is greedy (temperature 0) for every model except GPT-5.5, whose API only accepts its default sampling temperature; for the OpenQA run we keep the same decoding so that the only intentional change between MCQ and OpenQA is the answer format. Scoring. A model output is parsed by first stripping any <think> block, then taking the last explicit Answer:X if present, otherwise a leading bracketed or bare letter; outputs with no such letter are counted as incorrect. Per-tier, per-capability, per-source, and per-domain accuracies are computed by simple bucket-then-average over the parsed letters. Three openweight models (INTERNVL 3.5-8 B, MIMO - VL -7 B, LLAVA - VIDEO -7 B) are decoded with vLLM structured outputs that constrain the final token to a single letter, because their unconstrained outputs occasionally trail off into reasoning without committing to a choice. E.3

O PEN - ENDED QA J UDGE

Why an OpenQA pass. A multiple-choice task is easy to score but lossy: a model can recognise the right answer without being able to produce it. We therefore additionally run an OpenQA pass on the same 2,870 questions, asking each model to generate a free-form answer rather than pick a letter. Decoding stays greedy; the per-model token budget is raised to at least 256 tokens so that free-form answers are not truncated. Judge. Each generation is scored against the reference answer by GEMINI -2.5- FLASH , which sees only the question stem, the reference answer, and the candidate; no video, no audio, no chain-ofthought. The judge returns an integer 0–3 in the MMBench-Video style, and we report binarised accuracy at threshold ≥ 2 . The leaderboard is reproduced in Table 3 of the main text. 27

Table 8: Per-model evaluation configuration. Type is the backend used. TP is the tensor-parallel size on H200 GPUs (vLLM only). Frames is the per-question frame budget (native = native video input). Max tokens is the decoding cap. Sampling is greedy throughout. Model

Type

Closed-source proprietary gemini gemini openai openai openai openai openai

GEMINI -3.1- PRO GEMINI -2.5- FLASH GPT-5.5 GPT-4 O GPT-4 O - MINI NOVA -2- LITE SEED -2.0- LITE

TP Frames Max tokens Notes — — — — — — —

native native 256 64 128 32 32

1024 512 2048 1024 1024 1024 1024

video-native video-native forced frame mode forced frame mode forced frame mode forced frame mode forced frame mode

4 2 2 2 4 — 2 2 1 1 1 — — 1

16 12 16 12 16 1 16 16 16 16 16 16 32 native

512 512 512 2048 64 512 8192 512 512 512 4096 1024 1024 512

frame mode frame mode frame mode structured-output letter enable thinking=False 1-frame ceiling per the model card frame mode frame mode frame mode frame mode structured-output letter frame mode forced frame mode video-native; structured-output letter

Open-weight INTERNVL 3-78 B INTERNVL 3.5-38 B INTERNVL 3-8 B INTERNVL 3.5-8 B GLM -4.5 V GLM -4 V-9 B STEP 3- VL -10 B QWEN 3- VL -30 B - A 3 B QWEN 3- VL -8 B QWEN 2.5- VL -7 B MIMO - VL -7 B COGVLM 2- VIDEO -13 B GEMMA -4-31 B LLAVA - VIDEO -7 B

vLLM vLLM vLLM vLLM vLLM HF vLLM vLLM vLLM vLLM vLLM HF openai vLLM

F

F INE - TUNING R ECIPE

F.1

H YPERPARAMETERS

Table 9 summarizes every fine-tuning hyperparameter. The recipe deliberately stays close to a vanilla LLaMA-Factory qwen2 5 vl template: LoRA on the seven attention/MLP projections, frozen vision tower, single epoch, cosine schedule. The only non-default choice is the cutoff length, which is raised to 16,384 tokens to fit the long anchor-based prompts at 32 frames. Eval-time frame ablation. At eval time, the fine-tuned checkpoint is decoded with {16, 32, 64} uniformly subsampled frames per question; the headline numbers across these frame budgets and across the four evaluation benchmarks are reported in Table 4.

G

L IMITATIONS AND B ROADER I MPACT

G.1

L IMITATIONS

VLM bias. All annotations are produced by Qwen3-VL-8B-Instruct (Bai et al., 2025), so a systematic perception bias of that VLM also biases K AIROS descriptions; running the pipeline with a different VLM is straightforward and would help quantify the bias. Language-only entity matching. Cross-shot entity matching is language-only, so for visuallyhard-to-describe entities (e.g. visually similar but semantically distinct people in crowd scenes), label-exact matching splits an identity whenever the model paraphrases an established name instead of reusing it, and merges two mentions only when it assigns them the same label. This is addressable by adding a vision-only re-identification signal as a visual matching stage, which we leave to future work to keep the pipeline VLM-only. Residual leakage. Even after the multi-model audit and human review, the cross-benchmark leakage probe still finds +9.6 pp lift over random under GEMINI -3.1- PRO — the lowest of the ten benchmarks tested but not zero. We attribute this residual to the linguistic anchors that the ques28

Table 9: Fine-tuning hyperparameters. The recipe uses LLaMA-Factory’s qwen2 5 vl template; the seven LoRA target modules are the standard transformer projections. Setting

Value

Model and parameterisation Base model Q WEN 2.5-VL-7B-I NSTRUCT Framework LLaMA-Factory Adapter LoRA, rank 32 , α =64 LoRA target modules q,k,v,o,gate,up,down proj Frozen modules vision tower (multi-modal projector trainable) Precision bf16 Distributed DeepSpeed ZeRO-3 Hardware 4×H200 Optimization Optimizer Learning rate Schedule Weight decay Per-device batch size Gradient accumulation Effective batch size Epochs Cutoff length

AdamW (LLaMA-Factory default) 1×10−4 cosine, warmup ratio 0.03 0.01 1 8 32 1 16,384 tokens

Visual input Frame budget per sample Sampler Input resolution Frame source

32 uniform 420×420 pixels per frame (image and video pixels) pre-extracted 1 FPS JPEGs (no re-decoding)

tions include in lieu of numeric timestamps; removing those anchors would reduce leakage further but would also degrade question grounding for non-visual solvers, so we leave the trade-off explicit rather than tune it. Per-tier population imbalance. Two cells of the capability axis have small populations (D1 narrative summarisation 5 questions, C2 narrative transition 14 questions). Per-cell accuracy on those cells is therefore high-variance and should be read as a coarse signal rather than a precise number. The headline overall accuracy and the per-tier breakdown are unaffected. G.2

B ROADER I MPACT

Intended uses. K AIROS is intended for the development and evaluation of long-form video understanding systems in three modes: (i) as supervision data for video-language pretraining and finetuning; (ii) as a benchmark for long-form video QA; (iii) as a substrate for evaluating video-language alignment, entity tracking, and temporal grounding methods that go beyond clip-level decisions. Risks. The dataset contains only annotations of public, platform-distributed videos, with platformtakedown semantics preserved. Personally-identifying information that appears in the source videos (e.g. recognisable individuals in sports broadcasts) is not added by the annotation pipeline; faces and identities are referenced only at the level the source already exposes them. Misuse risks are those typical of any large video corpus; the annotation-side frames are not redistributed, which limits this surface.

29

Record · ID 668095 · SHA-256 39999881a32bee45
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.