Conceptio › Archive › arXiv CS
arXiv CSopen access

Concord: A Video Relational Algebra for Cross-Modal Query Optimization

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
data-managementdatabasesstorage
databases, sql, data management, storage

Concord: A Video Relational Algebra for Cross-Modal Query Optimization Sultan Muratbek∗

Charisse Ivana Yeung∗

Chanwut Kittivorawong

Alvin Cheung

UC Berkeley Berkeley, California, USA

UC Berkeley Berkeley, California, USA

UC Berkeley Berkeley, California, USA

UC Berkeley Berkeley, California, USA

arXiv:2609.05756v1 [cs.DB] 4 Sep 2026

Abstract Semantic video queries let users embed natural language prompts and use multimodal large language models (MLLMs) to interpret the video. Such queries are increasingly popular for querying video data. However, their expressiveness comes at a steep cost: an MLLM may process hours of media to return only seconds of relevant output, making naive execution slow, expensive, and inaccurate. We propose Concord, a system for expressing and optimizing semantic video queries. We makes three contributions. First, we introduce Video Relational Algebra (VRA), a nested algebra over videos, transcripts, frames, and object tracks that captures common semantic video operations. Second, we derive a set of approximate optimizations that rewrite VRA queries to reduce MLLM usage while improving result quality. For narrated video, Concord either processes transcripts instead of video or uses them to identify video clips for MLLM processing. For cross-camera queries without narration, detection and tracking replace a whole-video MLLM join with a track-level relational join. Third, we evaluate Concord on real-world videos. Across 4.59 hours of soccer broadcasts and 3.92 hours of lectures, transcript-tovideo queries send only 5.32% and 2.47% of source-video duration to the MLLM and reduce MLLM cost by up to 87%. In two five-second highway clips with 18 manually adjudicated cross-camera vehicles, a Detect-Track-Join query improves F1 from .364 to .813 while making no MLLM calls. See our project at concord-db.github.io.

1

Introduction

Multimodal large language models (MLLMs) let users ask semantic questions over videos: when a goal is scored in a soccer match, when an experiment visibly succeeds in a lab demo, or which vehicle appears across multiple camera views of the same highway. Naive execution sends complete, temporally redundant videos to an MLLM, incurring high cost and latency when relevant evidence is sparse. Furthermore, Video contains source-aligned modalities such as frames, audio, transcripts, and metadata. Many videos, including those on YouTube, already provide narration or transcripts. In narrated videos, transcripts offer a cheaper access path: they may answer a query directly or identify candidate source-time intervals for video inference. For example, a commentator announces a goal or a lecturer introduces a demo before it is visible. When the transcript lacks sufficient evidence to answer the query directly, it can still identify candidate intervals, source-time regions likely to contain an answer. The MLLM can then process these short regions instead of the complete video. Not all videos have useful text channels. For cross-camera vehicle association, a direct implementation sends the synchronized videos to an MLLM, which must discover vehicles, maintain their identities ∗ Equal contribution.

within each video, and associate them across camera views. Concord instead decomposes this semantic operation into a Detect–Track– Join query that derives entity tracks from pixels and joins them across cameras. This alternative changes the query’s processing granularity from whole videos to entity-level records. Current video analytics systems optimize video processing through sampling, proxies, indexes, materialized views, and relational hints [3, 10, 11, 16], while semantic data systems optimize programs whose operators may invoke foundation models [9, 13, 17]. What is missing is a compact algebra in which ordinary relational operators compose with media-specific representations. Such an algebra expresses video-only inference, transcript substitution, transcript-to-video pushdown, and track-level joins as alternative queries with the same intent. Towards that goal, we present Concord, a declarative system for multimodal analytic workflows. Concord provides a query API over records containing audio, video, transcripts, source-time views, frames, detections, and tracks. Users construct Video Relational Algebra (VRA) queries through this API. Concord then optimizes the VRA queries by rewriting them to using different media inputs, temporal extent, or processing granularity. In sum, this paper makes three contributions: • We define VRA: its media model, source-time views, deterministic and semantic functions, and operators for various video-related operations (§3). • We describe three VRA rewrites: modality substitution, crossmodal temporal pushdown, and the algebraic decomposition of whole-video semantic association into Detect–Track–Join (§4). • We evaluate Concord’s query rewrites on real-world videos. Transcript-to-video queries improve the cost–quality frontier for narrated event localization (§5.1 and §5.2), while Detect–Track– Join improves cross-camera F1 from .364 to .813 with no MLLM calls (§5.3). Through our case studies, Concord presents an opportunity for future semantic video processing systems, where modality, temporal extent, and processing granularity should be explicit query choices, subject to event- and entity-coverage constraints.

2

Background and Related Work

Video query processing. Video DBMSs reduce semantic-inference cost using specialized execution alternatives. NoScope [10] constructs model cascades and neural proxies; TASTI [11], EVA [21], and Seiden [3] use semantic indexes, materialized inference results, sampling, and temporal propagation. Spatialyze [12] exploits spatial and temporal metadata, while MIRIS [1] integrates query planning with object tracking. These systems optimize which models, frames, or stored results are evaluated. Concord builds on this

Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, and Alvin Cheung

by making media-derived representations composable within a relational query, enabling alternatives that change the representation, temporal extent, or record granularity supplied to semantic computation. Declarative video and multimodal queries. VIVA [16] is closest to Concord in making relationships between semantic functions declarative. Its relational hints specify when one registered model may replace or filter another, including the use of transcript search as a filter for visual recognition. VOCAL-UDF [22] supports compositional video queries by constructing missing program-based or distilled-model UDFs. VRA differs by representing aligned transcripts, temporal clips, detections, and tracks as query-visible values and relations that preserve source identity and time. Multimodal systems provide complementary interfaces. ThalamusDB [9] evaluates natural-language predicates over visual, audio, and textual data; CAESURA [18] generates executable multimodal queries from natural language; and KathDB [20] provides a unified relational interface for multimodal data. Systems such as Palimpzest [13], LOTUS [15], and DocETL [4, 17] optimize declarative semantic operators over unstructured data. Concord focuses specifically on algebraic rewrites that exchange or compose sourcealigned media representations while preserving the query’s output schema.

3

Concord and Video Relational Algebra

Users express semantic video queries in Concord by writing Video Relational Algebra (VRA) queries. A VRA query is an expression composed of VRA operators. Concord may rewrite an initial query into alternative queries that preserve the query intent and output schema. §4 describes these rewrites, which may change the input modality, temporal extent, or processing granularity.

3.1

Media Data Model

VRA extends the relational data model with first-class media values. A relation may contain scalar attributes, nested records, collections, and media-valued attributes. A video-valued attribute 𝑣 contains a reference, such as a local file path or URI, through which VRA operators (§3.2) access the corresponding video content. Each video is identified by a unique source_id that is retained by all derived values and records. While 𝑣 provides access to the video content, source_id identifies its origin, allowing independently processed clips and their results to be regrouped by source. Temporal representations derived from a video use source time. A source-time timestamp is measured on the timeline of the original video. Temporal information is represented using ordinary attributes, such as start and end for an interval. Derived media values, such as clips, retain their source-time intervals, so cliprelative results can be mapped back to the original video. Shared source identity and source time let VRA compose media-derived values and rewrite queries without losing their association with the originating video or its timeline. A VRA relation may represent media at different processing granularities: each tuple may correspond, for example, to a complete video, a clip, a frame, a detection, or a track. The operators in the following section transform both the values stored in these tuples and the granularity at which subsequent computation is performed.

3.2

VRA Operators

A VRA query transforms an input relation into an output relation by composing VRA operators. Relational operators such as Map, Filter, Reduce, Unnest, and Join transform and combine relations whose tuples may contain ordinary and media-valued attributes. Media operators expose common media transformations such as materialization, transcription, detection, and tracking. § 4 shows how Concord uses these transformations to construct alternative approximate queries. Media operators have tuple-wise Map semantics: they read designated attributes from each tuple, store their result in a designated output attribute, and preserve the remaining attributes. For example, View𝑣,𝑠,𝑒→𝑣𝑐 (𝑅) ≡ Mapview(𝑣,𝑠,𝑒 )→𝑣𝑐 (𝑅). This equivalence specifies the relational behavior of View, while its named form keeps materialization explicit in VRA queries and rewrites. Transcribe, Detect, and Track follow the same tuple-wise, attribute-preserving convention. Relational operators may be parameterized by deterministic or semantic functions. Deterministic functions execute ordinary nonMLLM code, whereas semantic functions are defined by naturallanguage instructions and a structured output schema and may invoke an MLLM. The operator determines how records are processed, while the function supplies the query-specific computation. Table 1 defines the VRA operators used in this paper. In the table, attributes and parameters before → are inputs, and the attribute after → stores the result. Operators preserve attributes not explicitly replaced. Together, these operators express queries ranging from semantic classification and aggregation to temporal localization and cross-video entity association.

3.3

Example Queries

We now show how to use VRA to write queries. 𝑄 1 and 𝑄 2 are temporal event-localization queries that motivate our cross-modal rewrites (§4.1, §4.2). 𝑄 3 is a cross-camera entity-association query that combines semantic processing with structured media records and motivates our semantic-to-structure rewrite (§4.3). 3.3.1 𝑄 1 : Soccer goal localization. A sports analyst wants to locate every goal in a collection of soccer broadcasts. For 𝑄 1 and 𝑄 2 , let 𝑅𝐸 (𝑠𝑜𝑢𝑟𝑐𝑒_𝑖𝑑, 𝑣, 𝑡) contain one tuple per input video. The attribute source_id identifies the originating video, 𝑣 is a media reference through which the video content can be accessed, and 𝑡 is its source-aligned timestamped transcript. Given an event description 𝑝, the semantic function localize(𝑣, 𝑝) returns a collection 𝐸 of event records represented as source-time points or intervals. The video-only query applies this function to 𝑣 and does not read the transcript 𝑡:   Unnest𝐸 Maplocalize(𝑣,𝑝 )→𝐸 (𝑅𝐸 ) . Map applies the localization function to each input video and stores the returned collection in 𝐸 while preserving the input attributes. Unnest then emits one tuple per event occurrence. For 𝑄 1 , 𝑝 specifies that an actual goal is scored, so the query returns one sourcetime timestamp for every goal in each soccer video.

Concord: A Video Relational Algebra for Cross-Modal Query Optimization

Table 1: Video Relational Algebra operator catalog. Operator Source operator Input(𝑢 )

Semantics Create a relation from records stored at 𝑢, including records containing media values.

Relational operators Map 𝑓 →𝐴 (𝑅) Apply 𝑓 to each tuple and store result in the field(s) specified by 𝐴. Filter𝑝 (𝑅) Retain tuples for which predicate 𝑝 holds. Reduce𝐾,𝑔→𝐴 (𝑅) Group tuples by 𝐾, produce one tuple per group, and store the result of 𝑔 in attribute 𝐴. Unnest𝐴 (𝑅) Expand collection-valued field 𝐴 into one tuple per element. Join𝑝,𝑠 (𝑅, 𝑆 ) Combine tuples satisfying predicate 𝑝 and optionally assign matching score 𝑠. Resolve 𝑓 (𝑅) Apply a set-level resolution function 𝑓 to remove, merge, or select among conflicting or equivalent records. Media operators View𝑣,𝑠,𝑒→𝑣𝑐 (𝑅) Materialize interval [𝑠, 𝑒 ) of 𝑣 and store the resulting clip in 𝑣𝑐 . Transcribe𝑣→𝑡 (𝑅) Derive source-aligned timestamped text from 𝑣 and store it in 𝑡 . Detect𝑣,𝐶→𝐷 (𝑅) Detect instances of classes 𝐶 in 𝑣 and store the timestamped detections in 𝐷. Track𝐷→𝑇 (𝑅) Link detections 𝐷 within each video source and store the resulting tracks in 𝑇 .

3.3.2 𝑄 2 : Lecture event localization. A student wants the exact portions of a lecture in which a physical demonstration occurs, excluding explanations, setup, and discussion of the demonstration. For example, 𝑝 may request every interval in which a sustained tone visibly causes a drinking glass to shatter. 𝑄 2 instantiates the localization expression above over lecture videos. Here, satisfying 𝑝 requires the co-occurrence of audible and visible evidence: a passage that discusses the experiment or contains the tone without the visible shattering does not qualify. The query therefore returns the source-time intervals containing the requested physical event. 3.3.3 𝑄 3 : Cross-camera vehicle association. A traffic analyst has synchronized videos from cameras observing overlapping regions of the same highway. The analyst wants to determine which vehicle observations across the camera feeds correspond to the same physical vehicle and track its movement across cameras. The desired output is one cross-camera trajectory record for each vehicle visible in multiple feeds. Let 𝑅𝑀 (source_id, 𝑣) contain one tuple per input camera video. 𝑄 3 returns one association record per physical vehicle visible in multiple cameras:

The following VRA query uses this association function to solve the cross-camera joining task:  Unnest𝐴 Reduce ∅,Sem_Associate→𝐴 (𝑅𝑀 ) . Sem_Associate returns a collection 𝐴 of cross-camera association records, and Unnest produces one tuple per associated vehicle. Unlike 𝑄 1 and 𝑄 2 , this query returns entities and their cross-source relationships as structured records. Downstream VRA operators can therefore filter, group, join, and process these records as normal relational data.

4

Optimizing VRA Queries

VRA exposes three properties of video processing as query choices: the representation (i.e., modality) supplied to a semantic function, the temporal extent processed by that function, and the granularity of its intermediate records. Concord uses these properties to rewrite queries via modality substitution, temporal candidate pushdown, and operator decomposition to improve query performance. Each rewrite specifies the properties that the rewritten query must satisfy and preserves its output schema, but may change result quality by altering the evidence presented to semantic functions. Quality may degrade when the new representation omits relevant evidence, or improve when it removes irrelevant or excessive context, since longer context can reduce model performance even when the relevant information is present [2]. We therefore call these rewrites approximate and use ⇝ to denote them below.

4.1

O1: Modality Substitution

Suppose relation 𝑅 contains representations 𝑥 and 𝑦 derived from the same media source. We call these representations source-aligned: they share the same source_id and, when temporal, express timestamps in source time. Let 𝑓 𝑥 and 𝑓 𝑦 be schema-compatible semantic functions over their respective representations, parameterized by task description 𝑝. O1 substitutes one representation for the other: Map 𝑓 𝑥 (𝑥,𝑝 )→𝐴 (𝑅) ⇝ Map 𝑓 𝑦 (𝑦,𝑝 )→𝐴 (𝑅).

(O1)

The surrounding operators remain unchanged because both implementations satisfy the same output schema. The rewrite is beneficial when 𝑦 is less expensive to process and contains sufficient evidence for the task. A materialized representation, such as a transcript, can also be reused across multiple queries. O1 is not restricted to temporal localization. It can be applied to any semantic operation admitting schema-compatible implementations over different modalities, including classification, extraction, and summarization. For 𝑄 1 and 𝑄 2 , O1 replaces the video eventlocalization function with a transcript event-localization function. We call the resulting alternative the transcript-only query.

(vehicle_id, attributes, timeline, match_score).

4.2

O2: Cross-Modal Temporal Candidate Pushdown

A semantic Reduce–Unnest query places all synchronized input videos in one group and invokes a semantic function Sem_Associate. The function discovers vehicles in each video, determines which observations correspond to the same physical vehicle across cameras, and returns a collection 𝐴 of cross-camera association records.

Even when a source-aligned representation cannot answer a semantic query directly, it may identify the source-time regions containing the required evidence. O2 uses a source-aligned representation to restrict the temporal extent processed by a video semantic function. Conceptually, it transforms a source-aligned representation

Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, and Alvin Cheung

into candidate intervals, materializes those intervals as video clips, evaluates the clips, and combines their results by source video. Let 𝑓 𝑉 (𝑣, 𝑝) denote a semantic function applied to video 𝑣 under task description 𝑝. A source-time interval is a temporal witness for a result produced by 𝑓 𝑉 when the evidence within that interval is sufficient to establish the result. O2 applies when the relevant results have bounded temporal witnesses and results obtained from separate windows can be reconciled into the function’s original output schema. It does not apply when the answer depends on the entire video, such as determining that an event never occurs. Candidate construction. Let 𝑅(source_id, 𝑣, 𝑥) contain one tuple per source video, where 𝑥 is a representation whose timestamps are expressed in the source time of 𝑣. The recall-oriented semantic function cand(𝑥, 𝑝) returns a collection 𝐶 of candidate intervals in source time. In our prototype, Concord implements cand using an MLLM prompt conditioned on 𝑝. Concord constructs the materialized candidate clips as follows:   𝑅𝐶 = Unnest𝐶 Mapcand(𝑥,𝑝 )→𝐶 (𝑅) ,   𝑅𝐼 = Resolveoverlap Mapwindow(𝑐 )→(𝑠,𝑒 ) (𝑅𝐶 ) , 𝑅𝑊 = View𝑣,𝑠,𝑒→𝑣𝑐 (𝑅𝐼 ). Unnest produces one tuple per candidate 𝑐. The deterministic function window adds temporal context and returns source-time boundaries (𝑠, 𝑒). Resolve coalesces overlapping windows from the same source, and View materializes every remaining interval as a standalone clip 𝑣𝑐 . By its tuple-wise Map semantics, View preserves the input attributes while adding 𝑣𝑐 . Every tuple in 𝑅𝑊 therefore retains its source_id and source-time boundaries (𝑠, 𝑒). Rewrite. The video semantic function is applied independently to each materialized clip, and the resulting clip-level records are grouped by source video:   𝑅𝐴 = Reducesource_id,reconcile 𝑓 →𝐴 Map 𝑓 𝑉 (𝑣𝑐 ,𝑝 )→𝐴𝑐 (𝑅𝑊 ) . The inner Map applies 𝑓 𝑉 to each materialized clip and stores its clip-level result in 𝐴𝑐 . Because one source video may produce several clips, the outer Reduce groups these results by source_id and applies reconcile 𝑓 . The reconciliation function combines the clip-level results and returns one value 𝐴 satisfying the output schema of full-video execution. O2 is therefore the rewrite Map 𝑓 𝑉 (𝑣,𝑝 )→𝐴 (𝑅) ⇝ 𝑅𝐴 .

(O2)

Both sides associate one result 𝐴 with each source_id; therefore, the operators following the rewritten expression remain unchanged. Example: event localization. For 𝑄 1 and 𝑄 2 , 𝑥 is the timestamped transcript, 𝑓 𝑉 is the video event localizer, and 𝐴 is the event collection 𝐸. The candidate function identifies transcript-grounded regions that may contain the requested event. The same video localizer used by the video-only query is then applied to each materialized clip, producing clip-relative event predictions 𝐸𝑐 . In this instance, reconcile 𝑓 converts each prediction from clip-relative time to source time and coalesces duplicate predictions. The unchanged downstream Unnest emits one tuple per event. This rewrite produces the transcript-to-video query evaluated in §5.

O2 is a cross-modal analogue of selection pushdown: it restricts the input to a cost-dominant semantic operation using evidence from a source-aligned representation. Because the retained intervals are generated approximately, the rewrite is characterized by two quantities. Candidate recall is the fraction of reference results whose temporal witnesses are covered by at least one retained window. Selectivity is the union duration of the retained windows divided by the duration of the source video. O2 is most effective when candidate recall is high and selectivity is low.

4.3

O3: Semantic Reduce–Unnest to Detect–Track–Join

For a query that associates the same physical entity across a pair of synchronized video sources, a semantic Reduce–Unnest query can pass both videos to one MLLM-backed function. The MLLM must simultaneously discover entities, maintain their within-video identities, associate entities across the two sources, and format the resulting associations. Although concise, this query provides no inspectable intermediate records. The Detect–Track–Join query decomposes the association into explicit operators and intermediate relations for detection, tracking, candidate matching, conflict resolution, and trajectory construction: 𝑅𝐷 = Detect𝑣,𝐶→𝐷 (𝑅),    𝑅𝑇 = Unnest𝑇 Track𝐷 ′ →𝑇 Mapembed◦NMS(𝐷 )→𝐷 ′ (𝑅𝐷 ) , 𝑅 𝐽 = Resolveone_to_one (Join𝑝,𝑠 (𝑅𝑇 , 𝑅𝑇 )), 𝑅𝐴 = Mapto_association→𝑎𝑠𝑠𝑜𝑐 (𝑅 𝐽 ). where 𝑎𝑠𝑠𝑜𝑐 = (vehicle_id, attributes, timeline, match_score). O3 is the rewrite:  Unnest𝐴 Reduce ∅,Sem_Associate→𝐴 (𝑅) ⇝ 𝑅𝐴 .

(O3)

Here, 𝑅 contains videos from two synchronized sources, 𝐶 is the set of relevant entity classes, and 𝑅𝐷 contains frame-level detections. Within Map, non-maximum suppression removes redundant boxes, while embed extracts an appearance vector from each entity crop, producing the enriched detection collection 𝐷 ′ . Track links detections across frames into tracks 𝑇 , and Unnest produces the relation 𝑅𝑇 with one tuple per within-video track. Join self-joins 𝑅𝑇 . Its predicate 𝑝 admits only tracks from different sources with compatible motion and imposes a canonical ordering on source pairs, thereby removing same-source and symmetric duplicate pairs. The scoring function 𝑠 ranks the remaining candidates using signals such as appearance similarity. Because one track may appear in multiple candidates, Resolve𝑜𝑛𝑒_𝑡𝑜_𝑜𝑛𝑒 greedily processes candidates in descending score order and retains a pair only if neither track has already been matched within that source pair. Finally, to_association maps each resolved match to exactly one tuple containing (vehicle_id, attributes, timeline, match_score); intermediate track and join fields are not returned. This rewrite replaces a monolithic MLLM operation with inspectable relations for detections, tracks, candidate pairs, resolved matches, and associations. These intermediates expose failure points and enable component-level optimization while eliminating MLLM

Concord: A Video Relational Algebra for Cross-Modal Query Optimization

calls from the rewritten query. Its quality depends on entity coverage—the fraction of reference vehicles represented by usable tracks— and pair coverage—the fraction whose correct cross-source match survives candidate pruning and one-to-one resolution.

5

Evaluation

We have built a prototype of Concord and evaluate the effectiveness of our rewrites using the queries defined above. We evaluate 𝑄 1 and 𝑄 2 using the video-only query, the transcript-only query (O1), and the transcript-to-video query (O2). We evaluate 𝑄 3 using the semantic Reduce–Unnest query and the Detect–Track–Join query (O3). For each event-localization workload, the video-only and transcript-to-video queries use identical task-specific videolocalization instructions and output schema. Soccer and lectures use different prompts. The shared video-localization stage receives a complete input video in the video-only query and transcriptselected materialized clips in the transcript-to-video query. Datasets. Soccer contains three SoccerNet matches with English commentary [5], represented as six half-match videos totaling 4.59 hours at 224p, with ten labeled goals. Lectures contains three 360p videos [14], totaling 3.92 hours and four manually labeled events. Highway uses the first five seconds of two synchronized camera feeds from the I24V multi-camera highway dataset [6]. We manually adjudicated a cross-camera reference containing 18 vehicles. Queries and prompts were frozen before evaluation and receive no reference labels. Runtime methodology. We report cold-start end-to-end query execution time. For the video-only query, this includes transferring each complete input video to the MLLM provider and performing video event localization. For the transcript-only query, it includes text-based event localization over the materialized transcript. For the transcript-to-video query, it includes applying the candidate function to the materialized transcript to select source-time windows, materializing the corresponding clips, transferring those clips to the MLLM provider, and performing video event localization. Transcripts are treated as materialized inputs generated offline (e.g., from YouTube), so their one-time construction cost is excluded. No previously uploaded media is reused in the reported measurements. Models and implementation. We use Gemini 3.1 Flash-Lite [7] for all MLLM-backed operators, including video event localization, transcript candidate generation, and the semantic cross-camera association baseline. We report token usage and estimate cost using the provider’s pricing at the time of evaluation. The Detect–Track– Join query executes locally on a CPU without GPU acceleration.

5.1

Soccer Goal Localization

Predictions are matched one-to-one with reference goals. A prediction is correct when its absolute timestamp error relative to the matched reference goal is at most 30 seconds. We additionally evaluate tolerances of 5, 10, and 60 seconds to measure sensitivity to temporal precision. Table 2 reports the results. The transcript candidate function retains 879 of 16,522 seconds of source video, corresponding to 5.32% selectivity. Its 13 candidate windows cover all ten reference goals, yielding 100% candidate recall.

At the primary 30-second tolerance, the transcript-to-video query improves F1 from .818 for video-only execution to .952. The transcriptonly query also achieves .952 F1 at this tolerance, but produces less precise timestamps: among matched goals, its mean and median absolute errors are 4.274 and 2.011 seconds, compared with .348 and .181 seconds for transcript-to-video. The F1 results at tighter tolerances confirm this difference. At tolerances of 5, 10, 30, and 60 seconds, transcript-to-video achieves .952 F1 throughout, whereas transcript-only achieves .762, .857, .952, and .952, respectively; video-only remains at .818 across all four tolerances. The video stage therefore refines transcript-derived candidate regions into more precise goal timestamps. Relative to video-only execution, the transcript-to-video query reduces total MLLM tokens by 85.6%, estimated MLLM cost by 85.0%, and cold-start wall time by 80.9%, from 712.3 to 135.8 seconds. Thus, after temporal pushdown is coupled with clip materialization, the rewrite improves both model cost and end-to-end latency while preserving candidate recall.

5.2

Lecture Event Localization

Video is presented to the MLLM at one frame per second. Predicted intervals are matched one-to-one with reference intervals using temporal intersection over union (tIoU), defined as the duration of their intersection divided by the duration of their union. A prediction is counted as correct at the primary threshold when its tIoU with the matched reference interval is at least .3. We additionally evaluate thresholds of .1 and .5 to measure sensitivity to temporal boundary accuracy. Table 2 reports the results at the primary tIoU threshold. The transcript candidate function retains 349 of 14,115 seconds of source video, i.e., 2.47% selectivity. Its three candidate windows fully cover all four reference events, yielding 100% candidate recall. At tIoU ≥ .3, the transcript-to-video query improves F1 from .500 for video-only execution to .889, with precision .800 and recall 1.000. Relative to video-only execution, it reduces total MLLM tokens by 87.4%, estimated MLLM cost by 87.2%, and cold-start query execution time by 74.6%, from 396.2 to 100.5 seconds. The improvement persists across temporal-overlap thresholds. At tIoU thresholds .1, .3, and .5, transcript-to-video achieves F1 scores of .889, .889, and .444, respectively. The corresponding videoonly scores are .750, .500, and .250, while the transcript-only scores are .857, .286, and .000. Video refinement therefore improves temporal localization rather than merely detecting whether the requested demonstration is present. The transcript-only query reaches only .286 F1 at the primary threshold because transcript evidence often identifies a broad semantic neighborhood rather than the number and precise boundaries of the visible events. For the query that localizes successful soap-bubble demonstrations, one transcript-only prediction spans approximately 57 seconds and covers both six-second reference events instead of returning a separate interval for each. This behavior differs from commentary displacement in soccer. A soccer goal is a point event whose verbal description may precede or follow the visible goal. In the lecture workload, the transcript identifies the relevant demonstration but may not determine its visual boundaries or distinguish multiple occurrences within the

Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, and Alvin Cheung

Table 2: Event-localization results at the primary thresholds (Soccer: 30 seconds; Lectures: tIoU ≥ .3). “Video” is the fraction of source duration reaching the MLLM; “MLLM Tokens” include both input and output tokens. workload

query

video (%)

precision

recall

F1

MLLM tokens

MLLM cost

query time

Soccer

video-only transcript-only transcript→video

100.00 0.00 5.32

.750 .909 .909

.900 1.000 1.000

.818 .952 .952

1.506M 0.132M 0.217M

$0.378 $0.034 $0.057

712.3s 6.3s 135.8s

Lectures

video-only transcript-only transcript→video

100.00 0.00 2.47

.500 .333 .800

.500 .250 1.000

.500 .286 .889

1.285M 0.128M 0.161M

$0.322 $0.032 $0.041

396.2s 3.6s 100.5s

Table 3: Highway cross-camera join over two five-second clips and 18 reference vehicles. query semantic Reduce→Unnest Detect→Track→Join

preds

TP/FP/FN

precision

recall

F1

timeline IoU

attribute exact

wall time (s)

MLLM calls

MLLM tokens

4 14

4/0/14 13/1/5

1.000 .929

.222 .722

.364 .813

.576 .762

.583 .538

52.724 99.559

1 0

1,515 0

same discussion interval. Transcript candidate localization and video event localization therefore perform complementary roles in the transcript-to-video query.

5.3

Cross-Camera Vehicle Trajectories

The semantic Reduce–Unnest baseline sends both clips to Gemini 3.1 Flash-Lite in one MLLM-backed Reduce and unnests the returned cross-camera trajectories. The Detect–Track–Join query uses YOLOE detections [19], motion-aware tracking, ResNet50 appearance embeddings [8], pair pruning, and greedy one-to-one matching. We score each prediction–reference pair as the fixed weighted sum of 0.7 tIoU + 0.3 Attr (coefficients chosen a priori to favor temporal overlap over noisy cross-camera attributes; not tuned on the reported runs), then assign matches greedily oneto-one. Here tIoU is the temporal IoU from earlier, averaged over shared cameras, and Attr is soft agreement on class, color, and subtype. A pair trajectory is accepted if this score is at least 0.70; the Join itself keeps candidate pairs with appearance cosine similarity at least 0.40. Table 3 summarizes the cross-camera association results. The semantic Reduce–Unnest baseline is precise for the four trajectories it returns, but its recall is .222. The Detect–Track–Join query produces 14 trajectories, matches 13 reference vehicles, and raises F1 from .364 to .813 while improving mean timeline IoU from .576 to .762. It also eliminates MLLM calls by replacing whole-clip inference with explicit Detect, Track, and Join operators. On these five-second clips, however, the Detect–Track–Join query is slower than the semantic Reduce–Unnest query (99.6 s vs. 52.7 s): the rewrite trades MLLM usage for higher coverage, not end-to-end latency on this short sample.

6

Conclusion

We described Concord, a multimodal video query processing system built on VRA over video inputs. By exposing different modalities within a single algebra, VRA allows users to express queries and also Concord to optimize them. Our experiments show that Concord’s optimizations can significantly reduce cost and query execution time while improving accuracy, demonstrating their promise for semantic video query processing.

References [1] Bastani et al. 2020. MIRIS: Fast Object Track Queries in Video. In SIGMOD. 1907–1921. [2] Du et al. 2025. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. In Findings of EMNLP. 23281–23298. [3] Bang et al. 2023. Seiden: Revisiting Query Processing in Video Database Systems. Proc. VLDB Endow. 16, 9 (2023), 2289–2301. [4] Wei et al. 2026. Multi-Objective Agentic Rewrites for Unstructured Data Processing. arXiv:2512.02289 [cs.DB] https://arxiv.org/abs/2512.02289 [5] Giancola et al. 2018. SoccerNet: A Scalable Dataset for Action Spotting in Soccer Videos. In CVPR Workshops. [6] Gloudemans et al. 2024. So You Think You Can Track?. In WACV. 4528–4538. [7] Google. 2026. Gemini 3.1 Flash-Lite. https://ai.google.dev/gemini-api/docs/ models/gemini-3.1-flash-lite. Accessed August 4, 2026. [8] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 770–778. doi:10.1109/CVPR.2016.90 [9] Saehan Jo and Immanuel Trummer. 2024. ThalamusDB: Approximate Query Processing on Multi-Modal Data. Proc. ACM Manag. Data 2, 3 (2024). [10] Kang et al. 2017. NoScope: Optimizing Neural Network Queries over Video at Scale. Proc. VLDB Endow. 10, 11 (2017). [11] Daniel Kang et al. 2022. TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data. In SIGMOD. 1934–1947. [12] Kittivorawong et al. 2024. Spatialyze: A Geospatial Video Analytics System with Spatial-Aware Optimizations. Proc. VLDB Endow. 17, 9 (2024), 2136–2148. [13] Liu et al. 2025. Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing. In CIDR. [14] MIT OpenCourseWare. 2016. Physics III: Vibrations and Waves. https://ocw. mit.edu/courses/8-03sc-physics-iii-vibrations-and-waves-fall-2016/. Fall 2016 course materials. [15] Patel et al. 2025. Semantic Operators and Their Optimization: Enabling LLMBased Data Processing with Accuracy Guarantees in LOTUS. Proc. VLDB Endow. 18, 11 (2025). [16] Romero et al. 2022. Optimizing Video Analytics with Declarative Model Relationships. Proc. VLDB Endow. 16, 3 (2022), 447–460. [17] Shankar et al. 2025. DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing. Proc. VLDB Endow. (2025). [18] Matthias Urban and Carsten Binnig. 2024. Caesura: Language Models as MultiModal Query Planners. In CIDR. [19] Ao Wang, Lihao Liu, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. 2025. YOLOE: Real-Time Seeing Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 24591–24602. [20] Xiao et al. 2026. KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration. In CIDR. [21] Zhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, and Umakishore Ramachandran. 2022. EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized Views. In Proceedings of the 2022 International Conference on Management of Data (Philadelphia, PA, USA) (SIGMOD ’22). Association for Computing Machinery, New York, NY, USA, 602–616. doi:10.1145/3514221.3526142 [22] Zhang et al. 2025. Self-Enhancing Video Data Management System for Compositional Events with Large Language Models. Proc. Manag. Data 3, 3 (2025).

Related documents

Record · ID 673629 · SHA-256 0787308e390b5856
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.