IEEE TRANSACTIONS ON IMAGE PROCESSING
1
Event-Aware Instructed Assistant for Referring Video Segmentation
arXiv:2606.26994v1 [cs.CV] 25 Jun 2026
Jinyu Liu, Henghui Ding, Shuting He, Yu-Gang Jiang IEEE Fellow
Abstract—Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple distinct events. Under such a mechanism, the model needs to directly understand all the complex content in the video and text, which can easily lead to confusion and hallucinations. To address this issue, we propose to decompose a video to a set of simple events by learnable Event Query, and understand complex video content in an event-by-event, easy-to-understand manner. This is based on the observation that natural language expressions often divide a video into distinct, text-related segments, each representing a separate event within a compound event. We introduce EVIS, an Event-Aware Video Instructed Segmentation Assistant, which utilizes text-guided Event Queries to partition a video into simple events, extracting event-aware visual-text features to achieve a hierarchical understanding of the video. Additionally, we propose Object-Pixel-Hybrid Learning, which enables the MLLMs to track targets in long-term videos by integrating fine-grained pixel features with prior object queries. Extensive experimental results on 5 public benchmarks demonstrate EVIS’s strong performance in addressing the referring video segmentation task.
"Panda moves up then descends, and finally disappears. "
"Panda moves up then descends, and finally disappears. "
"Panda moves up then descends, and finally disappears. "
Index Terms—Referring Video Object Segmentation, EventAware, Multi-Modal Learning.
I. I NTRODUCTION
"Panda moves up then descends, and finally disappears. "
EFERRING Video Object Segmentation (RVOS) [1]– [4] is an emerging and challenging task focused on segmenting target objects according to the given natural language description throughout a video. It has a wide range of applications in the real world, such as embodied perception and video editing. Recent datasets like MeViS [1] emphasize the temporal motion properties of videos, which are crucial in this field. With the rapid development of multi-modal Large Language Models (MLLMs) [5], [6], LLM-based methods like LISA [7] and GSVA [8] are introduced to referring image segmentation, showcasing the ability of MLLMs to localize and segment objects through advanced reasoning and understanding. Following LISA [7], VISA [9] and VideoLISA [10] extend the use of MLLMs to the referring video domain. However, existing approaches [1], [2], [9], [10], [12]–[14] often treat a video as a single event composed of multiple images, overlooking the fact that videos with temporal motion properties [1] typically contain multiple distinct events. Under such a mechanism, the model needs to directly understand all
Fig. 1. Illustration of the Event Taxonomy [11]. When focusing on different expression parts, trajectories of related objects in the video are various. Any length of trajectory for each object can be abstracted as a simple event. These simple events make up the compound event throughout the video.
R
Jinyu Liu and Henghui Ding are with the Institute of Big Data, College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China (e-mail: [email protected], [email protected]). Yu-Gang Jiang is the Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China (e-mail: [email protected]). Shuting He is with Shanghai University of Finance and Economics, China (e-mail: [email protected]). Corresponding authors: Henghui Ding and Yu-Gang Jiang This project was supported by the National Natural Science Foundation of China (NSFC) under Grant No. 62472104. Henghui Ding was supported by Xiaomi Young Scholars Program.
the complex content in the video and text, which can easily lead to confusion and hallucinations, especially in MLLM-based methods. Even for humans, understanding complex scenarios often involves breaking down the whole into its parts and then integrating these parts into a cohesive understanding [15], [16]. When processing complex information, such as videos, the human cognitive system decomposes the content into manageable segments, gradually assembling them to form an understanding of the complete scenario. Shipley et al. [11] propose the Event Taxonomy, hypothesizing that a compound event consists of the combined occurrence of two or more simple events, which occur when objects change or interact. Inspired by the Event Taxonomy [11], we propose to decompose the video to multiple simple events by a group of learnable Event Queries and understand complex video scenarios in an event-by-event manner. We have observed that natural language expressions often segment a video into distinct, text-related parts, each representing a separate event within a compound structure. As shown in Fig. 1, when considering the partial phrase “panda moves up”, two object trajectories from the earlier section align with it. Similarly, the phrase “panda descends” corresponds to the latter part of this case. However, for “panda moves up then descends”, all objects meet the
RVOS Segmenter "First bus stops, then starts moving."
"First bus stops, then starts moving."
a) Previous Methods
Event-Aware Module
...
Multiple Simple Events
RVOS Segmenter ...
...
...
"First bus."
...
...
...
...
...
...
Single Compound Event
2
...
...
IEEE TRANSACTIONS ON IMAGE PROCESSING
"First bus stops."
...
b) EVIS (ours)
Fig. 2. Comparison of previous methods and ours. a) Previous methods often struggle to directly understand the complex content within a single video. b) EVIS (ours) innovatively shifts from single compound-event video paradigm to hierarchical simple-event analysis.
criteria: the blue segments in the 3rd row of Fig. 1. This Merging Module (EAFM) combined with Event-Intra illustrates the importance of distinguishing between simple Attention and Event-Inter Attention to achieve compreevents within a video sequence for accurate interpretation hensive video-text alignment. of the full video content in relation to the given linguistic • We demonstrate impressive performance across five description. referring video segmentation datasets, notably achieving a With Event Query to decompose the given video into simple 46.8% J &F on the challenging MeViS dataset, validating events, we introduce EVIS, Event-Aware Video Instructed the effectiveness of the proposed approach. Segmentation Assistant, to address the challenging referring II. R ELATED W ORK video segmentation. The proposed EVIS leverages text-guided Event Queries to partition a video into distinct simple events, Referring Video Segmentation focuses on segmenting extracting event-aware visual-text features to enable a hierar- targets within a video according to the given expression. The chical understanding of the video content. The event-centric field has been significantly accelerated by several benchmark design distinguishes EVIS from previous RVOS methods, as datasets, e.g., A2D-Sentences [17], Ref-DAVIS17 [4], Refshown in Fig. 2. To capture events across frames throughout YouTube-VOS [3], and MeViS [1]. Some previous methods the video, EVIS incorporates an advanced Event-Aware Frame directly extend referring image segmentation to video scenes. Merging Module (EAFM), which uses text-guided Event For example, Khoreva et al. [4] incorporates the referring image Queries to merge objects from multiple frames into a set of segmentation methods MAttNet [18] to segment each frames specific events. Additionally, Event-Intra Attention and Event- and use post-processing to maintain temporal coherence. Other Inter Attention are integrated to the EAFM. The Event-Intra studies, such as ReferFormer [12] and MTTR [13], introduce a Attention enhances the model’s ability to capture the short-term DETR-like architecture for referring video object segmentation, spatial-temporal dynamics in each event, while the Event-Inter streamlining the segmentation process and achieving notable Attention supports long-term learning across multiple events. results. More recently, MeViS [1] dataset, based on the These blocks not only facilitate fine-grained learning within complex video object segmentation dataset MOSE [19], is events but also maintain a coherent understanding throughout developed to emphasize the importance of motion expressions the entire video, ensuring robust interactive alignment between and address current methods’ limitations in understanding video content and textual expressions. However, pixel features motion information. Some previous methods, such as VISA [9] tend to be highly redundant, while object features, though fewer and VideoLISA [10], leverage MLLMs for this task: VISA [9] in number, capture more abstract and meaningful information. extends the referring image segmentation task to the video To tackle this issue, we introduce Object-Pixel-Hybrid Learning, domain by propagating the first frame to the rest video frames which encourages the simultaneous learning of prior object through an object tracking method [20], while VideoLISA [10] queries and pixel-level features by a single [SEG] token for proposes a sparse dense sampling strategy to reduce the each video, achieving a multi-level feature interaction. This number of tokens, preserving temporal dynamics by applying learning strategy guarantees long-term object tracking while global average pooling on sparse frames to reduce them to maintaining computational efficiency. Overall, this simple-to- a lower resolution. Recently, Glus [21] adopts Global and compound learning procedure allows EVIS to preserve object Local consistency, a set of sparse context frames provides coherence in complex scenes. global information, while a stream of continuous query frames In summary, our main contributions are as follows: conducts local object tracking. Villa [21] proposes Key Segment Extractor to select the most query-relevant segments, and • We propose EVIS, an Event-Aware Video Instructed uses Context Synthesizer to aggregate text-related visual cues Segmentation Assistant that leverages the Event Query and for addressing long durations, multiple objects, rapid motion, EAFM module for hierarchical multi-modal understanding and heavy occlusions. Veason-R1 [22] explores the effect of in referring video segmentation. Group Relative Policy Optimization (GRPO) [23] for video • We design Event Query to decompose a video to a series of reasoning segmentation. In this study, we focus on enhancing simple events, promoting the model to understand complex the understanding of temporal motion cues in an event-by-event, video scenarios in an event-by-event, easy-to-understand easy-to-understand manner. manner. Referring Image Segmentation involves segmenting targets • To capture events, we develop an Event-Aware Frame within images based on text descriptions. Methods are mainly
IEEE TRANSACTIONS ON IMAGE PROCESSING
3
"First bus stops, then starts moving."
Multi-Modal Large Language Model
...
LoRA
Object Query
...
... ...
Mask Decoder
"Sure. It is
... ."
Detector
Visual Encoder
...
Visual Tokenizer
...
...
Event Query
...
...
...
...
...
Cross-Attention
Object-Pixel-Hybrid
xL
Event-Aware Frame Merging Module
Temporal
Fig. 3. Overview of the proposed Event-Aware Video Instructed Segmentation Assistant (EVIS). EVIS employs the event query and Event-Aware Frame Merging Module (EAFM) to learn hierarchical video features in an event-by-event manner. First, we decouple the visual tokens to pixel tokens Vp and object tokens Vo , where tokens are split along the temporal dimension. Using object queries Qo generated by the detector, EAFM module is applied to comprehend various objects, guided by the event query. Notably, the event query cross-attends with the text embeddings at the beginning of each EAFM block, ensuring full interactions between video and text and preventing text information from being lost. Finally, the pixel tokens Vp and updated event-level object queries Q̂o are fed into the MLLMs to generate the [SEG] token and predict the final mask, facilitating comprehensive video understanding.
categorized into one-stage [24]–[27] that perform end-to-end reducing token count. Recent studies [5], [36]–[39] have predictions and two-stage [18], [28], [29] that decompose the further integrated region-level image grounding and pixel-level task into instance segmentation and text-instance matching. Hu understanding into MLLMs. For example, InternVL [5] aligns et al. [25] exemplify the one-stage paradigm by integration of the representation of the scaled-up vision encoder with LLMs visual and linguistic features for mask prediction. MAttNet [18] and achieves surprising performance on image grounding task. demonstrates the two-stage way by initially employing instance segmentation, followed by text-guided object selection. The III. M ETHOD emergence of Transformer [30] has catalyzed significant advancement in referring image segmentation. Ding et al. [24] A. Preliminaries pioneers the application of Transformers in this domain through V : Visual tokens of the input video. Vp : Pixel tokens, input their Vision-Language Transformer (VLT). This breakthrough for the visual encoder. Vo : Object tokens, input for the detector. has precipitated the development of numerous Transformer- Qo : Object queries. Q̂o : Object queries updated by the EAFM based methods. Notable examples include LISA [7] and module. Qo′ : Object queries from all events. qo : Object queries GSVA [8]. LISA [7] introduces a specialized [SEG] token to from one frame. Fs : Sentence embedding. Qe : Event queries. interface with segmentation mask decoders such as SAM [31], Qg : Global object queries. k: Number of top event queries enabling MLLMs to generate precise masks. Building upon selected in the Frame Merging Block. A: Assign attention this foundation, GSVA [8] extends the functionality of [SEG] in the Frame Merging Block, which is not differentiable. Â: token while introducing [REJ] token to support generalized Assign attention, preserving gradient of A and is differentiable. referring expression segmentation (GRES) [32]. VRS-HQ [33] Ã: Attention between Ê and Q̂o . Wo : Learnable transformation proposes Temporal Dynamic Aggregation and Token-driven matrices for object queries in the Gumbel Softmax operation. Keyframe Selection modules achieving strong performance. We : Learnable transformation matrices for event queries in Multi-Modal Large Language Model. The rapid develop- the Gumbel Softmax operation. Wo′ : Learnable transformation ment of large language models (LLMs) motivates research into matrices for object queries in the process of global object extending their capabilities to the visual domain, breeding the queries generation. We′ : Learnable transformation matrices multi-modal large language models (MLLMs) [5], [6], [34]– for event queries in the process of global object queries [40]. Previous work like LLaVA [34], BLIP-2 [35], MiniGPT- generation. E: Object queries in single event. Ê: Object queries 4 [6] and Qwen-VL [37], empowering the MLLMs with ability in single event updated by the Event-Intra Attention. L: Number to handle complex visual tasks. In video scenarios, learning of stacking layers of the EAFM module. fp : Pixel features temporal information becomes crucial. A straightforward way generated by the ViT of MLLM. Nl : Number of frames in is to concatenate tokens from multiple frames, , though this is single video. No : Number of object queries in single frame. constrained by computational limits. To address this issue, No′ : Number of frames in single event. Nh : Number of event some works [40]–[43] employ token merging or pooling queries. i: Index of object query in single frame. l: Index strategies, which reduce token numbers at the cost of detail loss. of frame in single video. h: Index of event queries in single Other works such as BLIP-2 [35] use Q-Former architecture video. t: Index of set of object queries in single event. U : to extract abstract features, retaining rich information while Hybrid unit. To : Number of object queries in a hybrid unit.
IEEE TRANSACTIONS ON IMAGE PROCESSING
Tu : Number of pixel features in a hybrid unit. m: Ground truth segmentation mask. m̂: Predicted segmentation mask by the model. M: Event attention mask for object queries in the Event-Intra Attention. yt : Ground truth text answer. ŷt : Text answer generated by the LLM. Ollm : [SEG] token embedding of the outputs generated by the LLM. λm : Weight of mask loss Lm . λt : Weight of text generation loss Lt . λbce : Weight of BCE loss. λdice : Weight of DICE loss. An overview of the proposed EVIS approach is shown in Fig. 3. First, we decouple the visual tokens V into pixel tokens Vp and object tokens Vo using a visual tokenizer. To ensure dimensional consistency for input into the MLLM, we use a projection layer to adjust the feature dimension of Vo . Mask2Former [44] is used as the detector to extract object queries Qo of potential candidate objects. Text features are extracted by RoBERTa [45]. The event query Qe embeds the textual features by cross-attending to the sentence embedding Fs before being inputted into the Event-Aware Frame Merging Module (EAFM). EAFM is applied to the object queries Qo to progressively gather spatial-temporal information, generating event-level object queries Q̂o under the guidance of the event query Qe . EAFM module is stacking with L layers. Next, text tokens Vt , pixel tokens Vp , and object queries Q̂o are fed into the MLLMs to produce the [SEG] token. Finally, the last-layer embedding of the [SEG] token is decoded into a segmentation mask via the mask decoder. B. Easy-to-Understand: Event Query
4
...
Event-Inter Attention
Event-Intra Attention
Frame Merging Block ...
Event Query
Object Query
Fig. 4. Event-Aware Frame Merging Module (EAFM). The EAFM module effectively comprehends various objects in an event-by-event manner, guiding the MLLMs to capture event-intra and event-inter information in a video.
Merging Module. First, the frame merging block group object queries into distinct simple events. Subsequently, an eventintra attention mechanism is applied to capture fine-grained spatial-temporal information within each event. Following this, cross-attention is performed between the event query and the refined event-intra object queries to produce a global query for each event. These global object queries then undergo eventinter attention, allowing for the extraction of long-term object trajectory features, which generate the event-inter object queries. Finally, the resulting sequence of event-level object queries is computed as a weighted sum of event-intra and event-inter object queries.
Existing methods [1], [2], [9], [10], [12]–[14] for referring video segmentation often treat a video as a single event. These methods are not specifically designed for the complex and varied scenes within videos and primarily rely on referring image segmentation paradigms to perform frame-by-frame target object segmentation. We argue that the model can better understand complex, compound events in a video if it first comprehends the simple events within them. Shipley et al. [11] propose the Event Taxonomy, hypothesizing that a compound event consists of the combined occurrence of two or more simple events, Frame Merging Block groups object queries into distinct occurring when objects change or interact. Inspired by this events based on embedding similarity using differentiable toptaxonomy, we design the learnable Event Query to decompose k assignment. Rather than forwarding image tokens from all the video into simple events, based on the observation that Nl frames, we use object queries generated by the detector, natural language expressions often divide a video into distinct, Mask2Former [44], to ensure efficient interaction. As shown text-related segments, each representing a separate event within in Fig. 5, the Frame Merging Block receives the learned event a compound event. Notably, our method differs fundamentally query and object query as inputs. It merges all object queries from object trajectory re-grouping in both representation associated with the same event query into a new sequence of and supervision. Instead of hard-assigning objects to disjoint object queries, based on their similarity within the embedding temporal segments, we model events as dynamic, semantic space. attention aggregations over object queries, which allows the Formally, suppose there are Nh events in a video, each same object to simultaneously participate in multiple events. h indexed by h and with a set of learnable event queries {Qhe }N h=1 . The Event Query encourages the model to understand complex l Nl We denote the object queries from all the frames as {Qo }l=1 , video scenarios on an event-by-event, easy-to-understand basis. o the l-th frame’s object queries Qlo = {qoi }N i=1 , where Nl is the frame number and No is the object query number in a single h l Nl h C. Event-Aware Frame Merging Module frame. We simplify {Qhe }N h=1 to {Qe } and similarly{Qo }l=1 l h EAFM aggregates frame-wise object queries into temporal to {Qo }. The association between the event queries {Qe } and events and models both intra-event dynamics and inter-event object queries {Qlo } is established through a similarity matrix relationships. Fig. 4 shows the proposed Event-Aware Frame A inspired by [46], computed using the Gumbel Softmax [47]
IEEE TRANSACTIONS ON IMAGE PROCESSING
5
operation:
Through the propagation of the Event-Intra Attention and EventInter Attention, we capture both internal and external event (1) information. In order to integrate the features across various temporal dimensions, we conduct event integration between the where the We and Wo serve as learnable transformation event-intra object queries and event-inter global object queries. matrices that project the event and object queries into a During each block calculation, we record the position of the shared embedding space, while the stochastic variables {γ} initial object query at each step and average the intra-event are drawn independently from the Gumbel(0, 1) distribution. and inter-event queries, mapping the updated results back to We compute the event to assign an object query to by selecting the original input locations. the top-k events. Since the top-k assignment operation is not differentiable, to address this, we use a straightforward trick D. Object-Pixel-Hybrid Learning in [48] to compute as: While models are becoming increasingly capable with  = top-k(A) + A − sg(A), (2) advancements in MLLMs, the number of model parameters is where sg denotes the stop gradient operator. This formulation also growing. VideoLISA [10] employs a sparse dense sampling enables  to maintain the top-k discrete event assignments strategy to reduce the number of visual tokens, applying global while preserving gradient of A, ensuring the Frame Merging average pooling in sparse frames to lower their resolution. However, directly averaging high-dimensional features can Block remains end-to-end optimization. Event-Intra Attention. To capture the short-term spatial- cause information distortion and result in a loss of important temporal dynamics within each event via masked self-attention, details. To address this challenge, we propose Object-Pixel-Hybrid we propose the Event-Intra Attention block. This block applies Learning, which enables MLLMs to track targets in longa mask M to each event, ensuring that only the information term videos by integrating fine-grained pixel features with relevant to the specific event is considered, thereby isolating it contextual object queries. In a single frame, pixel features tend from external influences. We denote the object queries from Nh to be highly redundant, while object features, though fewer all events as Qo′ = {Eh }h=1 . Assuming that the h-th event ′ in number, capture more abstract and meaningful information. l +N ′ Eh = {Qto′ }t=l′ o begins at the l′ -th set of object queries, Additionally, object features can guide the language model where No′ is the number of object query sets in the current in filtering out redundant information from the pixel features. event. Event-intra Attention is calculated as: We define a hybrid unit composed of a single pixel feature ( n 1, if (Qm and multiple object queries, allowing for multi-level feature o′ ∈ Eh ) ∧ (Qo′ ∈ Eh ) Mh [m, n] = , (3) representation and interaction. Formally, we concatenate the −∞, otherwise pixel features fp = ViT(Vp ) with event-level object queries Eh QTo′ Êh = softmax Mh ⊙ √ Qo′ , (4) Q̂o , where ViT is the vision encoder of the MLLMs, and then C interlace them as input into the LLMs: where C is the number of channels, Mh is the event attention l+To U = Concat[fpl , Q̂l+1 ], (6) o , ..., Q̂o mask for the h-th event’s object queries Eh , and ⊙ denotes Ollm = LLMs(Concat[U1 , ..., UTu ]), (7) element-wise multiplication. Global Object Query aggregates intra-event object queries where To is the number of object queries in a hybrid unit, U into a compact event-level representation. Based on the object denotes the hybrid unit consisting of one single pixel feature queries updated by the Event-Intra Attention, we further fp and To object queries, Concat[ , ] is concatenation. Tu is the abstract the object queries into a global object query Qhg for number of hybrid units, and Ollm refers to the outputs of the the h-th event. This process is guided by the updated event LLMs. With this hybrid input, we aim to encourage the model query through cross-attention and a weighted sum operation: to learn features of varying dimensions, while mitigating the Pl′ +No′ negative influence of redundant pixel features. ′ ÃW E ′ o h Qhg = We′ t=l , (5) Pl′ +No′ à t=l′ where W ′ and W ′ are learned weights for the event and E. Training Objective n exp(We Qm e ·Wo Qo + γm ) Am,n = PNh , h n h=1 exp(We Qe ·Wo Qo + γh )
e
o
With Object-Pixel-Hybrid Learning, an apparent issue is the object queries to obtain merged features. Ã is the attention disparity between pixel and target features. To address this, we matrix between the updated event and object queries. Event-Inter Attention models long-range temporal dependen- select an equal number of pixel and object features to generate cies and object trajectories across different events. To capture masks for the loss calculation. Building upon the foundational long-term event trajectory features across all events, we further work LISA [7], the overall training objective L is formulated design an Event-Inter Attention block. The global object query as a weighted combination of the standard text generation loss generated by Event-Intra Attention is used as input to Event- Lt and segmentation mask loss Lm : Inter Attention, which is computed similarly to Event-Intra L = λt Lt + λm Lm , (8) Attention but without applying the event mask. Event Integration fuses intra-event and inter-event features where Lt is the auto-regressive cross-entropy (CE) loss for text and maps enriched representations back to original positions. generation, while Lm combines per-pixel binary cross-entropy
IEEE TRANSACTIONS ON IMAGE PROCESSING
6
TABLE I A BLATION STUDY OF OUR PROPOSED EVIS ON M E V I S DATASET. EAFM AND OPH REPRESENT E VENT-AWARE F RAME M ERGING M ODULE AND O BJECT-P IXEL -H YBRID L EARNING , RESPECTIVELY.
...
EAFM ✗ ✗ ✓ ✓
... Gumbel Softmax
J 36.5 38.1 42.1 43.7
F 42.3 44.3 47.7 49.9
TABLE II A BLATION STUDY OF E VENT-AWARE F RAME M ERGING (EAFM) MODULE ON M E V I S. FM, I NTRA , AND I NTER DENOTE F RAME M ERGING , E VENT-I NTRA ATTENTION , AND E VENT-I NTER ATTENTION , RESPECTIVELY.
... Event Query
J &F 39.4 41.2 44.9 46.8
OPH ✗ ✓ ✗ ✓
Object Query
Fig. 5. Frame Merging Block. We compute the event to assign an object query to by selecting the top-k Event Queries.
(BCE) loss and DICE loss [49], weighted by their respective coefficients λbce and λdice . Given the ground truth (ŷtxt , m̂) and model predictions (ytxt , m), Ltxt and Lm are defined as: Ltxt = CE(ŷtxt , ytxt ),
(9)
Lm = λbce BCE(m̂, m) + λdice DICE(m̂, m),
(10)
where ŷtxt , ytxt correspond to textual sequences and m̂, m denote binary segmentation masks.
FM ✗ ✓ ✓ ✓ ✓
Intra ✗ ✗ ✓ ✗ ✓
Inter ✗ ✗ ✗ ✓ ✓
J &F 41.2 45.8 46.1 46.5 46.8
J 38.1 43.1 43.3 43.5 43.7
F 44.3 48.5 48.9 49.5 49.9
TABLE III D IFFERENT NUMBER OF EVENT QUERIES IN THE EAFM MODULE . Nh 4 5 6 7
J &F 46.1 46.6 46.8 46.6
J 42.6 43.2 43.7 43.4
F 49.6 50.0 49.9 49.8
F. Re-implemented EVIS without LLM.
Decoder
Mask Feature
...
...
EAFM
annotations, and JHMDB-Sentences has 928 videos across 21 action categories. Evaluation Metrics. We use region similarity J , contour accuracy F, and their average J &F. For A2D-Sentences and JHMDB-Sentences, we use mAP, oIoU, and mIoU.
B. Implementation Details
...
"First bus stops, then starts moving." Mask2Former
...
...
...
Since the event-centric design is independent of the model’s reliance on LLMs, a natural question arises: can the EAFM module operate without LLM? To investigate this possibility, we build a simple baseline model by removing the LLM and SAM components to evaluate the capability of EAFM, as shown in Fig. 6. For experiments, we follow [2]. Tab. IX additionally compares the re-implemented EVIS with SOTA methods without LLMs.
Fig. 6. Model architecture of the re-implemented EVIS by removing LLM.
IV. E XPERIMENTS A. Datasets and Evaluation Metrics Datasets. The proposed EVIS approach is trained on a variety of segmentation datasets. For image datasets, we follow LISA [7]. For video datasets, we follow [10] to optimize the EAFM module for video-text dynamics. We evaluate across five RVOS benchmarks: MeViS [1], Ref-YouTubeVOS [3], Ref-DAVIS17 [4], A2D-Sentences [17], and JHMDBSentences [50]. MeViS is a recently established benchmark focusing on motion analysis and contains 2,006 videos with 28K annotations. Ref-YouTube-VOS is the most extensive RVOS dataset, containing 3,978 videos and 13K text annotations. Ref-DAVIS17 adds text descriptions to DAVIS17 [51] dataset. A2D-Sentences includes 3.7K videos with 6.6K action-specific
We implement our model using InternVL2 [5], a multimodal large language model built upon InternViT [52] and Qwen2 [53] with 1B parameters. The visual encoder and mask decoder are derived from SAM [31]. We perform joint training with both image and video datasets. In the pre-processing stage, we insert the original category name or referring expression from the dataset into a template. For example: “USER: <VIDEO> Can you segment {description} in this video? ASSISTANT: Sure, it is [SEG].”, where {description} serves as a placeholder for the specific text to fill. For the video data, we configure Tu = 4 and To = 8. For the EAFM module, we set L = 3, Nh = 6 and k = 2, respectively. In the case of image data, we replicate the images to create pseudo video sequences. We apply the AdamW [54] optimizer, setting the learning rate to 0.0012 and the weight decay to 0. The learning rate scheduler is WarmupDecayLR, with 100 warmup iterations. Both the weights for text generation loss λt and mask loss λm are set to 1.0, while the weights for BCE loss λbce and DICE loss λdice are set to 2.0 and 0.5, respectively. The per-device batch size is configured to 8. We use a total of 6,000 iterations to train the final model.
IEEE TRANSACTIONS ON IMAGE PROCESSING
7
C. Ablation Study We perform ablation studies on the challenging MeViS [1]. Module Effectiveness. We first conduct ablation experiments to evaluate the effectiveness of different aspects of the proposed EVIS. As shown in Tab. I, adding the Event-Aware Frame Merging Module (EAFM) yields a 5.5% huge improvement in J &F over the baseline, which is our re-implementation based on VideoLISA [10] integrated by InternVL2 [5]. The inclusion of EAFM strengthens the model’s capacity to comprehensively capture event-level video-text features, supporting improved multi-modal understanding. Subsequently, the Object-PixelHybrid Learning (OPH) is introduced to enable the MLLMs to track targets in long-term videos by integrating fine-grained pixel features with prior object queries, which contributed an additional 1.9% gain in J &F. Finally, EVIS achieve a stateof-the-art J &F of 46.8% on the challenging MeViS dataset, underscoring the efficacy of the proposed approach. Ablation of EAFM Module. We conduct ablation experiments on the effectiveness of EAFM module. As shown in Tab. II, incorporating the Frame Merging (FM) module results in a 4.6% improvement in J &F over the vanilla baseline. The FM module enables the model to split object queries into different simple events for afterward event-level understanding. Next, we introduce the Event-intra Attention module to capture fine-grained short-term spatial-temporal dynamics within each event. This addition boosts performance by 0.3% in J &F, highlighting the critical role of fine-grained eventintra understanding in referring video segmentation. Then we incorporate Event-inter Attention to further motivate the model learning object trajectory global representations between all of the events, which is essential for long-term understanding in videos. Inclusion of Event-inter Attention results in a 0.7% gain in J &F. Finally, when all components are integrated, EAFM module results in a performance improvement of 4.3%, which demonstrates the effectiveness of the proposed method. Number of Event Queries Nh . As mentioned above in Fig. 1, any length of trajectory for each object can be abstracted as a simple event. We utilize the event query to interact with object queries from a single video generating multiple simple events. The number of events is equal to the number of event queries. When a video is partitioned into a limited number of events, each event may encapsulate an overwhelming amount of information, making it difficult for the model to effectively process and interpret the content within an event. Conversely, when the video is segmented into a large number of smaller events, each individual event becomes more understandable, but the challenge shifts to synthesizing information across a substantial number of events. Consequently, it is crucial to determine an optimal event number that allows the model to strike a balance, facilitating both the comprehension of individual events and the integration of information across the entire sequence for optimal learning performance. The best result is achieved when Nh is set to 6, as shown in Tab. III. When Nh = 1, which means directly comprehending compound events, matching the setup in the second row of Tab. I (without EAFM but with OPH), still enables OPH for efficient learning. Influence of MLLM choice. As shown in Tab. IV. Using
TABLE IV A BLATION STUDY ON MLLM IMPACT. Method VideoLISA (n-frame) VideoLISA (SDS) EVIS VideoLISA (n-frame) VideoLISA (SDS) EVIS
Multi-Modal LLM LLaVA-Phi-3-3.8B LLaVA-Phi-3-3.8B LLaVA-Phi-3-3.8B InternVL2-1B InternVL2-1B InternVL2-1B
MeViS 43.2 44.4 46.3 39.4 40.9 46.8
Ref-Ytb-VOS 63.3 63.7 64.2 61.9 62.4 64.4
Ref-DAVIS17 68.6 68.8 69.3 65.7 66.4 68.8
TABLE V D IFFERENT NUMBER OF STACKING DEPTH L IN EAFM ON M E V I S [1]. L 1 2 3 4
J &F 43.1 46.4 46.8 46.8
J 39.8 43.1 43.7 43.8
F 46.4 49.7 49.9 49.8
TABLE VI D IFFERENT V I T BACKBONE OF SAM [31] ON M E V I S [1]. Visual Encoder SAM-ViT-B SAM-ViT-L SAM-ViT-H
J &F 42.4 45.3 46.8
J 40.3 43.0 43.7
F 44.5 45.6 49.9
TABLE VII D IFFERENT NUMBER OF OBJECT QUERIES IN M ASK 2F ORMER [44]. Number of Object Queries 100 50 20 10
J &F 43.9 45.7 46.8 46.4
J 41.0 42.6 43.7 43.2
F 46.8 48.8 49.9 49.6
alternative MLLMs may lead to subtle performance variations, we argue that the parameter size is not an absolute metric for evaluating an MLLM’s capabilities. Using internVL2-1B instead of LLaVA-Phi-3-3.8B improve EVIS by 0.5% and 0.2% J &F on MeViS and Ref-Youtube-VOS, respectively. However, under identical InternVL2-1B settings, EVIS outperforms VideoLISA by 5.9% J &F on MeViS. VideoLISA employs an identical configuration to the baseline model in Table I Line 1, utilizing n-frame sampling and InternVL2-1B. This alignment precisely highlights the substantial methodological innovations introduced by EVIS. Effect of Top-k Event Query Selection. In the Frame Merging Block of EAFM, as the number of object queries assigned to an event query increases, the complexity of the information encapsulated within the event query also grows. Conversely, when an event query is associated with fewer object queries, the informational overlap between distinct events becomes minimal, hindering information integration and propagation across consecutive events. As shown in Fig. 7, we conduct ablations on the effect of top-k event query selection. The performance deteriorates when the value of k is either excessively small or overly large. The best result is achieved when the number of k is set to 2. Impact of Stacking Depth L in EAFM. Tab. V shows the performance of EVIS with varying stacking depth L on MeViS [1] dataset. The results indicate that increasing the
IEEE TRANSACTIONS ON IMAGE PROCESSING
8
51
SemSeg ✓ ✓ ✓ ✓
Training Data RIOS RVOS ReasonSeg ✗ ✗ ✗ ✓ ✗ ✗ ✓ ✓ ✗ ✓ ✓ ✓
J 36.5 38.2 44.7 46.8
MeViS F J &F 35.3 37.7 36.3 40.1 42.0 47.4 43.7 49.9
MeViS Test Metrics
TABLE VIII A BLATION STUDY ON THE DIFFERENT TRAINING DATASETS .
49
47
45
43
number of stacked layers leads to improved performance. To 1 2 3 4 5 6 balance segmentation capability and computational efficiency, we select L = 3 as the optimal configuration. Different Backbones in SAM. We modify SAM’s ViT Fig. 7. Ablation on the effect of top-k event query selection on MeViS [1] architecture from H to L and B to evaluate the effectiveness of dataset. Metrics are region similarity J , contour accuracy F and their combined EVIS, demonstrating its stable progression across different average score J &F , respectively. backbone configurations. As shown in Tab. VI, when the backbone of SAM changes from ViT-H to ViT-L and ViTall evaluation metrics. On Ref-YouTube-VOS, EVIS with the B, the performance decreases by 1.5% J &F and 4.4% J &F, InternVL-1B achieves a J &F score of 64.4%, surpassing respectively, indicating the effectiveness of our method in fully VideoLISA-3.8B [10] by 0.7%. On the Ref-DAVIS17 dataset, leveraging the feature information provided by SAM, even EVIS achieves a 68.8% in J &F, surpassing the state-of-thewith varying backbone complexities. art traditional method [2] by 4.8%. It is crucial to emphasize Ablation studies for Mask2Former. The number of object that VISA-7B [9] achieves only a marginal 0.6% J &F queries has a direct impact on the model’s segmentation improvement over the proposed EVIS-1B. However, VISA performance, as shown in Tab. VII, the best result is achieved employs the Chat-UniVi-7B [40] as its MLLM, which has when number of object query is set to 20. We also conduct a significantly larger number of parameters compared to our ablation studies on whether to freeze Mask2Former. The method. These results underscore the effectiveness of EVIS in experimental results demonstrate that when Mask2Former is complex scenarios. trained in an unfrozen manner, the performance on MeViS decreases significantly from 46.8% to 42.3%. We hypothesize A2D-Sentences [17] and JHMDB-Sentences [50]. We folthis degradation stems from the disruption of Mask2Former’s low [2] to fine-tune the model on A2D-Sentences. In Tab. X, EVIS achieves new state-of-the-art results, outperforming the pre-trained query context representations. Training Datasets. We conduct ablation experiments on the LoSh [14], by 0.6% and 0.5% mAP on the A2D-Sentences training data to assess the influence of various datasets on and JHMDB-Sentences datasets, respectively. The relatively the proposed approach EVIS, as presented in Tab. VIII. The smaller performance improvements on these datasets, compared addition of RIOS, RVOS (shown in Tab. VIII), and ReasonSeg to MeViS, can be attributed to the presence of simpler, imageto the training set leads to respective improvements of 1.7%, level descriptions in the sentences, which lack the complexity of event variations. 6.5%, and 2.1% J &F, respectively. RefCOCO, RefCOCO+ [68] and RefCOCOg [69]. We evaluate the proposed EVIS on three benchmarks for the D. Comparison with State-of-the-Art Methods referring image segmentation task. As shown in Tab. XIII, With event-level analysis, EVIS achieves remarkable perfor- while our original focus is on addressing the referring video segmentation task, our approach demonstrates competitive mance across both traditional and LLM-based methods. MeViS [1]. In Tab. IX, we evaluate the proposed EVIS on performance on image tasks, highlighting the effectiveness the newly released MeViS dataset for the referring video of our method. segmentation task. Following the experimental setup in [10], ReVOS [9]. Reasoning Video Object Segmentation (Reasonour approach outperforms existing state-of-the-art methods, VOS) is a new and challenging task that differs from RVOS in achieving a significant improvement of 2.4% in J &F over the that it requires the ability to reason using world knowledge. In previous LLM-based best-performing method, VideoLISA [10]. RVOS, model might be tasked with identifying “a running cat”, Notably, our model is based on an MLLM with only 1B whereas in ReasonVOS, it needs to identify “the smartest cat”. parameters, significantly fewer than the VideoLISA-3.8B and We evaluate EVIS’s reasoning ability on the ReVOS [9] dataset, VISA-13B models, further validating the effectiveness of as shown in Tab. XI. Notably, EVIS is not trained on any Video event-by-event learning. Furthermore, EVIS† also achieves Reasoning datasets, yet it still outperforms VideoLISA [10] by competitive performance on MeViS, demonstrating the EAFM 0.8% J &F. module is compatible with both LLM-based and traditional ReasonSeg [7]. Reasoning capabilities are a hallmark of methods without reasoning capability. LLMs, enabling them to capture complex relationships and Ref-YouTube-VOS [3] and Ref-DAVIS17 [4]. In Tab. IX, contextual dependencies. To validate the image reasoning we present results on Ref-YouTube-VOS and Ref-DAVIS17 effectiveness of our approach, we conduct a comprehensive datasets. Our method outperforms most approaches across comparison with state-of-the-art methods on the ReasonSeg [7]
IEEE TRANSACTIONS ON IMAGE PROCESSING
9
TABLE IX Q UANTITATIVE EVALUATION RESULTS ON M E V I S [1], R EF -YOUTUBE -VOS [3], AND R EF -DAVIS17 [4]. † INDICATES OUR MODEL RE - IMPLEMENTED BY REMOVING THE LLM. B OLD INDICATES THE BEST SCORES . Method Traditional methods URVOS [3] LBDT [55] MLSA [56] MTTR [13] ReferFormer [12] HTML [57] R2 -VOS [58] SgMg [59] OnlineRefer [60] TempCD [61] SOC [62] VLT+TC [26] LMPM [1] LoSh [14] DsHmp [2] EVIS† (ours) LLM-based methods LISA-7B [7] TrackGPT-7B [63] VISA-7B [9] VideoLISA-3.8B [10] EVIS-1B (ours)
Backbone
J &F
MeViS J
F
Ref-Youtube-VOS J &F J F
Ref-DAVIS17 J &F J F
ECCV’20 CVPR’22 CVPR’22 CVPR’22 CVPR’22 ICCV’23 ICCV’23 ICCV’23 ICCV’23 ICCV’23 NeurIPS’23 TPAMI’23 ICCV’23 CVPR’24 CVPR’24 -
ResNet-50 ResNet-50 ResNet-50 Video-Swin-T Video-Swin-B Video-Swin-B Video-Swin-T Video-Swin-T Swin-B Video-Swin-T Video-Swin-T Video-Swin-B Swin-T Swin-T Swin-T/Video-Swin-T Swin-T
27.8 29.3 30.0 31.0 35.5 37.2 46.4 46.7
25.7 27.8 28.8 29.8 33.6 34.2 43.0 43.5
29.9 30.8 31.2 32.2 37.3 40.2 49.8 49.9
47.2 49.4 49.7 55.3 62.9 63.4 61.3 62.0 62.9 62.3 62.4 63.8 63.7 63.6 64.1
45.2 48.2 48.4 54.0 61.3 61.5 59.6 60.4 61.0 60.5 61.1 61.9 62.0 61.8 62.2
49.2 50.6 50.9 56.6 64.6 65.2 63.1 63.5 64.7 64.0 63.7 65.6 65.4 65.4 66.0
51.6 54.3 57.9 61.1 62.1 61.9 62.4 62.2 63.5 61.6 62.9 64.0 65.6
47.3 53.8 58.1 59.2 59.0 59.1 59.3 60.2 58.9 60.1 60.8 62.0
55.9 62.0 64.1 65.1 64.8 65.6 65.0 66.7 64.3 65.7 67.2 69.2
CVPR’23 arXiv’23 ECCV’24 NeurIPS’24 -
LLaVA-7B LLaVA-7B Chat-UniVi-7B LLaVA-Phi-3-3.8B InternVL2-1B
37.2 40.1 43.5 44.4 46.8
35.1 37.6 40.7 41.3 43.7
39.4 42.6 46.3 47.6 49.9
50.2 56.4 61.5 63.7 64.4
49.7 55.3 59.8 61.7 62.6
50.6 57.4 63.2 65.7 66.2
58.4 63.2 69.4 68.8 68.8
54.9 59.4 66.3 64.9 65.8
61.9 67.0 72.5 72.7 71.8
Reference
TABLE X R ESULTS ON A2D-S ENTENCES AND JHMDB-S ENTENCES .
Method ReferFormer [12] OnlineRefer [60] HTML [57] SOC [62] SgMg [59] DsHmp [2] LoSh [14] EVIS (ours)
A2D-Sentences [17] mAP oIoU mIoU 55.0 78.6 70.3 79.6 70.5 56.7 79.5 71.2 57.3 80.7 72.5 58.5 79.9 72.0 59.8 81.1 72.9 59.9 81.2 73.1 60.5 81.3 73.5
JHMDB-Sentences [50] mAP oIoU mIoU 43.7 73.0 71.8 73.5 71.9 44.2 44.6 73.6 72.3 45.0 73.7 72.5 45.8 73.9 73.0 45.7 74.5 73.4 46.2 75.5 74.2
TABLE XII R EASONING SEGMENTATION RESULTS ON R EASON S EG [7] BENCHMARK . val test overall short query long query overall Method gIoU cIoU gIoU cIoU gIoU cIoU gIoU cIoU OVSeg [64] 28.5 18.6 18.0 15.5 28.7 22.5 26.1 20.8 RELA [32] 22.4 19.9 17.6 15.0 22.6 23.8 21.3 22.0 X-Decoder [65] 22.6 17.9 20.4 11.6 22.2 17.5 21.7 16.3 25.5 21.2 20.1 11.5 25.6 20.8 24.3 18.7 SEEM [66] Grounded-SAM [67] 26.0 14.5 17.8 10.8 22.4 18.6 21.3 16.4 LISA-7B [7] 44.4 46.0 37.6 34.4 36.6 34.7 36.8 34.1 VISA-7B [9] 52.7 57.8 VideoLISA-3.8B [10] 61.4 67.1 43.8 42.7 56.9 57.7 53.8 54.4 EVIS-1B (ours) 55.4 61.3 42.5 43.0 44.3 43.7 43.8 44.6
TABLE XI R EASONING SEGMENTATION RESULTS ON R E VOS [9]. Method MTTR [13] LMPM [1] ReferFormer [12] LISA [7] TrackGPT [63] VISA [9] VideoLISA [10] EVIS (ours)
Backbone Video-Swin-T Swin-T Video-Swin-B LLaVA-7B LLaVA-7B Chat-UniVi-7B LLaVA-Phi-3-3.8B InternVL2-1B
J &F 21.0 18.8 23.4 36.1 39.0 39.2 39.5 40.3
J 20.4 13.3 21.3 33.8 36.8 36.7 37.3 38.0
F 21.5 24.3 25.6 38.4 41.2 41.7 41.7 42.6
benchmark. As shown in Tab. XII, our method achieves competitive performance across multiple metrics, demonstrating its superior ability to handle complex segmentation scenarios. These results demonstrate the robustness and generalizability of EVIS, confirming its advantages over traditional segmentation approaches. Qualitative results in Fig. 9. Training and inference costs comparisons. We conduct experiments using a single NVIDIA A6000 GPU on MeViS dataset. As shown in Tab. XIV, Compared to VideoLISA, EVIS has 58% fewer parameters and runs at 3.6× the inference speed. Furthermore, EVIS significantly outperforms DsHmp
TABLE XIII Q UANTITATIVE EVALUATION RESULTS ( C I O U) ON R EF COCO, R EF COCO+ [68] AND R EF COCO G [69]. B OLD INDICATES THE BEST SCORES . RefCOCO [68] RefCOCO+ [68] RefCOCOg [69] Method val testA testB val testA testB val(U) test(U) MCN [70] 62.4 64.2 59.7 50.6 55.0 44.7 49.2 49.4 VLT [24] 67.5 70.5 65.2 56.3 61.0 50.1 55.0 57.7 CRIS [71] 70.5 73.2 66.1 62.3 68.1 53.7 59.9 60.4 LAVT [27] 72.7 75.8 68.8 62.1 68.4 55.1 61.2 62.1 ReLA [32] 73.8 76.5 70.2 66.0 71.0 57.7 65.0 66.0 X-Decoder [65] 64.6 SEEM [66] 65.7 VISA-7B [9] 72.4 75.5 68.1 59.8 64.8 53.1 65.5 66.4 LISA-7B [7] 74.1 76.5 71.1 62.4 67.4 56.5 66.4 68.5 VideoLISA-3.8B [10] 73.8 76.6 68.8 63.4 68.8 56.2 68.3 68.8 EVIS-1B (ours) 74.2 76.8 68.9 63.9 70.1 56.2 68.5 69.7
in terms of trainable parameters, computational complexity, and inference speed. The trainable parameters of EAFM and OPH are approximately 14.45M, making them extremely lightweight compared to most existing models. We also conduct experiments on EAFM and OPH, as shown in Tab. XV.
IEEE TRANSACTIONS ON IMAGE PROCESSING
10
(a) "Smaller one of the two planes moving left and landing."
(b) "Zebra moving forward then jumping on another zebra."
"Smaller one of the two planes moving left and landing."
"Zebra moving forward then jumping on another zebra."
"Smaller one of the two planes moving left and landing."
"Zebra moving forward then jumping on another zebra."
(c) "sheep with black head standing still then turn right."
(d) "tiger moving forward then turning around to hit another tiger."
"sheep with black head standing still then turn right."
"tiger moving forward then turning around to hit another tiger."
"sheep with black head standing still then turn right."
"tiger moving forward then turning around to hit another tiger."
(e) "The fish swims from left to right, then turns around."
(f) "Horse moving to the right then stopping."
Fig. 8. Example success and failure cases of EVIS. The black font denotes text that is not visible to the model.
User: Which part of container User: What in the picture is used to User: If there is a fire in the is held to regulate the flow? prevent dog from wandering off ? bathroom, what should we use ?
User: What part of the sink can control the water flow?
Assistant: It is
Assistant: It is
.
Assistant: It is
.
Assistant: It is
.
.
Fig. 9. Qualitative results of EVIS on ReasonSeg [7] Dataset.
Qualitative Results. Fig. 8 displays some cases of the proposed EVIS. Example (a), (b), (c) and (d) show successful cases where EVIS accurately interprets partial phrase and full expressions, demonstrating its event-level video comprehension. Examples (e) and (f) are failure cases. In case (e), the target fish, temporarily occludes the head of another fish, causing the model to perceive them as a single object. Similarly, in case (f), the target horse is obscured by another horse, resulting in temporal occlusion that challenges the model’s
reasoning and self-correction capabilities. To further illustrate the model’s capability, we also compare EVIS with SOTA methods, including VideoLISA and DsHmp, in Fig. 11. We observe that prior methods struggle dramatically with queries involving multi-stage temporal dynamics. In contrast, our method correctly identifies and tracks the target throughout the partial or entire event sequence. Event Decomposition Visualization. We visualize cosine similarity between global object and event queries in EAFM,
IEEE TRANSACTIONS ON IMAGE PROCESSING
11
"Panda moves up then descends, and finally disappears. "
"Panda moves up then descends, and finally disappears. "
"Panda moves up then descends, and finally disappears. "
"Panda moves up then descends, and finally disappears. "
Ours
DsHmp
VideoLISA
Fig. 10. Visualization for event decomposition. Scores are calculated by cosine similarity between global object and event queries in EAFM.
VideoLISA
“The bull that falls down then rises up.”
Fig. 12. Visulization of global queries learned w/o training (left) and w/ training (right). Features are colored according to the similarity between global object and event queries. Best viewed in color.
Ours
DsHmp
TABLE XIV T RAINING AND INFERENCE COSTS COMPARISONS .
VideoLISA
“The bull that falls down then rises up.”
Method Backbone Training Params (M) FLOPs (G) Inference (FPS) VideoLISA LLaVA-Phi-3-3.8B ≈ 380 ≈ 8.25 ≈ 9.8 EVIS LLaVA-Phi-3-3.8B ≈ 395 ≈ 8.47 ≈ 9.2 EVIS InternVL2-1B ≈ 160 ≈ 3.15 ≈ 35.4 Swin-T ≈ 28 ≈ 0.21 ≈ 58.8 DsHmp EVIS† Swin-T ≈ 20 ≈ 0.19 ≈ 65.5 EVIS † Swin-B ≈ 63 ≈ 0.70 ≈ 31.5 Swin-L ≈ 142 ≈ 1.58 ≈ 12.4 EVIS †
Ours
DsHmp
TABLE XV A BLATION STUDY ON DIFFERENT CONFIGURATIONS OF EVIS COMPONENTS ON THE M E V I S DATASET. PARAMS REPRESENT THE TOTAL NUMBER OF PARAMETERS IN THE EAFM AND OPH MODULES .
“The bull that falls down then rises up.”
Fig. 11. Qualitative comparisons with VideoLISA [10] and DsHmp [2].
showing similarity dynamics across frames with frame numbers and similarity scores, which verify that our approach captures event-level semantics, as shown in Fig. 10. t-SNE Visualization. We Visualize global queries learned w/o training and w/ training. As shown in Fig. 12, features are colored according to the similarity between global object and event queries. Initially, the global queries exhibit a lack of clear structure. However, after training, the event-level
EAFM L Nh 3 6 4 6 5 6 3 6 3 6 3 6 3 6
OPH Tu To 4 8 4 8 4 8 2 4 4 8 6 12 8 16
Params (M)
FLOPs (G)
FPS
J &F
≈ 14.45 ≈ 19.17 ≈ 23.89 ≈ 8.33 ≈ 14.45 ≈ 20.70 ≈ 27.27
≈ 0.14 ≈ 0.20 ≈ 0.23 ≈ 0.09 ≈ 0.14 ≈ 0.19 ≈ 0.27
35.4 34.5 33.9 37.1 35.4 31.6 25.8
46.8 46.5 46.1 41.5 46.8 46.7 45.6
queries form well-separated clusters. This demonstrates EVIS’s capacity to learn hierarchical event semantics, rather than simply partitioning the video into temporal segments.
IEEE TRANSACTIONS ON IMAGE PROCESSING
V. L IMITATIONS Object Occlusion. As shown in Fig. 8 (f), excessive mutual occlusion, especially when objects are fully obscured, can lead to errors in model judgment. Current methods lack the ability to reason about absent objects, highlighting the need to incorporate spatial reasoning and temporal context. Segmentation Paradigm of LLMs. Existing LLM-based approaches typically follow the LISA [7] framework, which extends the LLM’s vocabulary by introducing a [SEG] token and employs the embedding-as-mask paradigm to facilitate segmentation. However, the features encoded by the [SEG] token capture limited coverage of the language model’s overall output. To fully leverage the reasoning capabilities of LLMs, it’s essential to enhance the utilization of hidden features. Furthermore, exploring new paradigms for LLM-based segmentation remains a promising avenue for future research. VI. C ONCLUSION We propose EVIS, an Event-Aware Video Instructed Segmentation Assistant, which leverages object features from simple events to extract event-aware object trajectories, enabling a hierarchical understanding of videos. We further introduce the Event-Aware Frame Merging Module (EAFM), which utilizes text-guided event queries to merge objects from multiple frames, generating distinct simple events. To enhance the model’s learning capabilities, we design event-intra attention for finegrained learning within events, and event-inter attention for long-term learning across events, ensuring effective alignment between video content and its corresponding expression. Moreover, we propose Object-Pixel-Mixing Learning, a strategy that accelerates target tracking over long temporal spans by integrating fine-grained pixel-level features with prior object tokens, thus enhancing the performance of MLLMs. Our approach demonstrates significant advancements in accuracy and efficiency in referring video segmentation. R EFERENCES [1] H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “MeViS: A large-scale benchmark for video segmentation with motion expressions,” in Int. Conf. Comput. Vis., 2023, pp. 2694–2703. [2] S. He and H. Ding, “Decoupling static and hierarchical motion perception for referring video segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 13 332–13 341. [3] S. Seo, J. Lee, and B. Han, “URVOS: unified referring video object segmentation network with a large-scale benchmark,” in Eur. Conf. Comput. Vis., 2020, pp. 208–223. [4] A. Khoreva, A. Rohrbach, and B. Schiele, “Video object segmentation with language referring expressions,” in ACCV, 2018, pp. 123–141. [5] Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, B. Li, P. Luo, T. Lu, Y. Qiao, and J. Dai, “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” arXiv preprint arXiv:2312.14238, 2023. [6] D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” in Int. Conf. Learn. Represent., 2024, pp. 18 378–18 394. [7] X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 9579–9589. [8] Z. Xia, D. Han, Y. Han, X. Pan, S. Song, and G. Huang, “GSVA: generalized segmentation via multimodal large language models,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 3858–3869. [9] C. Yan, H. Wang, S. Yan, X. Jiang, Y. Hu, G. Kang, W. Xie, and E. Gavves, “Visa: Reasoning video object segmentation via large language models,” in Eur. Conf. Comput. Vis., 2024, pp. 98–115.
12
[10] Z. Bai, T. He, H. Mei, P. Wang, Z. Gao, J. Chen, L. Liu, Z. Zhang, and M. Z. Shou, “One token to seg them all: Language instructed reasoning segmentation in videos,” in Adv. Neural Inform. Process. Syst., 2024, pp. 6833–6859. [11] T. F. Shipley and J. M. Zacks, Understanding events: From perception to action. Oxford University Press, 2008. [12] J. Wu, Y. Jiang, P. Sun, Z. Yuan, and P. Luo, “Language as queries for referring video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 4964–4974. [13] D. Wu, X. Dong, L. Shao, and J. Shen, “Multi-level representation learning with semantic alignment for referring video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 4986–4995. [14] L. Yuan, M. Shi, Z. Yue, and Q. Chen, “Losh: Long-short text joint prediction network for referring video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 14 001–14 010. [15] K. Koffka, Principles of Gestalt psychology. routledge, 2013. [16] R. C. Atkinson, “Human memory: A proposed system and its control processes,” The psychology of learning and motivation, vol. 2, 1968. [17] K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. M. Snoek, “Actor and action video segmentation from a sentence,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 5958–5966. [18] L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in IEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 1307–1315. [19] H. Ding, C. Liu, S. He, X. Jiang, P. H. S. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” in Int. Conf. Comput. Vis., 2023, pp. 20 167–20 177. [20] H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in Eur. Conf. Comput. Vis., 2022, pp. 640–658. [21] L. Lin, X. Yu, Z. Pang, and Y.-X. Wang, “Glus: Global-local reasoning unified into a single large language model for video segmentation,” in CVPR, 2025, pp. 8658–8667. [22] S. Gong, L. Zhang, Y. Zhuge, X. Jia, P. Zhang, and H. Lu, “Reinforcing video reasoning segmentation to think before it segments,” arXiv preprint arXiv:2508.11538, 2025. [23] Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu et al., “Deepseekmath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024. [24] H. Ding, C. Liu, S. Wang, and X. Jiang, “Vision-language transformer and query generation for referring segmentation,” in Int. Conf. Comput. Vis., 2021, pp. 16 301–16 310. [25] R. Hu, M. Rohrbach, and T. Darrell, “Segmentation from natural language expressions,” in Eur. Conf. Comput. Vis., 2016, pp. 108–124. [26] H. Ding, C. Liu, S. Wang, and X. Jiang, “VLT: vision-language transformer and query generation for referring segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6, pp. 7900–7916, 2023. [27] Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr, “Lavt: Language-aware vision transformer for referring image segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 18 155–18 165. [28] C. Liu, X. Jiang, and H. Ding, “Instance-specific feature propagation for referring segmentation,” IEEE Trans. Multimedia, vol. 25, pp. 3657–3667, 2022. [29] Y.-W. Chen, Y.-H. Tsai, T. Wang, Y.-Y. Lin, and M.-H. Yang, “Referring expression object segmentation with caption-aware consistency,” in Brit. Mach. Vis. Conf., 2019. [30] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst., 2017. [31] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” arXiv preprint arXiv:2304.02643, 2023. [32] C. Liu, H. Ding, and X. Jiang, “GRES: Generalized referring expression segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 23 592–23 601. [33] S. Gong, Y. Zhuge, L. Zhang, Z. Yang, P. Zhang, and H. Lu, “The devil is in temporal token: High quality video reasoning segmentation,” in CVPR, 2025, pp. 29 183–29 192. [34] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Adv. Neural Inform. Process. Syst., 2023, pp. 34 892–34 916. [35] J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML, 2023, pp. 19 730–19 742.
IEEE TRANSACTIONS ON IMAGE PROCESSING
[36] Y. Cao, P. Zhang, X. Dong, D. Lin, and J. Wang, “Dualfocus: Integrating macro and micro perspectives in multi-modal large language models,” arXiv preprint arXiv:2402.14767, 2024. [37] J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-vl: A frontier large vision-language model with versatile abilities,” arXiv preprint arXiv:2308.12966, 2023. [38] H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 286–26 296. [39] W. Hu, Y. Xu, Y. Li, W. Li, Z. Chen, and Z. Tu, “BLIVA: A simple multimodal LLM for better handling of text-rich visual questions,” in AAAI, 2024, pp. 2256–2264. [40] P. Jin, R. Takanobu, W. Zhang, X. Cao, and L. Yuan, “Chat-univi: Unified visual representation empowers large language models with image and video understanding,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 13 700–13 710. [41] X. Ma, Y. Zhou, H. Wang, C. Qin, B. Sun, C. Liu, and Y. Fu, “Image as set of points,” in Int. Conf. Learn. Represent., 2023. [42] Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” in Adv. Neural Inform. Process. Syst., 2021, pp. 13 937–13 949. [43] S. Ren, S. Chen, S. Li, X. Sun, and L. Hou, “TESTA: temporalspatial token aggregation for long-form video-language understanding,” in EMNLP, 2023, pp. 932–947. [44] B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Maskedattention mask transformer for universal image segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 1280–1289. [45] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019. [46] J. Xu, S. D. Mello, S. Liu, W. Byeon, T. M. Breuel, J. Kautz, and X. Wang, “Groupvit: Semantic segmentation emerges from text supervision,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 18 113–18 123. [47] E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in Int. Conf. Learn. Represent., 2017. [48] A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Adv. Neural Inform. Process. Syst., 2017. [49] F. Milletari, N. Navab, and S. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in 3DV, 2016. [50] H. Jhuang, J. Gall, S. Zuffi, C. Schmid, and M. J. Black, “Towards understanding action recognition,” in Int. Conf. Comput. Vis., 2013, pp. 3192–3199. [51] J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbeláez, A. Sorkine-Hornung, and L. V. Gool, “The 2017 DAVIS challenge on video object segmentation,” arXiv preprint arXiv:1704.00675, 2017. [52] Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, and et al, “How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,” arXiv preprint arXiv:2404.16821, 2024. [53] A. Yang, B. Yang, B. Hui, B. Zheng, and et al, “Qwen2 technical report,” arXiv preprint arXiv:2407.10671, 2024. [54] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Int. Conf. Learn. Represent., 2019. [55] Z. Ding, T. Hui, J. Huang, X. Wei, J. Han, and S. Liu, “Language-bridged spatial-temporal interaction for referring video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 4954–4963. [56] D. Wu, X. Dong, L. Shao, and J. Shen, “Multi-level representation learning with semantic alignment for referring video object segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 4986–4995. [57] M. Han, Y. Wang, Z. Li, L. Yao, X. Chang, and Y. Qiao, “HTML: hybrid temporal-scale multimodal learning framework for referring video object segmentation,” in Int. Conf. Comput. Vis., 2023, pp. 13 368–13 377. [58] X. Li, J. Wang, X. Xu, X. Li, B. Raj, and Y. Lu, “Robust referring video object segmentation with cyclic structural consensus,” in ICCV, 2023, pp. 22 179–22 188. [59] B. Miao, M. Bennamoun, Y. Gao, and A. Mian, “Spectrum-guided multigranularity referring video object segmentation,” in Int. Conf. Comput. Vis., 2023, pp. 920–930. [60] D. Wu, T. Wang, Y. Zhang, X. Zhang, and J. Shen, “Onlinerefer: A simple online baseline for referring video object segmentation,” in Int. Conf. Comput. Vis., 2023, pp. 2749–2758. [61] J. Tang, G. Zheng, and S. Yang, “Temporal collection and distribution for referring video object segmentation,” in Int. Conf. Comput. Vis., 2023, pp. 15 420–15 430. [62] Z. Luo, Y. Xiao, Y. Liu, S. Li, Y. Wang, Y. Tang, X. Li, and Y. Yang, “SOC: semantic-assisted object cluster for referring video
13
object segmentation,” in Adv. Neural Inform. Process. Syst., 2023, pp. 26 425–26 437. [63] J. Zhu, Z. Cheng, J. He, C. Li, B. Luo, H. Lu, Y. Geng, and X. Xie, “Tracking with human-intent reasoning,” arXiv preprint arXiv:2312.17448, 2023. [64] F. Liang, B. Wu, X. Dai, K. Li, Y. Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu, “Open-vocabulary semantic segmentation with maskadapted CLIP,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 7061–7070. [65] X. Zou, Z. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan, N. Peng, L. Wang, Y. J. Lee, and J. Gao, “Generalized decoding for pixel, image, and language,” in IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 15 116–15 127. [66] X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, and Y. J. Lee, “Segment everything everywhere all at once,” in Adv. Neural Inform. Process. Syst., 2023, pp. 19 769–19 782. [67] T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang, “Grounded SAM: assembling open-world models for diverse visual tasks,” arXiv preprint arXiv:2401.14159, 2024. [68] S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in EMNLP, 2014, pp. 787–798. [69] J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in IEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 11–20. [70] G. Luo, Y. Zhou, X. Sun, L. Cao, C. Wu, C. Deng, and R. Ji, “Multi-task collaborative network for joint referring expression comprehension and segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 10 031–10 040. [71] Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu, “CRIS: clip-driven referring image segmentation,” in IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 11 676–11 685.
Jinyu Liu received the M.S. degree from Fudan University, Shanghai, China, in 2023. He is currently a Ph.D. student at College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China. Prior to that, he was a Research Intern at Tencent. His research interests include computer vision and multi-modal learning.
Henghui Ding (Member, IEEE) received the B.E. degree from Xi’an Jiaotong University, China, in 2016, and the Ph.D. degree from Nanyang Technological University (NTU), Singapore, in 2020. He was a Research Scientist at ByteDance, a Postdoctoral Researcher at ETH Zurich and NTU. He is currently a Professor at Fudan University, Shanghai, China. He serves as an Associate Editor for IEEE Transactions on Image Processing (TIP) and Pattern Recognition (PR), and regularly serves/served as a Senior Area Chair or Area Chair of top conferences such as CVPR, NeurIPS, ICLR, ICML, ECCV, AAAI, and ACM MM. His research interests include computer vision and machine learning.
IEEE TRANSACTIONS ON IMAGE PROCESSING
Shuting He received the B.E. degree from Xiamen University (XMU), Xiamen, China, in 2018, and the Ph.D. degree from Zhejiang University (ZJU), Hangzhou, China, in 2023. She is currently a tenuretrack Assistant Professor with Shanghai University of Finance and Economics (SUFE). Prior to that, she was a Research Fellow with Nanyang Technological University (NTU), Singapore. She serves as an Area Chair for CVPR, NeurIPS, ICLR, and BMVC. Her research interests include computer vision and machine learning.
Yu-Gang Jiang (Fellow, IEEE) received the PhD degree in Computer Science from City University of Hong Kong in 2009 and worked as a Postdoctoral Research Scientist at Columbia University, New York, during 2009-2011. He is currently a Distinguished Professor of Computer Science at Fudan University, Shanghai, China. His research lies in the areas of multimedia, computer vision, embodied AI and trustworthy AI. His research has led to the development of innovative AI tools that have been used in many practical applications like defect detection for highspeed railway infrastructures. His open-source video analysis toolkits and datasets such as CU-VIREO374, CCV, THUMOS, FCVID and WildDeepfake have been widely used in both academia and industry. He currently serves as Chair of ACM Shanghai Chapter and Associate Editor of several international journals. For contributions to large-scale and trustworthy video analysis, he was elected to Fellow of IEEE, IAPR, and CCF.
14