Meta-CoT: Enhancing Granularity and Generalization in Image Editing Shiyi Zhang1,2,* , Yiji Cheng2,* , Tiankai Hang2,* , Zijin Yin2 , Runze He2 , Yu Xu2 , Wenxun Dai1,2 , Yunlong Lin2 , Chunyu Wang2 , Qinglin Lu2 , Yansong Tang1,† 1 Shenzhen International Graduate School, Tsinghua University 2 Hunyuan, Tencent {sy-zhang23@mails., tang.yansong@sz.}tsinghua.edu.cn
Meta-CoT
Granularity
arXiv:2604.24625v1 [cs.CV] 27 Apr 2026
Triplet Decomposition
“What does the interior of this car look like?”
Thinking...
CEC Reward
Generalization
Meta-task Decomposition
Ⅰ
Task
1
Camera Motion
ⅠⅠ
Target
2
Object Addition
ⅠⅠⅠ
Understanding
3
Object Deletion
Instruction
Meta-task 1: Camera Motion
Meta-CoT
Step 1: Task Summary The editing instruction is: What does the interior of …This is a Spatial Reasoning task.
Move the camera view from outside the car to inside the car.
Meta-task 2: Object Addition
Step 2: Task Thinking New subjects appearing: Car seats, gear shifts, steering wheels, and other in-car accessories… Old subjects disappearing or changes: …
MLLM
Add Car seats to the image; Add gear shifts to the image; Add steering wheels to the image; Add a sky window to the image …
Step 3: Target Traversal
Measure the consistency between CoT and Editing Results
Meta-task 3: Object Deletion
Car: Remove, move the camera viewpoint into the car; Trees: Remove; Bushes: Remove…
Remove the car in the image; Remove the trees in the image; Remove the bushes …
CEC Reward = 9.0
(a) Meta-CoT (left) and CoT-Editing Consistency Reward (right) 6.415
Meta-CoT + GRPO (CEC Reward)
Full-task SFT
Meta-CoT (full- task) Meta-CoT (5 meta-task)
5
5.307 5.4
6.011
15.8%
5.538
Train Edit Only Bagel Think
Meta-task SFT
6.174 m Co
r pa
a
e bl
6.2
3.71
Meta-CoT
Bagel Think
3.39
3.2
Bagel w/o Think
3.1
11.7%
3.43
Train Edit Only
20.1% 5.8
3.83
Meta-CoT + GRPO (CEC Reward)
3.3
13.0% 3.5
3.7
(b) Overall Score on 21-tasks Benchmark (c) Overall Score on ImgEdit Figure 1. (a) We propose Meta-CoT with a hierarchical decomposition design and the CoT–Editing Consistency Reward. On (b) 21-task Editing Benchmark and (c) ImgEdit, our method shows: (1) Stronger instruction-based editing. It achieves 15.8% and 11.7% gains over the Train-Edit-Only baseline trained on the same parameters and editing data (without Meta-CoT). (2) Generalization to unseen tasks. As shown in (b), meta-task training enables comparable performance to full-task training while training only on a small subset of meta-tasks.
Abstract Unified multi-modal understanding/generative models have shown improved image editing performance by incorporating fine-grained understanding into their Chain-of-Thought (CoT) process. However, a critical question remains underexplored: what forms of CoT and training strategy can jointly enhance both the understanding granularity and generalization? To address this, we propose Meta-CoT, a paradigm that performs a two-level decomposition of any single-image editing operation with two key properties: (1) Decomposability. We observe that any editing intention can be represented as a triplet — (task, target, required understanding ability). Inspired by this, Meta-CoT decomposes both the editing task and the target, generating task-specific CoT and traversing editing operations on all targets. This * Equal contribution. † Corresponding author.
decomposition enhances the model’s understanding granularity of editing operations and guides it to learn each element of the triplet during training, substantially improving the editing capability. (2) Generalizability. In the second decomposition level, we further break down editing tasks into five fundamental meta-tasks. We find that training on these five meta-tasks, together with the other two elements of the triplet, is sufficient to achieve strong generalization across diverse, unseen editing tasks. To further align the model’s editing behavior with its CoT reasoning, we introduce the CoT-Editing Consistency Reward, which encourages more accurate and effective utilization of CoT information during editing. Experiments demonstrate that our method achieves an overall 15.8% improvement across 21 editing tasks, and generalizes effectively to unseen editing tasks when trained on only a small set of meta-tasks. Our code, benchmark, and model are released at here.
Step 1: Task Summary <Determine task type> The editing instruction is: Zoom out the camera view. This is a Camera Motion task.
- Two stools: Add two stools under the countertop. They should have a round leather seat with wooden legs and metal springs. Position them symmetrically under the countertop.
Task
Step 2: Task Thinking <Generate specific thinking processes for different tasks >
New subjects appearing:
Old subjects disappearing or changes in visibility:
- No old subjects disappear completely, but all existing objects become smaller and occupy less space in the frame due to the zoom-out effect. Adjust their positions and sizes as specified above.
“Zoom out the camera view.”
Captioning Thinking...
Reasoning
···
Understand
Mix Train with Understanding Data
Target
Step 3: Target Editing Mode Traversal <Traverse all targets and determine whether to edit and how to edit> - The backpack: Reduce its size proportionally and move it slightly upward and backward to fit the new perspective. - The black rectangular object: Reduce its size proportionally and move it slightly upward and backward. - The glass vase: Reduce its size proportionally and move it slightly upward and backward. - The small objects near the glass vase: Reduce their size proportionally and move them slightly upward and backward. - The countertop: Extend its visible area both horizontally and vertically. - The cabinets above the countertop: Extend their visible area both horizontally and vertically. - The floor: Add the visible area below the countertop, including the stools. - The wall behind the countertop: Extend its visible area both horizontally and vertically.
Figure 2. Overview of Triplet Decomposition. Triplet Decomposition enables fine-grained reasoning over both the task and the target through three steps: (1) Task Summary, (2) Task Thinking, and (3) Target-wise Editing Traversal. This decomposition allows the model to optimize reasoning and editing along two elements: task and target. During training, we jointly incorporate diverse visual understanding tasks to capture all components of the (task, target, understanding capability) triplet, as illustrated in the bottom-right.
1. Introduction Recent studies show that Chain-of-Thought (CoT) can effectively enhance editing performance in unified generation/understanding models [8, 11]. By explicitly reasoning the editing process, the model activates its understanding ability and produces more accurate edits. We argue that an effective image editing CoT paradigm should possess two properties: (1) Stimulation of the model’s understanding ability. (2) Generalization across diverse editing tasks. Existing works mainly focus on the first property, such as introducing spatial localization cues or other explicit understanding information into the CoT [11]. However, such approaches often exhibit limited generality since CoTs tailored to specific understanding forms may not adapt well to broader editing tasks. For example, localization-based CoT performs poorly on tasks like style transfer or viewpoint transformation. Therefore, this paper aims to investigate a critical question: what forms of CoT and training strategy can simultaneously enhance understanding granularity and generalization capability during image editing? To address this, we first observe that any single-image editing operation can be defined by a triplet: (Task, Target, Required Understanding Ability). For example, the editing instruction “Change the number of puppies to three” involves the task type “Quantity Modification”, editing targets
“puppies”, and requires the model to possess understanding abilities in localization and counting. Building on this, as shown in Figure 2, we propose the first decomposition level in the Meta-CoT, Triplet Decomposition, which guides finegrained reasoning over both the editing task and the editing target. This paradigm comprises three steps: (1) Task Summary, (2) Task Thinking, and (3) Target-wise Editing Traversal. This design offers strong decomposability, enabling the model to optimize over task and target, thereby not only learning to comprehend diverse editing operations but also mastering how to apply distinct editing strategies to different entities. To capture the third element, required understanding ability, of the triplet, we incorporate data from a variety of visual understanding tasks during training, ensuring the model fully learns across all three elements. Furthermore, to achieve generalization, we analyze single-image editing tasks and identify a set of primitive, universal operations, termed “meta-tasks”. As shown in Table 1 and Figure 3, analogous to bases in a vector space, these meta-tasks (e.g., add, remove, replace) form a minimal set capable of spanning the entire editing operation space. We then propose the second decomposition level in Meta-CoT: Meta-task Decomposition. Specifically, we redefine the triplet element “Task” as one or more metatasks, and adapt the CoT paradigm by replacing “Task Summary” with “Meta-task Summary”. For example, “Quantity
Table 1. Any single-image editing operation can be represented as a triplet of Task (first column), Editing Target (third column), and Required Understanding Capability (fourth column). Each operation can be further decomposed into combinations of Meta-Tasks. We summarize five generic meta-tasks (top) and list their possible compositional combinations for different editing tasks (second column). Task Addition Deletion Replacement Camera Motion Position Change Style Transfer Tone Adjustment Text Editing Shape Modification Structural Change Color Change Quantity Change Specified Quantity Change Human Attribute Editing Material Change Motion Change Logical Reasoning Causal Reasoning Spatial Composition Temporal Reasoning Multi-Instruction Editing Input Image
Meta-Task
Target
Understanding
Addition Deletion Replacement Camera Motion Position Change Replacement Replacement Addition, Deletion, Replacement Replacement Addition, Deletion, Replacement, Camera Movement Replacement Addition, Deletion Any Addition, Deletion, Replacement Replacement Replacement Addition, Deletion, Replacement Addition, Deletion, Replacement Addition, Deletion, Position Change, Camera Movement Addition, Deletion, Replacement Any
Any Any Any Any Any Style Tone / Lighting Text Shape Artificial Structure Color Any Any Human Material Motion Shape, Symbol, Text Any Any Any Any
Localization Localization Localization Spatial Understanding Spatial Understanding Style Understanding Tone Understanding Image-text Understanding / OCR Shape Understanding Structural Understanding Color Understanding Counting Counting, Localization Human Attribute Understanding Material Understanding Motion Understanding Symbolic / Mathematical Reasoning Causal Reasoning Spatial Reasoning Temporal Reasoning Multi-instruction Understanding
Instruction
Meta-Task Decomposition
Please complete the answers to the questions
1+1=2. Thus, this task can be summarized as: Replace the question mark in the image with “2”.
What will he look like when he gets old?
This task can be summarized as: Replace black hair with white hair; add wrinkles and age spots to the face; add wrinkles to the hands.
What does the interior of this car look like?
This task can be summarized as: Move the viewpoint inside; add components like the steering wheel, seats, and dashboard display.
Edited Image
Figure 3. Examples of editing intent decomposed into meta-tasks.
Change” corresponds to the meta-tasks add and remove. The model learns to decompose instructions into simple meta-tasks, forming a reasoning chain of fundamental operations. This paradigm yields strong generalizability: by training on only a small number of meta-tasks (as few as five), the model can generalize to all other editing tasks via compositional reasoning. Thus, during training, we no longer need to cover every task individually; all remaining tasks can be expressed as combinations of these meta-tasks. We also incorporate diverse visual understanding tasks into the training data to ensure mastery of all triplet elements. We further observe the inconsistency between the CoT reasoning and the final edited result in existing models. To address this, we propose the CoT–Editing Consistency Reward. Specifically, we employ a VLM (e.g., Qwen2.5VL [2]) to assess the consistency between the CoT and the edited image from both the task and object perspectives, and provide corresponding rewards. Using this reward, we perform Flow-GRPO [37] and effectively improve the alignment between the CoT reasoning and editing outcomes. To validate our approach, we construct a benchmark cov-
ering 21 distinct editing tasks, including Gedit-Bench [38], RiseBench [84], ComplexEdit [73], and five additional benchmarks we developed to fill missing task types. Compared to existing benchmarks, our benchmark offers broader task coverage, enabling a more thorough evaluation. The benchmark data is distributionally independent from training data, ensuring fair evaluation. We also conduct experiments on ImgEdit [75] to demonstrate the effectiveness of our method. We conduct our experiments based on Bagel [8], a mainstream open-source unified model that supports CoT–based editing, making it well-suited for our comparative study. After applying Meta-CoT via SFT and the CoT–Editing Consistency Reward via GRPO optimization, our method achieves a 15.7% overall performance improvement across all 21 tasks compared to the base model. Moreover, experiments demonstrate that training on only a few meta-tasks is sufficient to achieve strong generalization to other editing tasks, yielding performance comparable to that achieved by training on the full set of task types. The contributions of this paper can be summarized as: • We propose Triplet Decomposition, which breaks down editing tasks into task, target, and understanding ability, enhancing the granularity during the reasoning process. • We introduce Meta-task Decomposition, which reduces tasks to fundamental meta-tasks, enabling strong generalization through the meta-task-based training strategy. • We design a CoT-Editing Consistency Reward to mitigate mismatches between CoT reasoning and editing results. • We construct a comprehensive image editing benchmark which covers a broader range of editing task types and empirically validate the effectiveness of our approach.
2. Related Work 2.1. CoT in Understanding/Generation Models Chain-of-Thought (CoT) reasoning, originally introduced to elicit step-by-step logical inference in large language models, has recently been extended to vision-language models (VLMs) to enhance multimodal reasoning [20, 40, 57, 59, 81]. Early efforts integrated visual information into textual CoT [40, 81], while subsequent work developed native multimodal CoT by interleaving textual reasoning with visual tokens [12, 14, 18, 29, 50, 63]. CoT has also been applied to image generation [10, 15, 21, 23] and editing [8, 11]. In image editing, several works have incorporated CoT to improve instruction following and spatial reasoning. For instance, Bagel [8] and GoT [11] generate intermediate reasoning steps before editing images, demonstrating improved consistency between visual understanding and generation. However, current CoT paradigms for editing are either too generic to stimulate the model’s understanding capacity [8] or too specialized to adapt across diverse tasks [11], motivating us to develop a paradigm that enhances both the understanding granularity and generalization in image editing.
2.2. Unified Understanding and Generation Models Recent research has increasingly focused on unifying multimodal understanding [1, 2, 78, 79] and generation [30, 31, 41, 43, 47, 68, 80, 86] within single models for joint understanding and synthesis [4, 7, 8, 16, 17, 22, 33, 35, 36, 53, 54, 60, 62, 66, 67, 69, 71, 72, 87]. GPT-4o [45] exemplifies this by fusing visual analysis and generation, outperforming earlier unified models. Current unified models can be organized into three categories: (1) Autoregressive approaches that apply next-token prediction to jointly generate textual and visual tokens [6, 39, 48, 54, 58, 60]; (2) Additional Diffusion frameworks that couple a pre-trained LLM backbone with an external diffusion module, where the language model derives semantic conditions to guide the diffusion process [9, 46, 56, 64]; and (3) Unified Integrated Transformer models that natively combine both LLM and diffusion mechanisms within a shared transformer architecture [8, 32, 42, 52, 85]. We build our method upon Bagel [8], a unified multimodal transformer with intrinsic CoT-based editing capabilities, making it a suitable foundation for CoT paradigms exploration in image editing.
3. Method 3.1. Theoretical Definition Triplet Decomposition. Let T denote the original CoT space, and let T1 , T2 , and T3 denote the task, target, and required understanding capability. Then the triplet space is Striplet = T1 × T2 × T3 . We further let H =
log |space| denote space complexity (i.e., entropy). We prove: H(T1 , T2 , T3 ) = log |Striplet | < log |T | = H(T ), indicating triplet decomposition lowers the editing complexity. Next, we use the mutual information per entropy I(T ;X ) G = H(T )tgt [55] to denote the understanding granularity in CoT, where Xtgt is the target image. We prove: G(T1 , T2 , T3 ) > G(T ) , indicating Meta-CoT achieves higher understanding granularity than classical CoT (noted as T ). See supplementary material for detailed derivation. Meta-task Decomposition. B = {t1 , . . . , tn } is a set of meta-tasks. We say B forms a basis of the task space T if ∀T ∈ T , ∃ ti1 , . . . , tik ∈ B s.t. T = ti1 ◦ ti2 ◦ · · · ◦ tik .
3.2. Triplet Decomposition As shown in Table 1, any single-image editing operation can be decomposed into a triplet comprising three fundamental elements: the editing task, the editing target, and the understanding capability required to accomplish the operation. Building on this insight, we propose the first decomposition level, Triplet Decomposition (illustrated in Figure 2), in Meta-CoT. This paradigm decomposes an editing instruction into task and target, guiding the model to explicitly learn to understand different editing tasks and master editing methods for various targets during training. To accommodate the diverse understanding capabilities required for different editing tasks, we incorporate a diverse set of visual understanding tasks during training, ensuring the comprehensive mastery of all three elements of the defined triplet. As shown in Figure 2, our Triplet Decomposition unfolds in three steps. (1) Task Summary. The model infers the task type inductively from the instruction. (2) Task Thinking. The model generates a task-specific reasoning process based on the task type. For instance, for style transfer, it analyzes the visual attributes of the target style; for camera motion, it identifies object appearance or disappearance; and for logical reasoning-based editing, it deduces the implicit operations suggested by the instruction (see the supplementary materials for more examples). (3) Target Editing Mode Traversal. The model traverses all targets in the image and reasons about whether and how each should be edited. This step ensures spatial and semantic consistency and provides fine-grained, interpretable editing guidance.
3.3. Meta-task Decomposition Furthermore, as shown in Table 1, we identify a set of generic operations within single-image editing, referred to as “meta-tasks”. Meta-tasks serve as a set of bases in the single-image editing operation space, capable of combining and generalizing to produce various complex editing operations. In the ideal case, training on these meta-tasks enables the model to handle other, more complex editing tasks. Building on this insight, we propose the second decomposition level, Meta-task Decomposition, in Meta-CoT. In
Trainable Stage I & II
Freeze Stage II
CEC Reward = 9.0
Trained in Stage I Frozen in Stage II
Check CoT-Editing Consistency with Gemini In-context Prompting
MLLM Meta-CoT
Target Image Instruction: Zoom out the view.
Und. Expert
Gen. Expert
Check
Parse Task
Qwen2.5-VL
In-context Prompting
Task: Camera Movement
Random Noise Text Token. Instruction
Und Encoder
Qwen2.5
Gen Encoder
In-context Prompting
Meta-CoT (Meta-)Task Summary: This is a Camera Motion task.
Task Thinking: New subjects appearing: - Two stools: Add two stools under the countertop … Old subjects disappearing or changes in visibility: - No old subjects disappear completely, but all …
Subject Editing Mode Traversal: - The backpack: Reduce its size proportionally … - The black rectangular object: Reduce its size and … - The glass vase: Reduce its size proportionally and …
Input Image
Figure 4. Training pipeline (self attention omitted). (1) Stage 1: SFT on both reasoning and editing. (2) Stage 2: GRPO on editing.
Figure 5. Meta-CoT Data Construction. This pipeline processes the source image, target image, and instruction to the Meta-CoT.
practice, we define five distinct meta-tasks (listed in Table 1). Accordingly, our triplet evolves into (meta-task, target, required understanding capability). Next, we replace the first step of Meta-CoT, Task Summary, with Meta-task Summary, decomposing the instruction into a combination of basic meta-tasks. For example, the style transfer task can be decomposed into a “replacement” operation on the style attribute of the editing target. During training, as noted in Section 3.2, we supervise all triplet elements by training on data of the five meta-tasks and incorporating diverse visual understanding tasks to achieve comprehensive mastery.
denoising timesteps, where semantic fidelity is most critical, and omit updates on later timesteps. Empirically, we find that reducing optimization on later timesteps also alleviates potential noise artifacts introduced by Flow-GRPO.
3.4. CoT-Editing Consistency Reward In our experiments, we observed that in certain editing scenarios, particularly when the editing instruction does not explicitly specify the operation, the model may fail to follow the reasoning outlined in the CoT, even if the correct editing operation is inferred. This misalignment between CoT reasoning and execution often results in incorrect editing. To address this, we introduce the CoT–Editing Consistency (CEC) Reward. Specifically, we design a consistency metric from both the task and target perspectives, where a VLM (Qwen2.5-VL [2]) evaluates whether the generated edit aligns with the CoT reasoning in terms of both operation and target, producing a score from 0 to 10. Before training, we conduct a correlation validation for the CEC Reward. Specifically, we use the Meta-CoT SFT model to generate 500 editing samples with CoT. Four annotators then score the CoT–editing consistency. For samples with scores ranging < 3, we average all scores. For those with a range ≥ 3, we average the three closest scores. We then iteratively adjust the VLM’s initial prompt, computing the Pearson correlation r and mean absolute error ϵMAE against human annotations on the 500 samples, until r ≥ 0.8 and ϵMAE ≤ 2.5 [25, 27, 74]. We provide the details for the CEC Reward in the supplementary material. We adopt Flow-GRPO [37, 51] to optimize the model with the CEC Reward. Since the CEC Reward measures semantic alignment, we only focus the optimization on early
3.5. Training Pipeline As shown in Figure 4, our training includes two stages. In the SFT stage, both the understanding expert and generation expert are tuned to train CoT reasoning and image editing, with the image understanding encoder also updated. In the subsequent RL stage, we freeze the image understanding encoder and train only the generation expert. This is motivated by two observations: (1) after SFT, the model already achieves highly accurate CoTs that should be preserved; (2) training both modules during RL causes unstable optimization and degrades the reasoning ability learned in SFT.
3.6. Meta-CoT Data Creation Pipeline As shown in Figure 5, for each editing sample, we first determine its editing task type with Qwen2.5 [70], based on carefully designed task definition rules and the instruction, followed by a consistency check between the predicted task type and instruction with Gemini-2.5-Flash. Next, we input the source image, target image, instruction, and task type into Qwen2.5-VL [2], which, guided by a carefully designed prompt, generates the (Meta-)Task Summary, Task Thinking, and Target Editing Mode Traversal. This process also includes an evaluation to verify the alignment between the generated Meta-CoT and the actual editing process.
4. Experiment 4.1. Implementation Details Benchmark and Metrics. To comprehensively evaluate our model’s performance across diverse editing tasks, we construct a benchmark comprising 21 editing tasks. Among them, 11 categories are inherited from and fully overlap with GEdit-Bench [38], and 4 logic-related categories are fully sourced from RiseBench [84]. The multi-instruction
Table 2. Comparison of Overall Scores on the 21-task benchmark. All metrics are evaluated using GPT-4.1. Train Editing Only denotes the setting trained with the same parameters and editing data as our method, but without Meta-CoT. Background
Color
Material
Action
Human Attribute
Style
Add
Remove
Replace
Text
Tone
Bagel(w/o think) Bagel(w think) Train Editing Only SFT(Meta-CoT)
6.172 6.686 6.743 7.173
6.668 6.469 6.537 7.201
6.024 5.969 6.052 6.493
4.182 3.997 4.155 4.470
4.933 3.771 4.257 4.760
6.760 6.156 6.571 6.965
7.244 7.338 7.363 7.640
6.319 6.272 6.296 7.833
6.774 6.430 6.687 6.931
5.170 2.124 2.651 3.271
6.323 6.511 6.547 7.236
Meta-CoT + RL (Ours)
7.251
7.323
6.636
4.574
4.956
7.035
7.762
8.129
7.065
3.328
7.318
Method / Task
Causal
Logical
Bagel(w/o think) Bagel(w think) Train Editing Only SFT(Meta-CoT)
4.628 5.561 5.710 6.647
2.994 3.146 3.217 3.526
4.706 4.718 4.683 4.858
3.982 5.477 5.663 6.233
6.364 5.572 5.964 6.876
6.492 5.188 5.629 7.309
5.838 4.549 5.177 5.862
5.581 4.530 4.957 6.468
5.177 4.909 5.093 6.019
6.799 6.066 6.349 6.936
5.673 5.307 5.538 6.224
Meta-CoT + RL (Ours)
6.953
4.014
5.046
6.376
7.100
7.557
6.185
6.712
6.322
7.077
6.415
Method / Task
Spatial Temporal Reasoning
Camera
Structure Position Quantity
Specified MultiAverage Quantity Instruction
Table 3. System comparison on ImgEdit. All metrics are evaluated by GPT-4.1. Overall denotes average score across all tasks. Method
Add
Adjust
Extract
Replace
Remove
Background
Style
Hybrid
Action
Overall ↑
4.26 4.57
4.57 4.93
3.68 3.96
4.63 4.89
4.00 4.20
1.81
3.19
2.68
2.76
2.77
1.75 1.44 2.24 2.83 3.21 3.08 3.16 4.38
2.38 3.55 2.85 3.76 4.19 3.84 4.63 4.81
1.62 1.20 1.56 1.91 2.24 2.04 2.64 3.82
1.22 1.46 2.65 2.98 3.38 3.68 2.52 4.69
1.90 1.88 2.45 2.70 2.96 3.05 3.06 4.27
Closed-source Models FLUX.1 Kontext [Pro] [26] GPT Image 1 [High] [44]
4.25 4.61
4.15 4.33
2.35 2.90
IEAP [19]
2.34
3.08
2.03
MagicBrush [77] Instruct-Pix2Pix [3] AnyEdit [76] UltraEdit [83] OmniGen [65] ICEdit [82] Step1X-Edit [38] Qwen-Image [61]
2.84 2.45 3.18 3.44 3.47 3.58 3.88 4.38
1.58 1.83 2.95 2.81 3.04 3.39 3.14 4.16
1.51 1.44 1.88 2.13 1.71 1.73 1.76 3.43
4.56 4.35
3.57 3.66
Multi-round Editing Models 3.62
3.45
Diffusion Only Models 1.97 2.01 2.47 2.96 2.94 3.15 3.40 4.66
1.58 1.50 2.23 1.45 2.43 2.93 2.41 4.14
Unified Models GoT [11] Ming-UniVision [22] BAGEL(w/o think) [8] UniWorld-V1 [34] BAGEL(think) [8] OmniGen2 [62] BLIP3o-NEXT [5]
3.61 3.55 3.56 3.82 3.65 3.57 4.00
2.94 3.14 3.31 3.64 3.53 3.06 3.78
1.35 1.52 1.70 2.27 2.03 1.77 2.39
2.78 3.25 3.38 3.47 3.60 3.74 4.05
2.57 3.29 2.62 3.24 3.03 3.20 2.61
2.29 2.77 3.24 2.99 3.45 3.57 4.30
3.51 3.99 4.49 4.21 4.43 4.81 4.64
1.75 2.74 2.38 2.96 2.59 2.52 2.67
2.66 3.91 4.17 2.74 4.22 4.68 4.13
2.61 3.06 3.20 3.26 3.39 3.44 3.62
Meta-CoT+RL(Ours) ∆ Over Base Model
3.87 +6.0%
3.91 +10.8%
2.40 +18.2%
4.22 +17.2%
3.74 +23.4%
3.98 +15.4%
4.80 +8.4%
3.26 +25.9%
4.33 +2.6%
3.83 +13.0%
editing task is fully drawn from ComplexEdit [73]. We additionally introduce 5 new task categories (each with 100 samples) built from data entirely independent of the training set. Following GEdit-Bench, we adopt the Overall Score from VIEScore [25], which jointly measures instruction following, subject consistency, naturalness, and artifacts (ranging from 0 to 10) as our evaluation metric. Following [8, 34, 38, 61], all metrics are evaluated using GPT-4.1. We also evaluate our method on ImgEdit [75], which encompasses nine representative editing tasks covering diverse editing categories, with a total of 734 real-world test cases. The evaluation metrics include instruction adherence, image editing quality, and detail preservation, each
scored from 1 to 5, with all scores assessed by GPT-4.1. Training Details During the SFT stage, we train 10k steps on 48 GPUs using a 1.5M image–instruction–CoT dataset built from open-source data [24, 49]. Imageinstruction pairs are created by (1) instructions generation with Gemini-2.5-Flash under our defined edit taxonomy, (2) image editing with [26, 44, 61], and (3) filtering with both VLM (Gemini-2.5-Flash, GPT-4.1) and human evaluation. The creation of Meta-CoTs follows Section 3.6. The joint 100k understanding data source from LLaVA-OV [28] and Mammoth-VL [13]. During the RL stage, we train for 500 steps on an additional 20K editing dataset using 32 GPUs. More details are provided in the supplementary material.
Table 4. Comparison of the four components that form the Overall Score in VIEScore across the 21-task benchmark. Method
Ins.
Con.
Nat.
Art.
Bagel(w/o think) Bagel(w think) Train Editing Only SFT(Meta-CoT)
6.76 6.30 6.61 7.23
7.73 8.44 8.22 8.53
7.01 6.98 7.18 7.26
7.94 7.71 8.06 8.25
SFT + RL (Ours)
7.44
8.53
7.31
8.34
4.2. Quantitative Evaluation As shown in Table 2 and Table 3, our method achieves notable improvements in the overall editing score, which considers instruction following, consistency, and visual quality. Compared to Bagel (no-think), it achieves +13.1% on the 21-task benchmark and +19.7% on ImgEdit. Relative to Bagel (think), gains reach 20.1% and 13.0%, respectively. To isolate the contribution of the Meta-CoT paradigm, we also train a variant using identical data and optimized with the same parameters, excluding the training of MetaCoT. As shown in Table 2, our method outperforms this setting by 15.8%, validating the effectiveness of Meta-CoT. The RL stage also further enhances alignment and stability. Table 4 provides a breakdown across the four dimensions of VIEScore, Instruction Following, Subject Consistency, Naturalness, and Artifacts, where our method consistently outperforms the baselines, with the largest improvement in Instruction Following. This suggests that Meta-CoT reasoning enhances semantic understanding of both editing operations and targets, leading to more instruction-faithful edits. At the per-task level (as shown in Table 2), our method improves performance on all tasks except text editing. We observe that the reasoning process tends to hinder text editing, likely because the extensive textual reasoning interferes with identifying the correct text to modify. Developing mechanisms to preserve accurate text perception during reasoning remains a promising direction for future work.
4.3. Abaltion Study We further investigate three critical questions related to our method and present the results in Table 5 and Table 6. Can the meta-task training paradigm enable generalization to unseen editing tasks, and how many metatasks are needed? As shown in Table 5, we replace the first step of the CoT (“Task Summary”) with Meta-task Summary and conduct training under five distinct settings, each corresponding to different definitions of meta-tasks and the number of training task types. Starting from the basic 3-meta-tasks setting (add, delete, replace), we gradually increase the number of meta-tasks. The results show that, first, the model trained only on the five meta-tasks already achieves performance comparable to the full-data model on the 21-task benchmark and significantly outper-
Table 5. Ablation study on (1) the number of meta-tasks defined and tasks trained, and (2) the Task Thinking in MetaCoT. (n meta) denotes defining n meta-tasks and training only on them. (5 meta, full-task) indicates defining 5 meta-tasks, training on full tasks, and decomposing each task’s data into meta-tasks. Method
Ins.
Con.
Nat.
Art.
Train Editing Only
6.61
8.22
7.18
8.06
SFT(3 meta) SFT(4 meta) SFT(5 meta) SFT(6 meta) SFT(5 meta, full-task)
6.75 6.93 7.09 7.13 7.20
8.33 8.44 8.48 8.51 8.49
7.15 7.17 7.20 7.22 7.23
7.86 7.94 8.10 8.07 8.12
SFT(w/o task think)
6.98
8.35
7.19
8.07
SFT(Meta-CoT) SFT + RL(Ours)
7.23 7.44
8.53 8.53
7.26 7.31
8.25 8.34
Table 6. Ablation study on the amount of visual understanding data mixed during training. w/o und. denotes training without mixing visual understanding data. CoT measures the completeness and accuracy of Task Thinking and Target Editing Traversal. Method
Ins.
Con.
Nat.
Art.
CoT
Train Editing Only
6.61
8.22
7.18
8.06
-
SFT(w/o und.) SFT(1k und.)
6.74 6.92
8.28 8.31
7.18 7.20
8.11 8.16
7.56 7.81
SFT(Meta-CoT)
7.23
8.53
7.26
8.25
8.89
forms the train-edit-only version. This demonstrates the strong generalization capability of the meta-task training strategy: training on a small set of meta-tasks while learning task decomposition suffices to generalize to unseen tasks. In other words, mastering universal meta-tasks and task decomposition reasoning enhances the model’s ability to generalize. Second, results show that our defined five metatasks strike a good balance between generalization and performance: defining fewer meta-tasks leads to a significant drop in instruction-following, while defining more metatasks provides little additional improvement across all tasks. Does the Task Thinking in Meta-CoT benefit the editing process? As shown in Table 5, we compare our method with a variant that removes the Task Thinking (the second step of Meta-CoT). Results show a significant drop in instruction-following performance, confirming that reasoning based on task characteristics is crucial for the editing. How does joint training with understanding data affect editing performance? In Table 6, we compare two reduced settings: (a) removing all understanding data and (b) using only 1K samples, against the default 100K. In addition to the four VIEScore components, GPT-4.1 also evaluates CoT quality in terms of the completeness and accuracy of Task Thinking and Target Editing Traversal in Meta-CoT. Results show that both reduced settings cause significant drops in editing performance, particularly in in-
Input Image
Bagel
Train Edit Only
Ours
Input Image
Bagel
Train Edit Only
… visual characteristics of red bricks are: They have a warm orange color. They appear solid, modular, and structured, with mortar lines …
Material
Build the horse using red bricks. … Target editing mode traversal: - The guitar: Remove the guitar and replace it with a can of Coca-Cola. – Coca-Cola: Add a can of Coca-Cola to the hand of the fox. The fox: No changes …
Replace
Action
Style
Transfer the image into a folded-paper origami art style. This is a Complex Instruction task, which includes: Replace the baseball glove … Target editing mode traversal: - The baseball glove: … - The baseball: … - The blanket: …
Replace the guitar in front of the fox with a can of Coca-Cola. This is a Motion Change task. Raise the woman’s slightly lowered head and adjust her gaze to look straight ahead …
Complex Instruction
Replace the baseball glove with a ruby-adorned smooth golden artifact of similar size, and add a glowing sparkle effect alongside faint golden specks in the air. Remove the baseball, replace the blanket with a fine marble-tiled table surface, and adjust the scene's lighting to convey a sophisticated ambiance.
Change the person's movements to look forward. … is a Quantity Change task … Target editing mode traversal: - Sky: Add three clouds, leaving the rest unchanged … - Ground: No changes. - Sea: No changes …
… Target editing mode traversal: - Domes: Replace the color of all domes with a gradient of blue shades. - Castle: No changes to any parts other than the domes. …
Color
Change the color of the domes to gradient of blue shades. This is a Background Change task … Replace the background with pure white. … editing mode traversal: - Man: No changes. - Microphone: No changes …
Extract
Quantity
Add 3 clouds to the sky. … Since 1+1=2, …the edit can be summarized as: Replace the question mark with number 2 …
Logical
Draw what it will look like after the question mark is replaced with the correct answer. This is a Camera Motion task … Global changes: Rotate the camera to the right side of the boat … New subjects that appear: - Man: Show his side view. - Box: Display it more completely …
Extract the human figure standing in the image along with their clothing and visible accessories, separating them from the background environment. … The tadpole will turn into a frog … the edit can be summarized as: replace the tadpole with a frog …
Temporal
Draw what it will look like three months later.
Viewpoint
Change the camera view to the side. … Target editing mode traversal: - The windows: Replace the windows with open ones, showing outside view, including trees, a lawn, and a house. - The bathtub: Replace the reflections on the water surface with reflections of the sky and window …
… Target editing mode traversal: - Top shoe: Rotate 180 degrees horizontally, with the toe pointing to the left. - Bottom shoe: No changes. - Background: No changes …
Spatial
Ours … visual characteristics of the stained-glass cathedral window are: objects made of angular folded planes, sharp creases, geometric facets, and hard-edged shadows along folds …
Rotate the top shoe 180 degrees horizontally so its toe faces the heel of the bottom shoe.
Structure
Open the shoji screens in the living room and bedroom areas to reveal the view outside.
Figure 6. Qualitative results across diverse editing tasks, including conventional editing, reasoning-based editing, and multi-instruction editing (Zoom in to view). We present a partial visualization of the Meta-CoT reasoning process. Meta-CoT can decompose instructions, categorize them into specific tasks, generate reasoning based on the task characteristics, and accurately determine whether each target should be edited or not, ultimately achieving better editing results. Please see the supplementary materials for additional task examples.
struction following, as limited understanding data weakens the model’s comprehension of editing instructions. This is further corroborated by the notable decline in CoT quality, underscoring the necessity of balancing all three triplet elements during training and demonstrating that higherquality Meta-CoT reasoning leads to better editing results.
4.4. Qualitative Evaluation As shown in Figure 6, we present comparisons across diverse editing tasks, including conventional, reasoningbased editing, and multi-instruction editing. Our method significantly improves instruction following, logical reasoning, and multi-instruction understanding compared with baseline methods. This demonstrates that our approach more effectively activates and leverages the model’s inherent understanding capability during the editing process.
5. Conclusion In this paper, we have investigated the problem of how to simultaneously enhance the understanding granularity and generalization capability of Chain-of-Thought (CoT)guided image editing. To address this, we have presented Meta-CoT, which first employs the Triplet Decomposition to stimulate the model’s reasoning ability from both task and target perspectives. Furthermore, we have proposed the Meta-task Decomposition, which endows Meta-CoT with strong generalization capability across diverse editing scenarios. To align the CoT reasoning with the editing behavior, we have introduced the CoT-Editing Consistency Reward. Extensive experiments on our proposed 21-task benchmark and ImgEdit have demonstrated that our method not only significantly improves editing performance but also exhibits strong generalization to unseen editing tasks.
Acknowledgments. This work was supported part by the Guangdong Natural Science Funds for Distinguished Young Scholar (No. 2025B1515020012).
References [1] Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023. 4 [2] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923, 2025. 3, 4, 5 [3] Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In CVPR, pages 18392–18402, 2023. 6 [4] Jiuhai Chen, Zhiyang Xu, Xichen Pan, Yushi Hu, Can Qin, Tom Goldstein, Lifu Huang, Tianyi Zhou, Saining Xie, Silvio Savarese, et al. Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset. arXiv preprint arXiv:2505.09568, 2025. 4 [5] Jiuhai Chen, Le Xue, Zhiyang Xu, Xichen Pan, Shusheng Yang, Can Qin, An Yan, Honglu Zhou, Zeyuan Chen, Lifu Huang, et al. Blip3o-next: Next frontier of native image generation. arXiv preprint arXiv:2510.15857, 2025. 6 [6] Xiaokang Chen, Chengyue Wu, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, and Ping Luo. Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811, 2025. 4 [7] Wenxun Dai, Zhiyuan Zhao, Yule Zhong, Yiji Cheng, Jianwei Zhang, Linqing Wang, Shiyi Zhang, Yunlong Lin, Runze He, Fellix Song, et al. Chatumm: Robust context tracking for conversational interleaved generation. arXiv preprint arXiv:2602.06442, 2026. 4 [8] Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al. Emerging properties in unified multimodal pretraining. arXiv preprint arXiv:2505.14683, 2025. 2, 3, 4, 6 [9] Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, et al. Dreamllm: Synergistic multimodal comprehension and creation. In ICLR, 2024. 4 [10] Chengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang, Linjiang Huang, Xingyu Zeng, Hongsheng Li, and Xihui Liu. Got-r1: Unleashing reasoning capability of mllm for visual generation with reinforcement learning. arXiv preprint arXiv:2505.17022, 2025. 4 [11] Rongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang, Hao Li, Shilin Yan, Hao Tian, Xingyu Zeng, Rui Zhao,
Jifeng Dai, et al. Got: Unleashing reasoning capability of multimodal large language model for visual generation and editing. arXiv preprint arXiv:2503.10639, 2025. 2, 4, 6 [12] Xingyu Fu, Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Jianwei Yang, Dan Roth, Dinei Florencio, and Cha Zhang. Refocus: Visual editing as a chain of thought for structured image understanding. arXiv preprint arXiv:2501.05452, 2025. 4 [13] Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, and Xiang Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. arXiv preprint arXiv:2412.05237, 2024. 6 [14] Tanmay Gupta and Aniruddha Kembhavi. Visual programming: Compositional visual reasoning without training. In CVPR, pages 14953–14962, 2023. 4 [15] Feng Han, Yang Jiao, Shaoxiang Chen, Junhao Xu, Jingjing Chen, and Yu-Gang Jiang. Controlthinker: Unveiling latent semantics for controllable image generation through visual reasoning. arXiv preprint arXiv:2506.03596, 2025. 4 [16] Runze He, Kai Ma, Linjiang Huang, Shaofei Huang, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, and Si Liu. Freeedit: Mask-free reference-based image editing with multi-modal instruction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 4 [17] Runze He, Yiji Cheng, Tiankai Hang, Zhimin Li, Yu Xu, Zijin Yin, Shiyi Zhang, Wenxun Dai, Penghui Du, Ao Ma, et al. Re-align: Structured reasoning-guided alignment for in-context image generation and editing. arXiv preprint arXiv:2601.05124, 2026. 4 [18] Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. NeurIPS, 37:139348–139379, 2024. 4 [19] Yujia Hu, Songhua Liu, Zhenxiong Tan, Xingyi Yang, and Xinchao Wang. Image editing as programs with diffusion models. arXiv preprint arXiv:2506.04158, 2025. 6 [20] Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. Large language models can self-improve. arXiv preprint arXiv:2210.11610, 2022. 4 [21] Wenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao, Shixiang Tang, Yufan Shen, Qingyu Yin, Wenbo Hu, Xiaoman Wang, Yuntian Tang, et al. Interleaving reasoning for better text-to-image generation. arXiv preprint arXiv:2509.06945, 2025. 4 [22] Ziyuan Huang, DanDan Zheng, Cheng Zou, Rui Liu, Xiaolong Wang, Kaixiang Ji, Weilong Chai, Jianxin Sun, Libin Wang, Yongjie Lv, et al. Ming-univision: Joint image understanding and generation with a unified continuous tokenizer. arXiv preprint arXiv:2510.06590, 2025. 4, 6 [23] Dongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong, Hao Li, Le Zhuo, Shilin Yan, Pheng-Ann Heng, and Hongsheng Li. T2i-r1: Reinforcing image generation with collaborative semantic-level and token-level cot. arXiv preprint arXiv:2505.00703, 2025. 4
[24] Ivan Krasin, Tom Duerig, Neil Alldrin, Andreas Veit, Sami Abu-El-Haija, Serge Belongie, David Cai, Zheyun Feng, Vittorio Ferrari, Victor Gomes, Abhinav Gupta, Chen Sun, Gal Chechik, Kevin Murphy, Dhyanesh Narayanan, Saurabh Shetty, Yang Song, Joseph Tighe, Andrea Vedaldi, Sudheendra Vijayanarasimhan, and Oriol Vinyals. Openimages: A public dataset for large-scale multi-label and multi-class image classification. Dataset available fromnhttps://storage.googleapis.com/openimages/web/index.html, 2017. https : / / storage . googleapis . com / openimages/web/factsfigures.html. 6 [25] Max Ku, Dongfu Jiang, Cong Wei, Xiang Yue, and Wenhu Chen. Viescore: Towards explainable metrics for conditional image synthesis evaluation. arXiv preprint arXiv:2312.14867, 2023. 5, 6 [26] Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, et al. Flux. 1 kontext: Flow matching for in-context image generation and editing in latent space. arXiv preprint arXiv:2506.15742, 2025. 6 [27] Seongyun Lee, Seungone Kim, Sue Park, Geewook Kim, and Minjoon Seo. Prometheus-vision: Vision-language model as a judge for fine-grained evaluation. In Findings of the association for computational linguistics ACL 2024, pages 11286–11315, 2024. 5 [28] Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 6 [29] Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vulić, and Furu Wei. Imagine while reasoning in space: Multimodal visualization-ofthought. arXiv preprint arXiv:2501.07542, 2025. 4 [30] Lingen Li, Guangzhi Wang, Zhaoyang Zhang, Yaowei Li, Xiaoyu Li, Qi Dou, Jinwei Gu, Tianfan Xue, and Ying Shan. Tooncomposer: Streamlining cartoon production with generative post-keyframing. arXiv preprint arXiv:2508.10881, 2025. 4 [31] Lingen Li, Zhaoyang Zhang, Yaowei Li, Jiale Xu, Wenbo Hu, Xiaoyu Li, Weihao Cheng, Jinwei Gu, Tianfan Xue, and Ying Shan. Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 777–787, 2025. 4 [32] Weixin Liang, Lili Yu, Liang Luo, Srinivasan Iyer, Ning Dong, Chunting Zhou, Gargi Ghosh, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, et al. Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models. arXiv preprint arXiv:2411.04996, 2024. 4 [33] Chao Liao, Liyang Liu, Xun Wang, Zhengxiong Luo, Xinyu Zhang, Wenliang Zhao, Jie Wu, Liang Li, Zhi Tian, and Weilin Huang. Mogao: An omni foundation model for interleaved multi-modal generation. arXiv preprint arXiv:2505.05472, 2025. 4 [34] Bin Lin, Zongjian Li, Xinhua Cheng, Yuwei Niu, Yang Ye, Xianyi He, Shenghai Yuan, Wangbo Yu, Shaodong Wang,
Yunyang Ge, et al. Uniworld-v1: High-resolution semantic encoders for unified visual understanding and generation. arXiv preprint arXiv:2506.03147, 2025. 6 [35] Yunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen, Zhongdao Wang, Xinghao Ding, Wenbo Li, et al. Jarvisart: Liberating human artistic creativity via an intelligent photo retouching agent. arXiv preprint arXiv:2506.17612, 2025. 4 [36] Yunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin, Kaixiong Gong, Wenbo Li, Bin Lin, Zhenxi Li, Shiyi Zhang, Yuyang Peng, et al. Jarvisevo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization. arXiv preprint arXiv:2511.23002, 2025. 4 [37] Jie Liu, Gongye Liu, Jiajun Liang, Yangguang Li, Jiaheng Liu, Xintao Wang, Pengfei Wan, Di Zhang, and Wanli Ouyang. Flow-grpo: Training flow matching models via online rl. arXiv preprint arXiv:2505.05470, 2025. 3, 5 [38] Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, Chunrui Han, et al. Step1x-edit: A practical framework for general image editing. arXiv preprint arXiv:2504.17761, 2025. 3, 5, 6 [39] Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi. Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action. In CVPR, pages 26439–26455, 2024. 4 [40] Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. NeurIPS, 35: 2507–2521, 2022. 4 [41] Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Siran Chen, Xiu Li, and Qifeng Chen. Follow your pose: Poseguided text-to-video generation using pose-free videos. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 4117–4125, 2024. 4 [42] Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Xingkai Yu, Liang Zhao, Yisong Wang, Jiaying Liu, and Chong Ruan. Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation. arXiv preprint arXiv:2411.07975, 2024. 4 [43] Yue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng, Hongyu Liu, Jiayi Guo, Mark Fong, Yuxuan Xue, Zixiang Zhao, Konrad Schindler, et al. Fastvmt: Eliminating redundancy in video motion transfer. arXiv preprint arXiv:2602.05551, 2026. 4 [44] OpenAI. Gpt-image-1, 2025. 6 [45] OpenAI. Introducing 4o image generation, 2025. 4 [46] Xichen Pan, Satya Narayan Shukla, Aashu Singh, Zhuokai Zhao, Shlok Kumar Mishra, Jialiang Wang, Zhiyang Xu, Jiuhai Chen, Kunpeng Li, Felix Juefei-Xu, Ji Hou, and Saining Xie. Transfer between modalities with metaqueries. arXiv preprint arXiv:2504.06256, 2025. 4 [47] William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter-
national conference on computer vision, pages 4195–4205, 2023. 4 [48] Liao Qu, Huichao Zhang, Yiheng Liu, Xu Wang, Yi Jiang, Yiming Gao, Hu Ye, Daniel K Du, Zehuan Yuan, and Xinglong Wu. Tokenflow: Unified image tokenizer for multimodal understanding and generation. arXiv preprint arXiv:2412.03069, 2024. 4 [49] Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS, 35:25278– 25294, 2022. 6 [50] Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. NeurIPS, 37:8612–8642, 2024. 4 [51] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. 5 [52] Weijia Shi, Xiaochuang Han, Chunting Zhou, Weixin Liang, Xi Victoria Lin, Luke Zettlemoyer, and Lili Yu. Llamafusion: Adapting pretrained language models for multimodal generation. arXiv preprint arXiv:2412.15188, 2024. 4 [53] Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Emu: Generative pretraining in multimodality. In ICLR, 2024. 4 [54] Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024. 4 [55] Henri Theil. On the estimation of relationships involving qualitative variables. American Journal of Sociology, 76(1): 103–154, 1970. 4 [56] Shengbang Tong, David Fan, Jiachen Zhu, Yunyang Xiong, Xinlei Chen, Koustuv Sinha, Michael Rabbat, Yann LeCun, Saining Xie, and Zhuang Liu. Metamorph: Multimodal understanding and generation via instruction tuning. arXiv preprint arXiv:2412.14164, 2024. 4 [57] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022. 4 [58] Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need. arXiv preprint arxiv:2409.18869, 2024. 4 [59] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. NeurIPS, 35:24824–24837, 2022. 4 [60] Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai
Yu, Chong Ruan, et al. Janus: Decoupling visual encoding for unified multimodal understanding and generation. In CVPR, pages 12966–12977, 2025. 4 [61] Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng-ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al. Qwen-image technical report. arXiv preprint arXiv:2508.02324, 2025. 6 [62] Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao, Xin Luo, Yueze Wang, Wanli Li, Xiyan Jiang, Yexin Liu, Junjie Zhou, et al. Omnigen2: Exploration to advanced multimodal generation. arXiv preprint arXiv:2506.18871, 2025. 4, 6 [63] Penghao Wu and Saining Xie. V?: Guided visual search as a core mechanism in multimodal llms. In CVPR, pages 13084–13094, 2024. 4 [64] Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. Next-gpt: Any-to-any multimodal llm. In Forty-first ICML, 2024. 4 [65] Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Chaofan Li, Shuting Wang, Tiejun Huang, and Zheng Liu. Omnigen: Unified image generation. In CVPR, pages 13294–13304, 2025. 6 [66] Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528, 2024. 4 [67] Jinheng Xie, Zhenheng Yang, and Mike Zheng Shou. Showo2: Improved native unified multimodal models. arXiv preprint arXiv:2506.15564, 2025. 4 [68] Yu Xu, Fan Tang, Juan Cao, Yuxin Zhang, Xiaoyu Kong, Jintao Li, Oliver Deussen, and Tong-Yee Lee. Headrouter: A training-free image editing framework for mmdits by adaptively routing attention heads. arXiv preprint arXiv:2411.15034, 2024. 4 [69] Yu Xu, Hongbin Yan, Juan Cao, Yiji Cheng, Tiankai Hang, Runze He, Zijin Yin, Shiyi Zhang, Yuxin Zhang, Jintao Li, et al. Tag-moe: Task-aware gating for unified generative mixture-of-experts. arXiv preprint arXiv:2601.08881, 2026. 4 [70] An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024. 5 [71] Shiyuan Yang, Xiaodong Chen, and Jing Liao. Uni-paint: A unified framework for multimodal image inpainting with pretrained diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia, pages 3190–3199, 2023. 4 [72] Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao. Direct-a-video: Customized video generation with user-
directed camera movement and object motion. In ACM SIGGRAPH 2024 Conference Papers, pages 1–12, 2024. 4 [73] Siwei Yang, Mude Hui, Bingchen Zhao, Yuyin Zhou, Nataniel Ruiz, and Cihang Xie. Complex-edit: Cot-like instruction generation for complexity-controllable image editing benchmark. arXiv preprint arXiv:2504.13143, 2025. 3, 6 [74] Michihiro Yasunaga, Luke Zettlemoyer, and Marjan Ghazvininejad. Multimodal rewardbench: Holistic evaluation of reward models for vision language models. URL https://api. semanticscholar. org/CorpusID, 276482127, 2025. 5 [75] Yang Ye, Xianyi He, Zongjian Li, Bin Lin, Shenghai Yuan, Zhiyuan Yan, Bohan Hou, and Li Yuan. Imgedit: A unified image editing dataset and benchmark. arXiv preprint arXiv:2505.20275, 2025. 3, 6 [76] Qifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, and Yueting Zhuang. Anyedit: Mastering unified high-quality image editing for any idea. In CVPR, pages 26125–26135, 2025. 6 [77] Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, and Yu Su. Magicbrush: A manually annotated dataset for instructionguided image editing. NeurIPS, 36:31428–31449, 2023. 6 [78] Shiyi Zhang, Wenxun Dai, Sujia Wang, Xiangwei Shen, Jiwen Lu, Jie Zhou, and Yansong Tang. Logo: A long-form video dataset for group action quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2405–2414, 2023. 4 [79] Shiyi Zhang, Sule Bai, Guangyi Chen, Lei Chen, Jiwen Lu, Junle Wang, and Yansong Tang. Narrative action evaluation with prompt-guided multimodal interaction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18430–18439, 2024. 4 [80] Shiyi Zhang, Junhao Zhuang, Zhaoyang Zhang, Ying Shan, and Yansong Tang. Flexiact: Towards flexible action control in heterogeneous scenarios. In Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, pages 1–11, 2025. 4 [81] Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. Multimodal chain-ofthought reasoning in language models. arXiv preprint arXiv:2302.00923, 2023. 4 [82] Zechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang, and Yi Yang. In-context edit: Enabling instructional image editing with incontext generation in large scale diffusion transformer. arXiv preprint arXiv:2504.20690, 2025. 6 [83] Haozhe Zhao, Xiaojian Shawn Ma, Liang Chen, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li, and Baobao Chang. Ultraedit: Instruction-based fine-grained image editing at scale. NeurIPS, 37:3058–3093, 2024. 6 [84] Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Xiaorong Zhu, Hao Li, Wenhao Chai, Zicheng Zhang, Renqiu Xia, Guangtao Zhai, Junchi Yan, et al. Envisioning beyond the pixels: Benchmarking reasoning-informed visual editing. arXiv preprint arXiv:2504.02826, 2025. 3, 5 [85] Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe
Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. arXiv preprint arxiv:2408.11039, 2024. 4 [86] Tianrui Zhu, Shiyi Zhang, Jiawei Shao, and Yansong Tang. Kv-edit: Training-free image editing for precise background preservation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16607–16617, 2025. 4 [87] Junhao Zhuang, Xuan Ju, Zhaoyang Zhang, Yong Liu, Shiyi Zhang, Chun Yuan, and Ying Shan. Colorflow: Retrievalaugmented image sequence colorization. arXiv preprint arXiv:2412.11815, 2024. 4