MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments Qingyun Liu1∗ , Jiwen Zhang1∗ , Jingyi Hu1 , Siyuan Wang3† , Zhongyu Wei1,2† 1 Fudan University, 2 Shanghai Innovation Institute, 3 The Chinese University of Hong Kong {qyliu25,jiwenzhang21,jyhu25}@m.fudan.edu.cn [email protected], [email protected]
arXiv:2606.31966v1 [cs.MA] 30 Jun 2026
Abstract
Goal: Find and put 2 forks, two waterglass and 1 wineglass on kitchen table
Single Agent
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark with an evaluation platform spanning diverse real-world tasks, two cooperation structures, and three collaboration modes. Through extensive experiments across various MLLMs, we summarize three key findings: (i) Collaboration generally improves embodied task completion, but its benefits depend on balancing collaborative gains against coordination complexity. (ii) Communication is essential to collaboration gains, while the best collaboration mode depends on team size and model capability. (iii) Moreover, collaboration improves robustness under noisy priors and exploration conditions. Generally, MECoBench provides a systematic testbed for understanding the mechanisms and limits of multimodal embodied collaboration. Code and dataset are available at https:/github.com/q-i-n-g/MECoBench.
1
Explored by agent Unexplored
❌ 60 1/5
Task progress
3-Agent Team
stuck in middle Explored by agent 0 Explored by agent 1 Explored by agent 2
5/5
Task progress
✅ 18
Figure 1: An illustration of realistic scenarios, where cross-modal multi-agent collaboration significantly improves the efficiency compared with single-agent.
on pure language-based multi-agent framework (Schmidgall et al., 2025; Hong et al., 2024) have proved that such cooperation could improve efficiency and overcome individual capability limits. However, the potential of multi-agent collaboration under multimodal embodied settings remains an underexplored area. Answering this is non-trivial, because embodied multi-agent collaboration differs fundamentally from text-based collaboration. Existing multiagent benchmarks rely on textual task descriptions, shared symbolic states, or pre-processed observations, reducing collaboration to plan coordination or dialogue (Zhu et al., 2025; Sun et al., 2025; Zhang et al., 2024a; Yang et al., 2026; Agashe et al., 2025). By contrast, embodied agents must coordinate from partial, dynamically changing visual observations while exploring environments, resolving perceptual inconsistencies, avoiding spatial conflicts, and acting under physical constraints (Shridhar et al., 2020). This tightly couples percep-
Introduction
Recent multimodal large language models (MLLMs) (Anthropic, 2026; OpenAI, 2026; Google DeepMind, 2026a) have demonstrated strong vision-language understanding and reasoning capabilities, making them promising foundations for embodied agents in interactive environments (Mu et al., 2023; Szot et al., 2025). While most existing MLLM-based studies focus on single-agent embodied intelligence (Yang et al., 2025; Liu et al., 2025; Zhang et al., 2025b), many real-world tasks such as household assistance require multiple agents to cooperate (Sycara, 1998), as shown in Figure 1. Previous works * Equal contribution. †
Team step Thr. 60 steps
Corresponding author.
1
Table 1: Comparison with related benchmarks. Multimodal: Whether agents receive multimodal observations rather than text-only observations. Collaboration structure (Structure): I and D denote independent and interdependent collaboration. Collaboration mode (Collab.): I, C, and D denote isolated, centralized, and decentralized collaboration. Communication medium (Comm.): T and V denote textual and visual communication. “–” indicates no explicit inter-agent communication. Benchmark
Env. Multimodal
Team size
FurnMove (Jain et al., 2020) CoELA (Zhang et al., 2024a) RoCoBench (Mandi et al., 2024) VillagerBench (Dong et al., 2024) PARTNR (Chang et al., 2025) TeamCraft (Long et al., 2024) Collab-Overcooked (Sun et al., 2025) MineCollab (White et al., 2025) COOP2 (Yang et al., 2026)
3D 3D 3D 3D 3D 3D 2D 3D 2D
✓ × × × × ✓ × × ×
Fixed (2) Fixed (2) Fixed (2–3) Variable (2–3) Fixed (2) Variable (2–4) Fixed (2) Variable (2–5) Variable (3/6)
D I I/D I/D I I/D D I I/D
D D D C C/D C/D D D I/C/D
– T T – – – T T T
1,000 44 6 225 1,000 950 30 184 36
MECoBench
3D
✓
Variable (1–5)
I/D
I/C/D
T/V
192
tion, exploration, communication, and coordination, leading to failure modes and emergent behaviors that text-based benchmarks cannot capture (Feng et al., 2026). Therefore, it is valuable to systematically study MLLM-based multi-agent collaboration in embodied settings, calling for a new benchmark that supports visually grounded interaction and diverse cooperation patterns.
Structure Collab. Comm. #Test cases
bust under challenging information conditions? To answer these questions, we evaluate a wide range of MLLMs across model families and parameter scales, and obtain the following findings: • Multi-agent collaboration generally improves embodied task completion, but its benefits depend on balancing gains against coordination complexity. Most MLLMs exhibit basic collaborative capability and benefit from a second agent participation. As the team grows further, however, performance saturates and then degrades, forming an inverted-U trend and indicating that moderate team sizes provide the best trade-off between parallelism and coordination overhead. Encouragingly, larger teams remain more robust as task complexity increases. • Communication is the key driver of collaboration gains, and the best mode is contingent on both team size and model capability. Removing communication consistently degrades performance, with substantially larger drops on sequential tasks and larger teams. Centralized coordination is most effective in small teams but saturates earlier than decentralized coordination as the team scales. Beyond textual messaging, shared memory yields more efficient collaboration and benefits loosely coupled parallel tasks, while vision-augmented leadership improves coordination efficiency—provided the leader model is capable enough to leverage visual input. • Collaboration further enhances robustness under challenging information conditions. Twoagent collaboration still improves task completion when location priors are removed and agents must explore from scratch, and yields its largest
To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark together with an evaluation platform for systematically analyzing visually grounded embodied multiagent cooperation. MECoBench contains diverse embodied tasks covering eight types of common real-world activities, and introduces two cooperation structures to support realistic evaluation. In parallel cooperation, agents can work on different subtasks simultaneously, whereas in spaceconstrained sequential cooperation, agents must coordinate in a chained manner to complete the task. Beyond task construction, our platform supports controlled variation along multiple dimensions, including collaboration mode, team size and exploration difficulty. This design enables fine-grained analysis of not only whether collaboration helps, but also when, why, and through which mechanisms it improves embodied task solving. Based on MECoBench, our work aims to answer three research questions progressively: (i) To what extent does multi-agent collaboration benefit MLLMs on embodied tasks, and how does this benefit change as the team size scales? (ii) How do different collaboration modes and communication mechanisms shape collaboration outcomes? (iii) Does collaboration remain effective and ro2
Task template
1
Work at desk: mouse, mug, cellphone and keyboard on the desk
Task Scene Grounding Legal placement Surface: coffee table, sofa … Container: cabinet, fridge, bookshelf, microwave, …
Goal State Goal sampling
Initial Scene
1 mouse, 0 cellphone, 1 keyboard, 2 mug
Parallel
Agents move freely in all rooms
Scene assets
Task Goal
Object Location
Task Scene Collaboration Structure
8 different apartments living room/kitchen/ bedroom/ bathroom
Agent 1: kitchen, bedroom Agent 2: bathroom, living room Kitchen living room
Sequential
2 Task Configuration Figure 2: Data construction pipeline of MECoBench. Each task is first grounded from a high-level task into a concrete scene, then set collaboration configuration for parallel or sequential execution.
relative gains under noisy priors, showing that multi-agent teams can compensate for misleading information through communication and distributed exploration. Available task information further amplifies the benefits of collaboration.
2
no comm
isolated
discuss
decentralized
one as leader
centralized
Figure 3: Illustration of four protocols under three collaboration modes.
MECoBench plate, and then randomly place the selected goal objects at legal initial locations. We define a broad range of surfaces and containers to enable diverse and realistic object distributions (listed in Table 5). After a task scene is instantiated, we further assign its collaboration structure. Under sequential settings, the apartment is partitioned into disjoint room zones based on the initial locations of task objects and their target goal locations, and different zones are assigned to different agents. Overall, MECoBench consists of 96 tasks, each evaluated under both parallel and sequential setups, yielding 192 test cases in total. Each task contains 2–7 subgoals. The scenes are uniformly distributed over both the number of subgoals and the eight task templates to ensure balanced task complexity and scenario diversity. Details of construction and statistic are in Appendix A.
To investigate MLLM-based multi-agent collaboration under embodied settings, we propose MECoBench, a comprehensive benchmark with an interactive evaluation platform. MECoBench is built upon VirtualHome (Puig et al., 2018), a realistic household simulator covering a wide range of everyday scenarios, and spans different collaboration structures, variable team sizes, diverse collaboration modes and communication mechanisms. Table 1 compares MECoBench with existing embodied multi-agent collaboration benchmarks. 2.1
broadcast
Benchmark Construction
To cover common real-world activities, we define eight semantic task templates (detailed definition in Table 4), including five adapted from WAH (Puig et al., 2021) and three newly designed. Beyond semantic diversity, we introduce two collaboration structure to capture different forms of spatial accessibility in the real world. In the parallel setting, all agents can freely access all rooms and complete subgoals independently and concurrently, modeling cases where individuals operate in a shared space. In the sequential setting, agents are assigned to disjoint room zones, making inter-agent object transfers necessary and reflecting situations where individuals have access to different regions and must coordinate across spatial boundaries. The final task is constructed through a two-stage pipeline shown in Figure 2. In the scene grounding stage, we firstly sample subgoals from a task tem-
2.2
Evaluation Platform
The platform supports a flexible number of agents and provides three collaboration modes with four basic protocols to enable systematic evaluation. Collaboration Mode To analyze how different coordination structures shape multi-agent performance, we design three collaboration modes implemented with four protocols as summarized in Figure 3. The isolated mode disables communication, requiring agents to act independently. The decentralized mode enables peer-level coordination through broadcast or discussion. In broadcast, 3
Object location
Observation
front
Object appearance
Agent 1
90°
Action trajectory [walk] [grab]
right
1
back
No comm
I'll open the fridge to check juice and wine….
Broadcast
I'll check the bedroom cabinet for the pudding
living room Legal action walk, grab, open, close, put on/ inside, handover, wait
Comm. summary
Discuss
4
Global info
1 4
3
Shared memory block Action exec
Open the grey fridge Walk to the wooden cabinet near the bed walk to kitchen
I’ll go to the kitchen to check the table…
Leader
1 2
Init location
bedroom
Task object record
Dialogue window
90°
Agent 2
kitchen
Room exploration
…
Width: 10 rounds left
Agent 0
Memory
History
1.8m
Initialize Place pudding, cupcake, juice and wine on the coffee table
[walk]<kitchen>(11)
1 Communicate
2 Think & Act
✅
success
❌
fail
Task progress
1 /4
Hand occupy null Agent location
Feedback
Virtual Home
Figure 4: Overall workflow of evaluation. Agents receive task goals and prior information, perceive the environment, communicate under different protocols, reason with history and memory, and execute actions with feedback.
where πi denotes the action policy of agent i. The selected actions are executed in simulator, which returns execution feedback Ft and resulting environment changes. Finally, the feedback is used to update the history and memory. This workflow continues until the task is completed or the maximum number of steps is reached.
each agent shares one message with all others before action; in discussion, agents communicate sequentially and can continue for another round when consensus is not reached. The centralized mode assigns one agent as leader to aggregate worker reports and assign actions to the team. Evaluation Pipeline As illustrated in Figure 4, we formulate the evaluation as an iterative observe– communicate–act workflow. At the beginning of each task, the environment is initialized with the task scene graph and N agents. Each agent is placed in its assigned initial room with the provided task goal G, the available prior information I, and the shared action space A (see Appendix B.3). At each timestep t, each agent i obtains a panoramic environment observation Oti from its first-person view. To support long-horizon execution over dozens of steps, the agent reasons not only over its current observation Oti , but also over its action history hia,t , recent dialogue history hd,t , self-maintained memory Mti , and initial information I. Depending on the chosen collaboration protocol P , each agent i first generates the communication message cit as cit = Pi (Oti , hia,t , hd,t , Mti , G, I).
Experiment Setup
3.1
Implementation Details
For parallel tasks, we compare single-agent and two-agent execution under the same effective 60step budget, defined as the total number of steps across all agents. Thus, for an n-agent team, each agent is allocated 60/n steps. For sequential tasks, we use a fixed two-agent setup with an 80-step budget. Unless otherwise specified, we use the broadcast collaboration protocol in all multi-agent settings. By default, we provide the task object appearance and possible location list to agents as the prior information I at initialization. Since MECoBench focuses on collaboration rather than object search, this prior reduces localization difficulty and better isolates cooperative ability. More detials are in Appendix B.1.
(1)
The messages from all agents form the communication context Ct = {c1t , c2t , . . . , cN t }. After communication, each agent generates its next action conditioned on the current communication context: ait = πi (Oti , Ct , hia,t , Mt , G, I, A),
3
3.2
Evaluation Metrics
We evaluate multi-agent collaboration from three aspects: (1) Effectiveness: Success Rate (SR), defined as the proportion of successful cases, and
(2) 4
Completion Rate (CR), defined as the average fraction of completed subgoals across cases. (2) Efficiency: Step-CR AUC (AUC), measuring the area under the completion-rate curve over equivalent steps, and Average Token Cost per Step. (3) Collaboration quality: Division of Labor (DOL), Conflict Action Rate (CAR), and Handover Failure Rate (HFR). DOL measures the distribution of subgoal completion across agents, calculated as P 1 − i s2i counti P si = . (3) , DOL = 1 − n1 j countj
0.80
Sequential 2-agent SR
Q75
0.00
4.1
Do Models Know How To Collaborate?
0.20
0.40
Q25
0.60
Parallel 1-agent SR
Q75
0.80
1.00
Figure 5: Comparison of model performance between parallel single-agent and sequential two-agent settings. Dashed lines indicate quartile thresholds. Percentage Points
20
Closed source
Open source
SR gain (positive) SR gain (negative)
CR gain (positive) CR gain (negative)
10 0 10 ini 5.4 pro b-vl b-vl b-vl -9b 7b 6b 1b 8b 1b a-4 .6v ash -5-m gpt- ini-3.1- en3-8 n3-32 3-235 en3.5 en3.5-2 ma4-2 ma4-3vl-3.5-3l-3.5-24 llam glm-4 4.6v-fl e n qw qw gem gem ern rnv w glm gem q qw qwe int inte
gpt
Figure 6: Performance change from 1-agent to 2agent under parallel settings. Bars show the absolute change in SR and CR (percentage points)
We evaluate a diverse set of both closed-source and open-source models. The closed-source models include state-of-art model GPT-5.4 (OpenAI, 2026), Gemini-3.1-Pro (Google DeepMind, 2026a), and GPT-5-mini (OpenAI, 2025) The open-source models include Qwen3-VL (Bai et al., 2025), Qwen3.5 (Qwen Team, 2026), Gemma4 (Google DeepMind, 2026b), InternVL 3.5 (Wang et al., 2025), Llama4 (Meta AI, 2025) and GLM 4.6V (Team et al., 2025), covering model sizes from 8B to 235B parameters. Appendix B.2 lists all models evaluated.
Will Multi-agent Collaboration Benefit Embodied Task Completion?
0.20 Q25
Evaluated Models
4
0.40
0.00
where n > 1 is the number of agents, and si denotes the fraction of subgoals completed by agent i among all completed subgoals. CAR measures the average fraction of conflicting actions between agents, such as grabbing or opening the same object. HFR measures the average fraction of failed handovers in sequential tasks. 3.3
0.60
performance, and Gemma4 series performs particularly well on collaboration tasks. In contrast, InternVL3.5 performs well individually but struggles to cooperate, revealing a gap between individual execution and group collaboration. Can models benefit from collaboration when it is optional? Compared with single agent, Figure 6 shows that most models benefit from the participation of a second agent. The improvement is particularly large for weak and mid-level models, likely because their single-agent baselines leave more room for gains. Notably, InternVL3.5 series exhibits a clear performance drop, further confirming its weakness in multi-agent coordination. Detailed results and analyses are provided in Appendix D.1.
Can models collaborate when collaboration is necessary? We compare the model performance under single- and two-agent settings, using parallel single-agent and sequential two-agent results as proxies for individual and collaborative ability, respectively. As demonstrated in Figure 5, singleagent performance on parallel tasks correlates positively with two-agent collaboration performance on sequential tasks, while the sequential success rates are generally lower. This suggests that spatial constraints make collaboration harder than individual execution, while most models still exhibit basic collaborative capability. Among these models, GPT-5.4 and GPT-5-mini achieve the best overall
4.2
Do More Agents Always Help?
To investigate this, we increase the number of agents from one to five under parallel settings with fixed effective step budget. Results are shown in Figure 7(a), where collaboration performance generally follows an inverted-U-shaped trend: adding agents initially improves performance, but further scaling leads to saturation and then degradation. Overall, moderate team sizes strike the best 5
Table 2: Fine-grained collaborative behaviors across all models. Freq. denotes the percentage of cases containing each behavior. ∆SR and ∆CR are percentagepoint differences between cases with and without the behavior. Parallel
Behav. (a)
Freq.
(b)
Info-Plan Act-Del. Info-Loc. Info-Obj. Coord-Q. Correct. Task-Assign
Figure 7: Team size scaling effect. (a) Performance curve of different team size. (b) SR of Qwen3-32B-VL versus the #objects, with fitted trend lines for different team sizes.
balance between collaboration benefits and costs. We further analyze the effect of task complexity using object count as a proxy. Figure 7(b) shows that success rates decline with increasing object count, due to longer trajectories and greater planning difficulty (see Appendix D.6 for full results across all models and task conditions). Nevertheless, multi-agent settings maintain more robust performance under high-complexity tasks, highlighting the value of collaboration in complex embodied scenarios. 4.3
∆SR
Sequential ∆CR
Freq.
∆SR
∆CR
75.6% +4.6↑ +2.8↑ 86.3% +16.9↑ +18.0↑ 57.1% +12.2↑ +7.7↑ 54.3% +13.1↑ +19.2↑ 47.3% +9.1↑ +6.9↑ 68.0% +7.6↑ +13.3↑ 25.9% +5.0↑ +3.9↑ 16.3% -6.7↓ -1.6↓ 17.6% -10.1↓ -1.8↓ 48.5% -15.5↓ -6.7↓ 16.9% -6.7↓ +1.6↑ 6.9% -11.3↓ +7.7↑ 15.7% +1.9≈ +2.9↑ 8.5% -11.5↓ -4.9↓
teammates’ states or locations. It occurs more often in sequential tasks where object transfer requires repeated position checks and serves as an indicator of task and coordination difficulty. Correct. denotes behavioral correction through selfreflection or peer feedback. Despite its negative SR correlation, it consistently improves CR in both settings, suggesting that correction serves as an error-recovery mechanism that improves task progress even when full success is unrecoverable. More analysis provided in Appendix E.1.
How Do Agents Collaborate?
We delve deeper into fine-grained collaborative behaviors to understand how agents collaborate in practice. These behaviors are identified from the agents’ communication and action trajectories using keyword-based matching rules, and are organized into three categories, as reported in Table 2. Examples are provided in Appendix E.1.
5
What Makes Multi-agent Collaboration Effective?
We answer this question along two dimensions: collaboration mode and communication mechanism. We first examine which collaboration mode is most effective. We then study the role of communication by testing whether it is necessary and whether alternative mechanisms further improve collaboration.
Information Sharing Agents share their intended plans (Info-Plan), current locations (InfoLoc), and detected task-relevant objects (Info-Obj). Plan and location sharing are frequent and consistently beneficial, while object discovery reporting shows mixed effects.
5.1
Which Collaboration Mode Works Best?
Two-Agent Collaboration. We compare different collaboration modes beyond basic broadcast mechanism. As shown in Figure 8(a), leader-based collaboration achieves the best overall performance for most models, suggesting that centralized collaboration helps small teams allocate tasks and avoid spatial conflicts. However, GPT-5.4 remains competitive with simple broadcast communication, indicating that stronger models may rely less on structured collaboration mode.
Task Coordination Action-level delegation (ActDel) issues concrete executable requests, such as asking another agent to grab, check, or place, and brings the strongest gains in both parallel and sequential tasks. In contrast, task assignment (TaskAssign) provides coarser responsibility divisions, such as assigning rooms, areas, or subtasks, and offers only marginal benefit in parallel tasks while hurting sequential ones. More specific and actionable communication is easier to translate into successful execution.
Scaling to Larger Teams. We further examine how collaboration modes scale with team size on parallel tasks. Figure 8(b) illustrates that centralized coordination is effective in small teams but saturates as team size increases, while decentralized
Alignment and Correction Coordination query (Coord-Q.) refers to questions that synchronize 6
(b)
(a)
Figure 8: Comparison of collaboration modes. (a) Performance profiles in 2-agent setup on parallel and sequential tasks. (b) SR and CR trends with increasing team size under decentralized broadcast and centralized leader-based.
(a)
(a)
(b)
Figure 9: Ablation study on communication. (a) Performance change when remove communication. (b) Performance across different team sizes.
5.3
Is Textual Communication Enough?
While explicit textual communication is essential, it may be insufficient for embodied collaboration, where agents must share evolving task states and visually grounded observations. We therefore explore two alternative communication mechanisms: shared memory and vision-augmented leadership.
coordination becomes more competitive by avoiding the bottleneck of relying on a single leader to aggregate information and assign actions. Moreover, under centralized coordination, model capability matters, as Qwen3-32B-VL benefits from three agents, whereas Qwen3-8B-VL peaks at two agents and then degrades. Therefore, the best collaboration mode depends on both team size and model capability. Full results across all metrics and team sizes are provided in Appendix D.3. 5.2
(b)
Figure 10: Comparison between broadcast and shared-memory. (a) Relative changes in effectiveness and efficiency over broadcast of Qwen3-VL. (b) Average completion progress over steps under broadcast and shared-memory of Qwen3-32B-VL.
Shared memory leads to efficient and effective collaboration. In shared memory, agents skip explicit message exchange and coordinate through a shared memory block. As shown in Figure 10, this mechanism improves the performance and greatly reduces token consumption, indicating more efficient collaboration (absolute results in Figure 31). It consistently improves success rate on parallel tasks, but degrades on sequential tasks, where precise timing and spatial alignment may still require explicit communication. Nevertheless, shared memory accelerates early progress, reducing more conflicts than broadcast under smaller step budgets.
Is Communication Necessary?
We ablate communication from the broadcast protocol under both parallel and sequential settings. As shown in Figure 9, disabling communication consistently degrades performance across models, with a much larger drop on sequential tasks. When the number of agents grows, the gap between communication and no communication widens. These results indicate that communication is essential for effective collaboration, especially when tasks require stronger coordination or involve larger teams. More details in Appendix D.2.
Vision-augmented leadership helps when the leader is capable. To give the leader more grounded context, in vision-augmented leaderworker collaboration, we allow workers to send 7
Score (%)
Success Rate (SR) 92 90 88 86 84 82
Completion Rate (CR) 100
3.12pp 89.6
98 1.04pp
83.3 82.3
4-31Bwen3-32B-VL Q
leader (text only)
1.74pp
93.8
94
92.0
92
4-31Bwen3-32B-VL Q
Gemma
leader-mm (+visual)
gpt-5-mini qwen3-8b-vl qwen3-32b-vl qwen3-235b-vl qwen3.5-9b gemma4-31b internvl-3.5-38b internvl-3.5-241b llama4
1.37pp
76
97.4 96.4
96
86.5
Gemma
1.04pp
AUC
78
75.0 73.6
74
0.58pp
72
71.2 70.6
70
4-31Bwen3-32B-VL Q
Gemma
Gemma4-31B
0.3
2-agent
gain
+7.7 +21.1 +6.8 +2.8
0.4 0.0 loss
+10.4 +1.5 +1.7 0.2 0.4
0.6
Completion Rate (CR)
Closed source
0.8
Open source
Figure 12: Performance change under parallel settings without prior location information.
both textual reports and current visual observations to the leader. As shown in Figure 11, additional visual information improves AUC for both models, while its impact on SR and CR varies across models. This difference may stem from models’ different abilities to effectively use visual cues. Models with stronger visual capabilities (Gemma4-31B), benefit from the additional visual input, whereas weaker models (Qwen3-32B-VL) drop.
(a)
(b)
Figure 13: Team size scaling effect without prior location information. (a) Performance curve of different team size. (b) SR of Qwen3-32B-VL versus the #objects, with fitted trend lines for different team sizes.
Is Multi-Agent Collaboration Robust?
Table 3: Performance under noisy and incomplete task information. We compare three settings: noisy location priors (W/ Noise), clean location priors (W/o Noise), and no location priors (W/o Loc.). Each cell reports relative SR gain on top, with the corresponding 1-agent SR → 2-agent SR shown below.
We examine robustness from two perspectives: the sensitivity of collaboration to imperfect prior information, and the failure modes that arise under collaboration. Additional prior information ablations and mixed-model team experiments are deferred to Appendix D.4 and D.5. 6.1
+6.2 +4.2 -2.1 0.0 0.1 0.2
Success Rate (SR)
1-agent
+6.1 +3.7
+4.4 +16.5 +0.0 -1.0
Qwen3-32B-VL
Figure 11: Comparison in leader-based collaboration with and without visual augmentation.
6
+2.1 -1.0
Model GPT-5-mini
Does Collaboration Remain Effective Under Imperfect Information?
Qwen3-32B
At initialization, we provide all task-relevant object locations in the prior information I, which is hard to obtain in real scenarios. Therefore, we further investigate the collaboration under two realistic settings: (1) Missing location priors: object locations are removed from the prior information, requiring agents to find targets through visual search. (2) Noisy location priors: an additional 30% false object locations are injected into the provided prior information. We extend the effective step budget to 80 steps to allow sufficient search, and apply parallel settings with basic broadcast protocol. As shown in Figure 12, two-agent collaboration improves the success rate and reduces the conflicts, indicating that collaboration remains useful when task information is incomplete. When the number of agents is further increased, a similar inverted-Ushaped trend emerges, as shown in Figure 13(a), while performance remains more stable on complex tasks. Moreover, Table 3 demonstrates that collaboration achieves the largest relative perfor-
Qwen3.5-9B
W/ Noise W/o Noise W/o Loc. +8.22%
+1.16%
+6.47%
76.04→82.29 90.62→91.67 32.29→34.38
+28.83%
+28.57%
+18.88%
54.17→69.79 58.33→75.00 21.88→26.04
+4.01%
+3.52%
+0.00%
78.12→81.25 85.42→88.54 29.17→29.17
mance gains when the information contains noise, suggesting that agents can compensate for misleading information through communication and distributed exploration. These results together validate that multi-agent collaboration is robust to insufficient and noisy information. 6.2
Where Does Collaboration Break Down?
We identify two major failure modes in multi-agent collaboration. These failures suggest that robust multi-agent embodied collaboration requires conflict resolution and belief verification mechanisms. Examples and more statistics in Appendix F. Multiple agents can introduce conflicts, including duplicate grabs, repeated labor, and redundant exploration. Although communication mitigates these conflicts for most models, InternVL3.5-241B still exhibits the highest duplicate-grab rates among 8
2-agent setup of 21.9%, partly explaining its performance drop in Figure 6. This also explains why performance drop with large team size: for Qwen332B-VL, scaling to five agents increases duplicate grabs to 57.4%. Hallucinated beliefs can propagate across agents. When one agent falsely believes the task is complete and broadcasts it, the team may stop prematurely. Although rare (1%-1.5%), such hallucinated completion is highly destructive, causing a 73.7% SR drop in parallel tasks and a 40.9% drop in sequential tasks. Object confusion can similarly cause agents to skip unfinished subtasks.
7
ration (Kang et al., 2025; Puig et al., 2021; Jain et al., 2020; Carroll et al., 2019; Agashe et al., 2025). Many are built on 2D games (Sun et al., 2025; Yang et al., 2026), offering limited support for physical interaction and spatial reasoning. Existing 3D benchmarks often convert visual observations into textual descriptions (Zhang et al., 2024a; Mandi et al., 2024; Chang et al., 2025; Dong et al., 2024; White et al., 2025), decoupling visual grounding and exploration from collaborative decision-making and overlooking behaviors unique to multimodal embodied interaction. Some recent benchmarks support multimodal observations (Kang et al., 2025; Fan et al., 2025a; Zha et al., 2026), but mainly evaluate visual question answering or offline reasoning rather than closedloop interaction. TeamCraft (Long et al., 2024) takes an initial step toward evaluating embodied MLLM collaboration, but does not systematically study how communication mechanisms, coordination modes, and team size shape collaboration.
Related Work
Multi-Agent System Inspired by human society, multi-agent systems (MAS) have been studied as a way to improve task efficiency and effectiveness (Stone and Veloso, 2000). Recent work has explored LLM-based MAS across diverse domains, including automated research (Schmidgall et al., 2025; Su et al., 2025), programming (Hong et al., 2024; Islam et al., 2025), medical decisionmaking (Fan et al., 2025b; Chen et al., 2025), and social simulation (Piao et al., 2025). Other studies further examine their scalability and emergent collaborative behaviors (Qian et al., 2025; Chen et al., 2024). Although recent work has begun to explore embodied multi-agent scenarios (Zhang et al., 2025a; Liu et al., 2024), multimodal embodied collaboration remains underexplored (Yu et al., 2024).
8
Conclusion
We empirically investigate how MLLM-based agents collaborate in embodied environments by constructing MECoBench, a multimodal embodied cooperation benchmark. Through controlled experiments and analyses, our results demonstrate that effective coordination can substantially improve both task performance and efficiency. By systematically examining collaboration modes and communication mechanisms, we reveal how these factors shape collaborative performance. We further evaluate collaboration under noisy and incomplete prior information, demonstrating its robustness in more challenging settings. Beyond aggregate performance, we identify emergent collaborative behaviors and failure patterns, explaining both the benefits of collaboration and the new challenges it introduces. Overall, our study provides a deeper understanding of MLLM-based embodied collaboration and introduces a benchmark for evaluating the collaborative capabilities of future MLLMs.
Embodied Agent Recent studies have explored LLM- or MLLM-based embodied agents for a wide range of tasks, including object manipulation (Yang et al., 2024; Mon-Williams et al., 2025), navigation (Zhang et al., 2024b), complex perceptioninteraction tasks (Szot et al., 2024), and embodied human-AI collaboration (Chang et al., 2025). Several benchmarks have also been proposed to evaluate embodied agents across diverse environments and capabilities (Deitke et al., 2022; Savva et al., 2019; Yang et al., 2025; Cheng et al., 2025). However, these benchmarks predominantly focus on single-agent settings, providing limited support for studying collaboration among multiple embodied agents.
Limitations In this work, we propose MECoBench and conduct systematic experiments to study cooperative behaviors of MLLM-based agents in embodied environments. Our analysis reveals several non-trivial patterns in multi-agent collaboration, but there remain several limitations for future work.
Embodied Multi-Agent Collaboration Benchmarks Several benchmarks have been proposed to evaluate embodied multi-agent collabo9
Expanding tasks and scenarios. MECoBench and corresponding evaluation platform is currently built on VirtualHome, an indoor simulator focused on household scenarios. As a result, our tasks are mainly limited to everyday indoor activities and object interactions. Future work can extend the benchmark to more diverse environments, such as openworld or outdoor scenes, and incorporate broader task types such as navigation and long-horizon exploration.
Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John Turner, Eric Undersander, and Tsung-Yen Yang. 2025. Partnr: A benchmark for planning and reasoning in embodied multi-agent tasks. In International Conference on Learning Representations, volume 2025, pages 65205–65268. Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. 2024. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors. In International Conference on Learning Representations, volume 2024, pages 20094–20136.
Data scale. MECoBench contains 96 cases from 8 task types under two major task setups. This scale is sufficient for controlled evaluation and for revealing key collaboration patterns, but it remains relatively small compared with the diversity and complexity of real-world embodied tasks. Future work can expand the number of task instances, object categories, scene layouts, and collaboration requirements.
Xi Chen, Huahui Yi, Mingke You, WeiZhi Liu, Li Wang, Hairui Li, Xue Zhang, Yingman Guo, Lei Fan, Gang Chen, Qicheng Lao, Weili Fu, Kang Li, and Jian Li. 2025. Enhancing diagnostic capability with multiagents conversational large language models. npj Digital Medicine, 8(1):159.
Agent scaling. Due to spatial constraints in the environment and the complexity of task execution, our experiments scale the team size up to 5 agents. Larger teams and more complex tasks are therefore left unexplored.
Zhili Cheng, Yuge Tu, Ran Li, Shiqi Dai, Jinyi Hu, Shengding Hu, Jiahao Li, Yang Shi, Tianyu Yu, Weize Chen, Lei Shi, and Maosong Sun. 2025. Embodiedeval: Evaluate multimodal llms as embodied agents. Preprint, arXiv:2501.11858. Matt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs, Kiana Ehsani, Jordi Salvador, Winson Han, Eric Kolve, Aniruddha Kembhavi, and Roozbeh Mottaghi. 2022. Procthor: Large-scale embodied ai using procedural generation. In Advances in Neural Information Processing Systems, volume 35, pages 5982–5994. Curran Associates, Inc.
References Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang. 2025. LLM-coordination: Evaluating and analyzing multi-agent coordination abilities in large language models. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 8053–8072, Albuquerque, New Mexico. Association for Computational Linguistics.
Yubo Dong, Xukun Zhu, Zhengzhe Pan, Linchao Zhu, and Yi Yang. 2024. VillagerAgent: A graph-based multi-agent framework for coordinating complex task dependencies in Minecraft. In Findings of the Association for Computational Linguistics: ACL 2024, pages 16290–16314, Bangkok, Thailand. Association for Computational Linguistics.
Anthropic. 2026. Claude Opus 4.7 System Card. https://www.anthropic.com/system-cards. System card PDF available at https: //cdn.sanity.io/files/4zrzovbb/website/ 037f06850df7fbe871e206dad004c3db5fd50340. pdf. Accessed May 23, 2026.
Xianzhe Fan, Xuhui Zhou, Chuanyang Jin, Kolby Nottingham, Hao Zhu, and Maarten Sap. 2025a. Somitom: Evaluating multi-perspective theory of mind in embodied social interactions. In Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc.
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025. Qwen3-vl technical report. Preprint, arXiv:2511.21631.
Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. 2025b. AI hospital: Benchmarking large language models in a multi-agent medical interaction simulator. In Proceedings of the 31st International Conference on Computational Linguistics, pages 10183–10213, Abu Dhabi, UAE. Association for Computational Linguistics.
Micah Carroll, Rohin Shah, Mark Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan. 2019. On the utility of learning about humans for human-ai coordination. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
10
Zhaohan Feng, Ruiqi Xue, Lei Yuan, Yang Yu, Ning Ding, Meiqin Liu, Bingzhao Gao, Jian Sun, Xinhu Zheng, and Gang Wang. 2026. Multi-agent embodied ai: advances and future directions. Science China Information Sciences, 69(5):151202.
Qian Long, Zhi Li, Ran Gong, Ying Nian Wu, Demetri Terzopoulos, and Xiaofeng Gao. 2024. Teamcraft: A benchmark for multi-modal multi-agent systems in minecraft. Preprint, arXiv:2412.05255. Zhao Mandi, Shreeya Jain, and Shuran Song. 2024. Roco: Dialectic multi-robot collaboration with large language models. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 286–299.
Google DeepMind. 2026a. Gemini 3.1 Pro Model Card. https://deepmind.google/models/ model-cards/gemini-3-1-pro/. Published February 2026. Accessed May 23, 2026. Google DeepMind. 2026b. Gemma 4 model card. https://ai.google.dev/gemma/docs/ core/model_card_4. Accessed: 2026-05-26.
Meta AI. 2025. Llama 4: Model cards and prompt formats. https://www.llama.com/docs/ model-cards-and-prompt-formats/llama4/. Accessed: 2026-05-26.
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, zili wang, Steven Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. Metagpt: Meta programming for a multi-agent collaborative framework. In International Conference on Learning Representations, volume 2024, pages 23247–23275.
Ruaridh Mon-Williams, Gen Li, Ran Long, Wenqian Du, and Christopher G. Lucas. 2025. Embodied large language models enable robots to complete complex tasks in unpredictable environments. Nature Machine Intelligence, 7(4):592–601. Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo. 2023. Embodiedgpt: Visionlanguage pre-training via embodied chain of thought. In Advances in Neural Information Processing Systems, volume 36, pages 25081–25094. Curran Associates, Inc.
Md. Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez. 2025. CodeSim: Multiagent code generation and problem solving through simulation-driven planning and debugging. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 5128–5154, Albuquerque, New Mexico. Association for Computational Linguistics.
OpenAI. 2025. Gpt-5 mini model. https: //developers.openai.com/api/docs/models/ gpt-5-mini. Accessed: 2026-05-26.
Unnat Jain, Luca Weihs, Eric Kolve, Ali Farhadi, Svetlana Lazebnik, Aniruddha Kembhavi, and Alexander Schwing. 2020. A cordial sync: Going beyond marginal policies for multi-agent embodied tasks. In Computer Vision – ECCV 2020, pages 471–490, Cham. Springer International Publishing.
OpenAI. 2026. Gpt-5.4 thinking system card. https://openai.com/index/ gpt-5-4-thinking-system-card/. Accessed: 2026-05-26.
Li Kang, Xiufeng Song, Heng Zhou, Yiran Qin, Jie Yang, Xiaohong Liu, Philip Torr, LEI BAI, and Zhenfei Yin. 2025. Viki-r: Coordinating embodied multiagent cooperation via reinforcement learning. In Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc.
Jinghua Piao, Yuwei Yan, Jun Zhang, Nian Li, Junbo Yan, Xiaochong Lan, Zhihong Lu, Zhiheng Zheng, Jing Yi Wang, Di Zhou, Chen Gao, Fengli Xu, Fang Zhang, Ke Rong, Jun Su, and Yong Li. 2025. Agentsociety: Large-scale simulation of llm-driven generative agents advances understanding of human behaviors and society. Available at SSRN.
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles.
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. 2018. Virtualhome: Simulating household activities via programs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Song XiXuan, Yifan Xu, Shudan Zhang, Hanyu Lai, Jiadai Sun, Xinyue Yang, Yu Yang, Zehan Qi, Shuntian Yao, Xueqiao Sun, Siyi Cheng, Qinkai Zheng, Hao Yu, Hanchen Zhang, Wenyi Hong, and 9 others. 2025. Visualagentbench: Towards large multimodal models as visual foundation agents. In The Thirteenth International Conference on Learning Representations.
Xavier Puig, Tianmin Shu, Shuang Li, Zilin Wang, Yuan-Hong Liao, Joshua B. Tenenbaum, Sanja Fidler, and Antonio Torralba. 2021. Watch-and-help: A challenge for social perception and human-{ai} collaboration. In International Conference on Learning Representations. Chen Qian, Zihao Xie, YiFei Wang, Wei Liu, Kunlun Zhu, Hanchen Xia, Yufan Dang, Zhuoyun Du, Weize Chen, Cheng Yang, Zhiyuan Liu, and Maosong Sun. 2025. Scaling large language model-based multiagent collaboration. In International Conference
Xinzhu Liu, Di Guo, Xinyu Zhang, and Huaping Liu. 2024. Heterogeneous embodied multi-agent collaboration. IEEE Robotics and Automation Letters, 9(6):5377–5384.
11
on Learning Representations, volume 2025, pages 41488–41505.
Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Rin Metcalf, Walter Talbott, Natalie Mackraz, R Devon Hjelm, and Alexander T Toshev. 2024. Large language models as generalizable policies for embodied tasks. In The Twelfth International Conference on Learning Representations.
Qwen Team. 2026. Qwen3.5: Towards native multimodal agents. Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. 2019. Habitat: A platform for embodied ai research. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, and 69 others. 2025. Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning. Preprint, arXiv:2507.01006.
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. 2025. Agent laboratory: Using LLM agents as research assistants. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5977– 6043, Suzhou, China. Association for Computational Linguistics.
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, and 56 others. 2025. Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. Preprint, arXiv:2508.18265.
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Isadora White, Kolby Nottingham, Ayush Maniar, Max Robinson, Hansen Lillemark, Mehul Maheshwari, Lianhui Qin, and Prithviraj Ammanabrolu. 2025. Collaborating action by action: A multi-agent llm framework for embodied reasoning. Preprint, arXiv:2504.17950.
Peter Stone and Manuela Veloso. 2000. Multiagent systems: A survey from a machine learning perspective. Autonomous Robots, 8(3):345–383.
Hanqing Yang, Narjes Nourzad, Shiyu Chen, Marie Siew, Jingdi Chen, and Carlee Joe-Wong. 2026. Coop2 : Defining, observing, and repairing cooperation in llm multi-agent systems. Preprint, arXiv:2603.00349.
Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nanqing Dong. 2025. Many heads are better than one: Improved scientific idea generation by a LLMbased multi-agent system. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 28201–28240, Vienna, Austria. Association for Computational Linguistics.
Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, Heng Ji, Huan Zhang, and Tong Zhang. 2025. Embodiedbench: Comprehensive benchmarking multi-modal large language models for visiondriven embodied agents. In Forty-second International Conference on Machine Learning.
Haochen Sun, Shuwen Zhang, Lujie Niu, Lei Ren, Hao Xu, Hao Fu, Fangkun Zhao, Caixia Yuan, and Xiaojie Wang. 2025. Collab-overcooked: Benchmarking and evaluating large language models as collaborative agents. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 4922–4951, Suzhou, China. Association for Computational Linguistics.
Yijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao, Lusong Li, Li Shen, Xiaodong He, Jing Jiang, and Yuhui Shi. 2024. Embodied multi-modal agent trained by an llm from a parallel textworld. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26275–26285.
Katia P Sycara. 1998. Multiagent systems. AI magazine, 19(2):79–79.
Xianhao Yu, Jiaqi Fu, Renjia Deng, and Wenjuan Han. 2024. Mineland: Simulating large-scale multi-agent interactions with limited multimodal senses and physical needs. Preprint, arXiv:2403.19267.
Andrew Szot, Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira, and Alexander Toshev. 2025. From multimodal llms to generalist embodied agents: Methods and lessons. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10644–10655.
Jirong Zha, Yuxuan Fan, Tianyu Zhang, Geng Chen, Yingfeng Chen, Chen Gao, and Xinlei Chen. 2026. Aircopbench: A benchmark for multi-drone collaborative embodied perception and reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 1507–1515.
12
Hongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Joshua B Tenenbaum, Tianmin Shu, and Chuang Gan. 2024a. Building cooperative embodied agents modularly with large language models. In International Conference on Learning Representations, volume 2024, pages 19373–19401. Hongxin Zhang, Zeyuan Wang, Qiushi Lyu, Zheyuan Zhang, Sunli Chen, Tianmin Shu, Behzad Dariush, Kwonjoon Lee, Yilun Du, and Chuang Gan. 2025a. Combo: Compositional world models for embodied multi-agent cooperation. In International Conference on Learning Representations, volume 2025, pages 49996–50019. Jiazhao Zhang, Kunyu Wang, Rongtao Xu, Gengze Zhou, Yicong Hong, Xiaomeng Fang, Qi Wu, Zhizheng Zhang, and He Wang. 2024b. Navid: Video-based vlm plans the next step for vision-andlanguage navigation. Robotics: Science and Systems. Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang, and Xipeng Qiu. 2025b. Vlabench: A large-scale benchmark for languageconditioned robotics manipulation with long-horizon reasoning tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11142–11152. Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, and Jiaxuan You. 2025. MultiAgentBench : Evaluating the collaboration and competition of LLM agents. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8580–8622, Vienna, Austria. Association for Computational Linguistics.
13
Table 4: Task templates used in our benchmark. Each task is defined by a set of target predicates, where ON(obj, rec) and IN(obj, rec) denote placing a surface on or inside a container, respectively.
A
Task Name
Predicate Set
Self-caring
ON(toothbrush, bathroomcounter), ON(toothpaste, bathroomcounter), ON(towel, bathroomcounter), ON(barsoap, bathroomcounter)
Work on desk
ON(keyboard, desk), ON(mouse, desk), ON(cellphone, desk), ON(mug, desk)
Gaming setup
ON(boardgame, coffeetable), ON(remotecontrol, coffeetable), ON(magazine, coffeetable), ON(toy, coffeetable), ON(book, coffeetable)
Prepare afternoon tea
ON(cupcake, coffeetable), ON(juice, coffeetable), ON(wine, coffeetable), ON(pudding, coffeetable), ON(apple, coffeetable)
Wash dishes
IN(plate, dishwasher), IN(waterglass, dishwasher), IN(wineglass, dishwasher), IN(cutleryfork, dishwasher)
Prepare food
ON(cupcake, kitchentable), ON(juice, kitchentable), ON(pancake, kitchentable), ON(poundcake, kitchentable), ON(wine, kitchentable), ON(pudding, kitchentable), ON(apple, kitchentable), ON(coffeepot, kitchentable)
Put food inside fridge
IN(cupcake, fridge), IN(juice, fridge), IN(pancake, fridge), IN(poundcake, fridge), IN(wine, fridge), IN(pudding, fridge), IN(apple, fridge)
Setup dinner table
ON(plate, kitchentable), ON(waterglass, kitchentable), ON(wineglass, kitchentable), ON(cutleryfork, kitchentable)
Find and put 1 magazine on coffeetable. Possible locations: One magazine is on the cabinet, the cabinet is inside the bedroom. Appearance: Glossy comic-style magazine with orange-and-blue cover, vertical rectangular shape, bold white title text, and a blue robotic hero illustration. coffeetable appearance: Square low coffee table with light gray wood top, dark charcoal recessed base, and a centered four-panel top divided by thin seams.;Find and put 1 toy on coffeetable. Possible locations: One toy is on the bathroomcounter, the bathroomcounter is inside the bathroom. / One toy is on the sofa, the sofa is inside the kitchen. Appearance: Colorful wooden toy truck with blue box body, green cab, yellow base, and red wheels.;Find and put 1 book on coffeetable. Possible locations: One book is on the kitchentable, the kitchentable is inside the kitchen. Appearance: Purple hardcover book with white front cover, Cisco Unity Connection title, gold page edges, rectangular shape.;Find and put 2 remotecontrols on coffeetable. Possible locations: One remotecontrol is on the desk, the desk is inside the bedroom. / One remotecontrol is on the desk, the desk is inside the livingroom. / One remotecontrol is inside the fridge, the fridge is inside the kitchen. Appearance: Gray plastic TV remote, rounded rectangular shape, slightly tapered ends, white button face with black and red buttons.;Find and put 1 boardgame on coffeetable. Possible locations: One boardgame is on the sofa, the sofa is inside the livingroom. / One boardgame is inside the fridge, the fridge is inside the kitchen. Appearance: Square yellow board with dark gray play area, raised center platform, and small red, yellow, green, and blue tokens.;coffeetable is in livingroom.
MECoBench
This section presents detailed information about MECoBench, covering its task definitions, benchmark statistics, and construction details. A.1
Task definition
Table 4 shows the detailed templates of the eight task types designed for MECoBench. Self-caring, work-on-desk, and gaming setup are three newly introduced types, while prepare afternoon tea, wash dishes, prepare food, put food inside the fridge, and setup dinner table are modified from WAH. Figure 14 gives an example of final task goal of gaming setup. A.2
Benchmark construction details
The construction of MECoBench tasks follows a two-step pipeline. In step one, task templates are instantiated into concrete scenes with randomized object placement and validated for feasibility.
Figure 14: Example of task goal content. The goal specifies target objects, their possible locations highlighted in blue, and appearance descriptions highlighted in orange.
Goal Sampling. For a given task type, we sample subgoals from predefined object-count ranges, specifying the target objects, required counts, and optional room or surface constraints. To mitigate visibility issues such as occlusion, we add 0–2 redundant instances for each target object category
during scene instantiation. For large objects that occupy more surface space, such as keyboards, boardgames, and magazines, the redundancy is lim14
ited to 0–1 instance. Scene Instantiation and Object Placement. Once a template is sampled, the target objects are assigned to legal initial locations based on the predefined legal placement (Table 5). Objects are placed randomly while respecting constraints: • No object is placed on the same surfaces or containers that are reserved for other goal locations. • Surface space is checked to ensure sufficient area for the object. • Containers are verified for available slots using their bounding boxes.
Figure 15: Three-layer sunburst chart illustrating the composition of eight household tasks used in our evaluation. The inner ring represents the eight task types; the middle ring shows the objects involved in each task; the outer ring indicates the object category.
This procedure ensures that the initial scene is diverse yet consistent with the task template. Table 5: Legal placement locations used for randomized object initialization in MECoBench. Type Container Surface
Legal Locations fridge, stove, dishwasher, microwave, cabinet, bookshelf sofa, cabinet, bathroomcounter, nightstand, bench, bed, desk, tvstand, kitchencounter, coffeetable, stove, kitchentable
Validation and Goal Specification. After placement, the VirtualHome environment is used to verify scene feasibility:
Figure 16: Data distribution grouped by task type and number of subgoals.
• The environment graph is updated to reflect the actual object positions.
Figure 16 shows the data distribution of MECoBench across different task types and task complexity levels, where task complexity is measured by the number of involved objects.
• Target object counts are checked to ensure that all task objects exist in the scene. • Predicates from the template are instantiated with the real node IDs in the scene.
B
Evaluation
B.1
Evaluation Setup
• Tasks failing these checks (e.g., insufficient instances or invalid placements) are discarded, and a new attempt is generated.
Each observation view is rendered at 256 × 256 pixels, resulting in a four-view concatenated observation of 256 × 1024. The four views are arranged in the order: front, back, left, and right. Both horizontal and vertical fields of view are set to 90◦ . The agent model height is 1.8 m, and the camera is mounted at the same height with a forward offset of 0.15 m relative to the agent’s position. The camera is pitched downward by 30◦ to capture more of the nearby surfaces and objects. Figure 17 shows an example of the resulting concatenated observation.
A.3
Benchmark Statistic
MECoBench involves 25 common everyday objects, which are grouped into six categories: food and drinks, tableware, electronics, hygiene items, entertainment items, and office items. Figure 15 illustrates the correspondence between the eight task types and their associated object categories. 15
Table 6: Full names of MLLMs used in our evaluation. Model Name
Creator
Full Name / Identifier
GPT-5-mini GPT-5.4 Gemini-3.1-Pro
OpenAI OpenAI Google
gpt-5.4-mini gpt-5.4 gemini-3.1-pro
Qwen3-8B-VL Qwen3-32B-VL Qwen3-235B-VL Qwen3.5-9B Qwen3.5-27B
Qwen Qwen Qwen Qwen Qwen
Qwen/Qwen3-VL-8B-Instruct Qwen/Qwen3-VL-32B-Instruct Qwen/Qwen3-VL-235B-A22B-Instruct Qwen/Qwen3.5-9B Qwen/Qwen3.5-27B
Gemma4-26B Gemma4-31B
Google Google
google/gemma-4-26b-it google/gemma-4-31b-it
InternVL-3.5-38B InternVL-3.5-241B
OpenGVLab OpenGVLab
OpenGVLab/InternVL3_5-38B OpenGVLab/InternVL3_5-241B-A28B
Llama-4 Scout
Meta
meta-llama/Llama-4-Scout-17B-16E-Instruct
GLM-4.6V GLM-4.6V-Flash
Z.ai Z.ai
zai-org/GLM-4.6V zai-org/GLM-4.6V-Flash
most two objects. When both hands are occupied, the agent cannot execute grab or open. For put_* and handover, the object to be placed or handed over must already be held by the acting agent.
Figure 17: Example observation image.
• Visibility: The target objects of walk, grab, and put_* actions must be visible in the agent’s current observation.
For all evaluated models, we adopt the recommended generation configurations whenever available. For closed-source models, we use the official APIs with their default generation settings. For open-source models, we follow the generation configurations specified in the corresponding model cards or vLLM (Kwon et al., 2023) documentation to ensure strong and reliable performance. B.2
• Container state: An agent cannot execute open on a container that is already open. • Room accessibility: For walk_*, grab, and put_* actions, the target must be located in one of the rooms accessible to the agent.
Model Versions
Table 6 summarizes the model versions and full identifiers used in our experiments. Proprietary models are accessed through their official APIs, while open-source models are locally deployed using vLLM (Kwon et al., 2023). B.3
• Invalid actions: Any action name not listed in Table 7 is considered invalid. Multi-Agent Interaction Constraints When multiple agents act simultaneously, additional interaction constraints are applied:
Action Space
Table 7 lists the action space for agents when evaluation for both parallel and sequential tasks. Parallel tasks include basic navigation and object manipulation actions, while sequential tasks introduce cross-room handover actions.
• Conflict resolution: The conflict set includes grab and open. If multiple agents attempt the same conflict action on the same object at the same time, one agent is randomly selected as the winner, while the others remain idle.
Single-Agent Action Legality Each agent action must satisfy the following legality constraints:
• Handover distance: A handover action is valid only when the giver and receiver are within 2 meters of each other.
• Hand occupancy: Each agent can hold at 16
Table 7: Action space for agents in MECoBench. Parallel tasks include basic navigation and object manipulation actions. Sequential tasks introduce cross-room handover actions. Action
Description
Parameters Parallel Task Actions
walk walk_to_room grab open put_on put_in
Move towards a visible object, furniture, or door Move directly to another room Pick up a visible and grabbable object Open a container Place the held object on a visible surface Place the held object inside a visible container
object bedroom / livingroom / kitchen / bathroom object container object (holding), surface object (holding), container
Sequential Task Additional Actions walk_to_door handover receive
Move to a specific door in the environment topology Give a held object to another agent when close (<2m) Wait to receive an object
C
{
"action": "grab", "parameter_1": {"name": "juice", "description": "Square beige plastic juice bottle with purple cap and label, Naked branding, fruit graphics, rectangular sides."}, "parameter_2": null, "room": null }
door_id object(holding), recipient agent name none
Core Prompt Templates
This section presents the core prompt templates designed for agents in the MECoBench evaluation platform. C.1
{"id": "376", "class_name": "juice", "bbox": […]}, {"id" : "378", "class_name": "wine", … }, … VLM Decoder
Communicate Phase
Below we list prompts for broadcast, discuss and leader-based communication protocol (Figure 19, 20, 21, 22, 23).
[grab] <juice> (376)
Figure 18: Example of action decoding and grounding.
You are an embodied agent in a 3D simulated environment. At this stage, based on the current observation, memory, history, and task goal, you must decide whether to send ONE helpful message to other agents.
B.4
## Communication Send a message if it is helpful for task progress and remain silent if the message would be redundant, speculative, or misleading. Your goal is to complete the task with the **smallest number of action steps**. Therefore, avoid waiting or letting others wait. We should all actively engage and search for tasks to contribute to. The valid actions we can take are: $VALID_ACTIONS$. Each agent can only perform one action at a time, and can hold at most two objects at a time. \n ### MESSAGE Output Format - If silent, output exactly: <NO_MESSAGE> - Otherwise, output the message text ONLY.
Action Decoding and Grounding
Agents output semantic actions based on their observation rather than executable simulator commands. Thus, each generated action is decoded and grounded before execution. Actions that do not require object grounding, such as receive, walk_to_room/door, and handover, are resolved by rule-based mappings. For object-centric actions, including walk, grab, open, put_on, and put_in, semantic object descriptions are grounded to concrete simulator object identifiers.
## Output Format [THINKING] Brief reasoning for the message decision. [MESSAGE] <NO_MESSAGE> or message text
As shown in Fig. 18, the top panels present the agent’s first-person observation and generated semantic action. The decoder then uses visible object candidates, their bounding boxes, and metadata, together with both the original and bbox-annotated views, to select the target simulator ID and produce the final executable action.
## Inputs $INPUT_BLOCK$
Figure 19: Broadcast protocol communication prompt. This module asks each agent to decide whether to send a concise, useful message before physical action planning.
17
You are an embodied agent in a 3D simulated environment. This is the FIRST discussion round for the current step. Based on the current observation, memory, dialogue, and task goal, decide whether more discussion is needed before acting.
You are an embodied agent in a 3D simulated environment. At this stage, you are acting as the leader for this step. Based on your current observation, memory, dialogue, and the workers' reports, assign exactly ONE next-step instruction to each agent, including yourself. You have access to a combined image that shows every agent's current observation views (your own view on top, followed by each worker's view, each row labeled with the agent name). Together with your memory, dialogue, and the workers' text reports, use this shared visual information to assign exactly ONE next-step instruction to each agent, including yourself.
You are an embodied agent in a 3D simulated environment. This is a FOLLOW-UP text-only discussion round for the current step. Use the dialogue so far and your private local context from the first round to decide whether more discussion is needed before acting. ## Discussion You may send ONE helpful discussion message to the other agents. Your goal is to complete the task with the **smallest number of action steps**. Therefore, avoid waiting idly or letting others wait. We should all actively engage and search for tasks to contribute to. The valid actions we can take are: $VALID_ACTIONS$. Each agent can only perform one action at a time, and can hold at most two objects at a time.\n READY rules: - Output <READY> if, from your perspective, no more discussion is needed before agents take actions. - Output <CONTINUE> if more discussion is still needed.\n MESSAGE rules: - If you have nothing useful to add, output exactly <NO_MESSAGE>. - Otherwise output one concise, task-relevant message. - Do not fabricate visual facts.
## Assignment You must output exactly one assignment for each agent in this list: $AGENT_NAMES$ Your goal is to complete the task with the **smallest number of action steps**. Therefore, avoid letting any agent stay idle or wait for other agents.\n Rules: - Use the exact agent names from the list. - Set "type" to "assignment" for every item. - Give concise, actionable instructions for THIS step only. - The valid actions we can take are: $VALID_ACTIONS$. Each agent can only perform one action at a time, and can hold at most two objects at a time. Get full use of all agents' actions and hands to minimize the total number of action steps. - Do not assign an agent to do multiple unrelated things in one message. - If the worker reports failure, you should assign a different action. Do not let the worker retry the failed action if the environment has not changed.
## Output Format [THINKING] Brief reasoning for the discussion decision. [READY] <READY> or <CONTINUE> [MESSAGE] <NO_MESSAGE> or message text [LOCAL_CONTEXT] A short private visual grounding note for your own later text-only discussion rounds.
## Output Format [THINKING] Brief reasoning for the assignments. [MESSAGES] A JSON array like: [ {"receiver": "<exact agent name>", "type": "assignment", "content": "..."} ]
## Inputs $INPUT_BLOCK$
Figure 20: Discussion protocol prompt. The first round prompt is highlighted in blue, and the follow-up rounds prompt is highlighted in orange.
Figure 22: Centralized mode: leader assignment prompt. Text-only communication is highlighted in blue, and the vision-augmented is highlighted in orange.
You are an embodied agent in a 3D simulated environment. At this stage, you must send ONE report message to the leader agent: $LEADER_NAME$ and wait for the leader's assignment. Base the report on your current observation, memory, dialogue, and task goal.
**Current memory:**\n $MEMORY$ **Historical dialogue:**\n $HISTORICAL_DIALOGUE$ **Current round dialogue:**\n $CURRENT_ROUND_DIALOGUE$
## Report Report only task-relevant observations, status, useful feedback and progress to the leader (Agent 0). If any action fails, you should report the failure and possible reason to the leader.\n MESSAGE rules: - If you have nothing useful to report, output exactly <NO_MESSAGE>. - Otherwise output one concise report message to the leader. - Do not invent visual facts.
**Successful actions history:**\n $ACTION_HISTORY$ **Last action feedback:**\n $FEEDBACK$ **The last action (executed vs intended):**\n $LAST_ACTION$ **Current room:**\n You are now standing in the $CURRENT_ROOM$. Most of your current observation images are in this room.
## Output Format [THINKING] Brief reasoning for the report. [MESSAGE] <NO_MESSAGE> or message text
**Holding objects:**\n $HOLDING_OBJECTS$ **Task goal:**\n $TASK_GOAL$ **Current completion progress:** \n $TASK_PROGRESS$
Figure 21: Centralized mode: worker report prompt.
Figure 23: Input block for communicate phase.
18
C.2
## Physical Actions ### Valid Actions - put one object you are holding on a surface or into a container - open an enclosure - grab an object - walk to something visible - walk to another room (allowed rooms:$ALLOWED_ROOMS$)\n ### Interaction Rules - When an action involves any object, specify the name and a clear visual/location description. - You may interact (walk to, grab, put, open) with objects ONLY when they are currently visible. - If the object is not visible, first walk to the room includes it. If still occluded, walk to a nearby visible landmark (e.g., table/counter) until the target becomes visible. - You can hold at most TWO objects at the same time. If both hands are full, put one object down before grabbing another object or opening an enclosure. - Do NOT repeat a failed action unless the environment has changed. Do NOT repeating the same action consecutively. - You may walk directly to any room using the room name except the current one.\n ## ACTION Output Format Output a single JSON object, with no extra text: { "action": "<one of: walk / walk_to_room / grab / put_in / put_on / open >", "parameter_1": {"name": "<class name>", "description": "<short description>"} | null, "parameter_2": {"name": "<class name>", "description": "<short description>"} | null, "room": "<bedroom|livingroom|kitchen|bathroom>" | null }\n Parameter rules: - walk/open/grab: use "parameter_1" only. - put_in/put_on: "parameter_1" = held object, "parameter_2" = container/surface. - walk_to_room: set "room" and set both parameters to null. - Actions with no parameters: set "parameter_1", "parameter_2", and "room" to null.
Act Phase
This section presents the prompts used for action decision-making, memory update and action decoding (Figure 24, 25, 26, 27, 28, 29). You are an embodied agent in a 3D simulated environment. At this stage, based on the current observation, memory, current communication context, and task goal, you must produce TWO outputs in order: 1) Decide the NEXT single physical action. 2) Update the memory. You coordinate with other agents ONLY through a shared memory object. There is no current-round dialogue in this mode. At this stage, based on the current observation, shared memory, previous action feedback, and task goal, you must produce TWO outputs in order: 1) Decide the NEXT single physical action. 2) Update the shared memory with observed facts, your current plan, and your own status. ## Physical Actions $ACTION_RULES$ ## Memory Update $MEMORY_RULES$ ## Output Format [THINKING] Brief reasoning for the next action. [ACTION] The JSON action object. Only one action block is allowed. [THINKING] Brief reasoning for the memory update. [MEMORY] The updated memory JSON.
Figure 24: Action and memory update prompt. Normal memory is highlighted in blue and shared memory is highlighted in orange.
Figure 26: Action rules block.
You are resolving object_ids for an action based on the class name and descriptions. You are given: - An action with object names and descriptions. - A list of visible candidate objects with ids and class names. - A list of holding objects with ids and class names. - Two images of the same views: 1) Raw RGB image. 2) Bbox-annotated image with class names and object_ids. Rules: 1. Only choose ids from candidates or holding_objects. 2. Match by class name and description. ... 3. For put_in / put_on, parameter_ids[0] MUST be the held object, parameter_ids[1] is the target container/surface. 4. If multiple candidates match, choose the best and most reasonable fit ... 5. If the action is walk_to_room, return the room name in parameter_ids. 6. If unsure, choose the closest matching candidate rather than returning null. Output JSON only (no extra text): { "action": "<action>", "parameter_ids": [<int>, <int>] OR [<int>] OR ["<room>"] OR [] } Input JSON: $RESOLVE_INPUT$
**Current memory:**\n $MEMORY$ **Current round dialogue:**\n $CURRENT_ROUND_DIALOGUE$ **Successful actions history:**\n $ACTION_HISTORY$ **Last action feedback:**\n $FEEDBACK$ **The last action (executed vs intended):**\n $LAST_ACTION$ **Allowed rooms for you:**\n $ALLOWED_ROOMS$ **Department topology:**\n $DEPARTMENT_TOPOLOGY$ **Current room:**\n You are now standing in the $CURRENT_ROOM$. Most of your current observation images are in this room. **Holding objects:**\n $HOLDING_OBJECTS$ **Task goal:**\n $TASK_GOAL$ **Current completion progress:** \n $TASK_PROGRESS$
Figure 27: Input block for the act phase. For sequential tasks, the two additional parts are highlighted in blue.
Figure 25: Prompt for action decoder.
19
## Shared Memory Update Every agent reads and writes to this SAME memory object. You can use this memory to coordinate with other agents.\n ### Rules: - Store only task-relevant objects, containers, and target locations. - Update memory using the current observation and previous action feedback. Do NOT record the action you are about to output as completed yet. - If new information conflicts with existing memory, correct the memory to reflect the latest belief. - Keep memory concise and consistent. Memory represents the current belief state, not a full history. Do not invent objects. - Keep agent_locations, agent_plans, agent_status, agent_holding, task_progress, and claims updated so other agents can coordinate. - Read shared memory before acting to avoid redundant exploration and action conflicts. - Prefer unassigned subtasks instead of repeating work another agent is already doing. If all visible subtasks are claimed, deliver any object you hold to its target, explore a different unchecked place, or wait instead of duplicating the same target. - Track multi-count goals by required, satisfied, and remaining counts. Mark a subtask satisfied only when the environment feedback or satisfied task list confirms the object is at the target. - Use claims to reserve one concrete next subtask per agent, including the object name, source location, target location, and state: "planned", "holding", "delivering", or "done". - Do not update or guess another agent's status, holdings, or completed actions.\n ### MEMORY Output format { "livingroom": { ... }, "kitchen": { ... }, "bedroom": { ... }, "bathroom": { ... }, "global": { "notes": [], "agent_locations": {"<agent_name>": "", ...}, "agent_plans": {"<agent_name>": "", ...}, "agent_status": {"<agent_name>": "", ...}, "agent_holding": {"<agent_name>": [{"name": "...", "source": "...", "target": "..."}], ...}, "task_progress": {"<goal_key>": {"required": 0, "satisfied": 0, "remaining": 0}}, "claims": [{"agent": "<agent_name>", "object": "...", "source": "...", "target": "...", "state": "planned|holding|delivering|done"}], "coordination_board": ["", ...] } }
## Memory Update ### Rules: - Store only task-relevant objects, containers, and target locations. - Update memory using the current observation, executed actions, feedback, and reliable messages from other agents. - If new information conflicts with existing memory, correct the memory to reflect the latest belief. - Keep memory concise and consistent. Memory represents the current belief state, not a full history. Do not invent objects. - Use global.communication_summary to keep a concise running summary of still-useful communication facts, assignments, and commitments. Do not copy raw dialogue into it.\n ### MEMORY Output format { "livingroom": { "exploration": "unexplored|partially_explored|explored", "objects": [{"name": "...", "description": "...", "location": "...", "status": "..."}], "containers": [{"name": "...", "description": "...", "location": "...", "status": "unchecked|checked|open|closed"}], "notes": [] }, "kitchen": { ... }, "bedroom": { ... }, "bathroom": { ... }, "global": { "notes": [], "agent_locations": {"<agent_name>": "", ...}, "communication_summary": "” } }
Figure 28: Memory update rules block.
D
Experiment Results
In this section, we provide detailed results for the experiments in Sections 4, 5, and 6, as well as some additional experimental results. D.1
Performance of Different Models in Parallel and Sequential Tasks
Table 8 shows the detailed result of figure 5 and 6. Overall, GPT-5.4 achieves the best performance. GPT-5-mini and Gemma4-31B are the secondtier models, showing consistently strong results across the two collaboration structures. In addition, Qwen3.5, Gemma4-26B, and Gemini-3.1-Pro also exhibit relatively strong performance. Qwen3VL-32B and Qwen3-VL-235B fall into the middleperformance group, with relatively balanced results across parallel and sequential tasks. Qwen3-VL8B is a relatively weak model, but its performance remains comparatively balanced between the two coordination structures. In contrast, InternVL-3.5 and GLM-4.6V struggle particularly on sequential tasks, where precise collaboration is required. Finally, Llama-4-Scout and GLM-4.6V-Flash perform poorly under both collaboration structures.
Figure 29: Shared memory update rules block.
ter within the same family. However, scaling tends to amplify existing task-specific patterns rather than change them. Some families remain balanced across parallel and sequential tasks, while others are consistently stronger in single-agent or parallel settings, or particularly weak on sequential tasks requiring precise coordination. For example, Gemma4 benefits from scaling especially on sequential tasks, whereas InternVL-3.5 achieves stronger single-agent parallel performance at larger scale but degrades in sequential collaboration. This suggests that stronger individual capability does not necessarily translate into better communication and coordination. Based on model performance and cost-
Model scaling effects. Overall, except for InternVL-3.5, larger models generally perform bet20
Table 8: Performance comparison under parallel and sequential normal settings. Avg. SR and Avg. CR denote the average SR and CR across parallel and sequential tasks under the 2-agent setup. Best results are shown in bold, and second-best results are underlined, respectively for 2-agent and 1-agent setup. Model
#Agents
Parallel Task SR
CR
Sequential Task
AUC
SR
Avg.
CR
AUC
SR
CR
– 0.901 – 0.914 – 0.824
– 0.624 – 0.655 – 0.472
– 0.818 – 0.834 – 0.672
– 0.942 – 0.942 – 0.881
– 0.332 – 0.620 – 0.675 – 0.844 – 0.872 – 0.839 – 0.904 – 0.429 – 0.267 – 0.225 – 0.283 – 0.090
– 0.208 – 0.393 – 0.451 – 0.558 – 0.587 – 0.556 – 0.620 – 0.280 – 0.180 – 0.163 – 0.188 – 0.060
– 0.266 – 0.584 – 0.511 – 0.777 – 0.818 – 0.750 – 0.828 – 0.365 – 0.375 – 0.146 – 0.365 – 0.063
– 0.551 – 0.777 – 0.786 – 0.915 – 0.921 – 0.892 – 0.925 – 0.630 – 0.591 – 0.384 – 0.589 – 0.273
Closed-source MLLMs GPT-5-mini GPT-5.4 Gemini-3.1-Pro
1 2 1 2 1 2
0.906 0.917 0.854 0.896 0.802 0.771
0.974 0.982 0.960 0.969 0.927 0.937
0.793 0.776 0.788 0.731 0.645 0.606
– 0.719 – 0.771 – 0.573
Open-source MLLMs Qwen3-8B-VL Qwen3-32B-VL Qwen3-235B-VL Qwen3.5-9B Qwen3.5-27B Gemma4-26B Gemma4-31B InternVL-3.5-38B InternVL-3.5-241B Llama-4 GLM-4.6V GLM-4.6V-Flash
1 2 1 2 1 2 1 2 1 2 1 2 1 2 1 2 1 2 1 2 1 2 1 2
0.354 0.458 0.583 0.802 0.656 0.667 0.854 0.938 0.844 0.938 0.760 0.812 0.833 0.833 0.677 0.594 0.833 0.719 0.250 0.250 0.594 0.719 0.094 0.125
0.620 0.770 0.753 0.934 0.855 0.896 0.924 0.985 0.950 0.969 0.867 0.944 0.921 0.945 0.875 0.830 0.946 0.914 0.467 0.543 0.825 0.895 0.351 0.456
0.419 0.535 0.542 0.694 0.648 0.649 0.721 0.764 0.760 0.726 0.636 0.644 0.703 0.688 0.659 0.553 0.728 0.594 0.345 0.380 0.501 0.600 0.271 0.316
effectiveness, we select Qwen3-VL-32B and Qwen3-VL-8B as two representative models for studying team-size scaling and related factors in subsequent experiments. For strong models, we conduct most experiments with GPT-5-mini and Gemma4-31B, and include GPT-5.4 when necessary.
– 0.073 – 0.365 – 0.354 – 0.615 – 0.698 – 0.688 – 0.823 – 0.135 – 0.031 – 0.042 – 0.010 – 0.000
Table 9: Results on Parallel tasks for Qwen3-VL models with one to five agents. Comparison of performance with (broadcast) and without communication. The best results of both models are shown in bold. #Ag.
With comm. SR
Team size scaling effects. Table 9 shows the detailed result for Figure 7 (a) and 9(b). The results show that: (1) Simply increasing the number of agents does not necessarily improve performance. Without communication, adding more agents can even introduce disorder and redundant actions, leading to lower efficiency and worse task completion. However, when agents are allowed to communicate, multi-agent collaboration becomes more effective in coordinating actions and completing tasks. (2) More agents are not always better. Qwen3-VL-32B benefits from additional agents
CR
AUC
Without Comm. SR
CR
AUC
0.620 0.546 0.655 0.642 –
0.419 0.339 0.411 0.409 –
0.753 0.802 0.705 0.720 –
0.542 0.567 0.481 0.468 –
Qwen3-VL-8B 1 2 3 4 5
0.354 0.458 0.521 0.469 0.458
0.620 0.770 0.837 0.794 0.794
0.419 0.535 0.558 0.501 0.468
0.354 0.271 0.396 0.323 –
Qwen3-VL-32B 1 2 3 4 5
21
0.583 0.802 0.833 0.844 0.740
0.753 0.934 0.943 0.954 0.892
0.542 0.694 0.656 0.658 0.567
0.583 0.583 0.500 0.469 –
Table 10: Success rate and completion rate by number of objects for Qwen3-VL models on Parallel tasks. Best results for each model are shown in bold, and second-best results are underlined. Model
#Objects
1 Agent
2 Agents
3 Agents
4 Agents
5 Agents
SR
SR
SR
SR
SR
CR
CR
CR
CR
CR
Qwen3-32B-VL
2 3 4 5 6 7
0.875 0.938 0.875 0.938 0.938 0.969 0.938 0.969 1.000 1.000 0.765 0.863 1.000 1.000 1.000 1.000 0.941 0.980 0.875 0.938 0.625 0.812 0.938 0.984 1.000 1.000 1.000 1.000 0.750 0.922 0.438 0.650 0.688 0.838 0.688 0.888 0.812 0.938 0.800 0.867 0.389 0.630 0.667 0.935 0.778 0.944 0.667 0.935 0.500 0.880 0.385 0.604 0.615 0.901 0.583 0.905 0.750 0.964 0.667 0.917
Qwen3-8B-VL
2 3 4 5 6 7
0.750 0.875 0.875 0.938 0.938 0.969 0.812 0.906 0.875 0.938 0.588 0.784 0.824 0.922 0.882 0.941 0.824 0.922 0.706 0.863 0.312 0.562 0.500 0.828 0.500 0.828 0.500 0.828 0.438 0.766 0.250 0.562 0.375 0.713 0.438 0.800 0.333 0.747 0.500 0.787 0.111 0.444 0.056 0.685 0.167 0.713 0.278 0.741 0.111 0.704 0.077 0.472 0.077 0.483 0.154 0.769 0.000 0.637 0.077 0.692
Table 11: Results for 2-agent broadcast and no-communication protocols on Parallel and Sequential tasks. NA indicates that no handover attempt was made. Model
Mode
Parallel Task SR
CR
Sequential Task
AUC CAR
SR
CR
AUC HFR
100
100
75
75
SR (%)
CR (%)
Broadcast 0.833 0.945 0.689 0.019 0.823 0.904 0.620 0.322 Gemma4-31B No Comm 0.719 0.863 0.558 0.023 0.177 0.464 0.308 0.486 Broadcast 0.802 0.934 0.694 0.020 0.365 0.620 0.393 0.294 Qwen3-32B-VL No Comm 0.583 0.802 0.567 0.028 0.073 0.325 0.232 0.733 Broadcast 0.938 0.985 0.764 0.008 0.615 0.844 0.558 0.400 Qwen3.5-9B No Comm 0.865 0.939 0.680 0.036 0.062 0.325 0.231 0.903 Broadcast 0.458 0.770 0.535 0.010 0.073 0.332 0.208 0.662 Qwen3-8B-VL No Comm 0.274 0.552 0.342 0.017 0.042 0.182 0.132 NA
50 25 0
sults in figure 7(b). This indicate that multi-agent collaboration is especially helpful for more complex tasks with more target objects. For 7-object tasks, the SR nearly doubles for both models: from 0.385 to 0.750 for Qwen3-VL-32B, and from 0.077 to 0.154 for Qwen3-VL-8B. A task progress comprison of different team size shown in Figure 30.
50 25
0
20
40
Equivalent Steps 1 Agent
60
2 Agent
0 3 Agent
0
20
40
Equivalent Steps 4 Agent
60
D.2
5 Agent
Role of communication
Table 11 shows the detailed results in figure 9(a) when ablating communication from collaboration. Removing communication degrades both task performance and collaboration quality. In parallel tasks with a shared workspace, it generally increases CAR, indicating more conflicts actions between agents. In sequential tasks, the impact is more severe, as agents fail to coordinate crossregion handovers, causing HFR to rise substantially.
Figure 30: Task progress over steps for 1 to 5 agents under the broadcast protocol on the parallel task. Results are obtained with Qwen3-VL-32B.
up to four agents, while Qwen3-VL-8B reaches its best performance with three agents and drops afterward. This suggests that weaker models are more affected by coordination overhead. (3) Collaboration can partially compensate for limited individual model capability, but cannot fully remove the capability gap. With communication, three Qwen3-VL-8B agents approach the single-agent performance of Qwen3-VL-32B, but the final performance remains constrained by the underlying model capability. Table 10 shows the detailed re-
D.3
Comparison of Different Collaboration Modes
Table 12 gives the detailed results of Figure 8(a). The results show that centralized collaboration generally improves division of labor and task effi22
Table 12: Results of different communication protocol under parallel and sequential tasks across models. Spatial Err denotes the frequency of spatial errors, where an agent attempts to move to a location outside its assigned region. Best results within each model are shown in bold, and cross-model best results are highlighted in yellow. Model
Mode
Parallel SR
CR
Sequential
AUC DOL Tokens CAR
SR
CR
AUC Spatial Err Tokens HFR
Closed-source MLLMs Broadcast 0.900 0.961 0.700 0.773 4331.9 0.015 0.767 0.904 0.640 Discuss 0.767 0.932 0.730 0.717 6266.6 0.001 0.767 0.911 0.652 Leader 0.784 0.934 0.752 0.883 6558.6 0.009 0.767 0.870 0.639
GPT-5.4
0.017 0.021 0.023
7328.1 0.228 7619.3 0.177 8085.5 0.223
Open-source MLLMs Broadcast 0.833 0.945 0.689 0.794 Discuss 0.802 0.855 0.638 0.731 Leader 0.865 0.964 0.736 0.840 No Comm 0.719 0.863 0.558 0.706
5381.8 6852.5 6863.0 3577.4
0.019 0.823 0.904 0.620 0.001 0.812 0.922 0.625 0.003 0.854 0.963 0.652 0.023 0.177 0.464 0.308
0.038 0.025 0.016 0.297
8042.4 8978.3 8723.5 4539.7
0.322 0.224 0.138 0.486
Broadcast 0.802 0.934 0.694 0.802 Discuss 0.808 0.929 0.694 0.818 Qwen3-32B-VL Leader 0.833 0.938 0.706 0.855 No Comm 0.583 0.802 0.567 0.629
4682.0 6351.7 5964.7 3017.7
0.020 0.365 0.620 0.393 0.003 0.333 0.583 0.374 0.012 0.396 0.622 0.408 0.028 0.073 0.325 0.232
0.166 0.100 0.046 0.451
6749.4 8324.3 7516.5 3559.0
0.294 0.304 0.378 0.733
Broadcast 0.458 0.770 0.535 0.676 Discuss 0.458 0.783 0.494 0.702 Leader 0.490 0.785 0.548 0.642 No Comm 0.274 0.552 0.342 0.357
4585.4 7206.1 5701.5 3310.1
0.010 0.073 0.332 0.207 0.011 0.021 0.283 0.195 0.005 0.052 0.332 0.227 0.017 0.042 0.182 0.132
0.312 0.319 0.223 0.613
6788.2 0.662 8934.5 0.601 7161.4 0.658 3475.1 NA
Qwen3-8B-VL
Success Rate (SR)
Parallel
100%
AUC +3.6%
70%
80% 60% 40%
Sequential
80%
+8.3% +1.2%
+5.2% +6.2%
Token Consumption
+2.9% +4.9%
60%
+4.6%
50% 32B N=232B N=4 8B N=2 8B N=4
32B N=232B N=4 8B N=2 8B N=4
-7.3%
40%
30%
+1.0%
32B N=2
+0.6%
40%
20%
6.0k 5.5k 5.0k 4.5k 4.0k
+0.9%
20%
8B N=2
32B N=2
8B N=2
Broadcast
-0.8k -0.3k
-1.4k
Shared Memory
-1.8k
83.3
80
32B N=232B N=4 8B N=2 8B N=4 7.0k 6.0k 5.0k 4.0k
100
-0.7k
-2.5k
Score (%)
Gemma4-31B
90.5
93.8 85.5
78.5
66.7
60
54.8
49.0
61.6
40 20
32B N=2
8B N=2
0 All Weak (8B+8B)
SR
CR
Strong Leader (32B+8B)
AUC
All Strong (32B+32B)
100
100
75
75
SR (%)
CR (%)
Figure 31: Comparison of explicit broadcast and implicit shared memory Figure 32: Effect of leader capability in protocol. The models are from the Qwen3-VL family. centralized collaboration. The models are from the Qwen3-VL family.
50 25 0
models, centralized collaboration also helps reduce spatial cognition errors in sequential tasks, suggesting that leader-based coordination can provide useful guidance when agents have limited planning or spatial reasoning ability. Figure 33 shows the comparison of task progress under single agent and different communication protocol in two-agent team.
50 25
0
20
40
Equivalent Steps 1 Agent
No Comm
60 Broadcast
0
0 Discuss
20
40
Equivalent Steps Leader
60 Leader-MM
Shared Memory. Figure 31 reports the absolute values of success rate, AUC, and token consumption for broadcast and shared memory, providing a more detailed version of Figure 10.
Figure 33: Task progress over equivalent steps under different communication protocols on parallel tasks. Results are obtained with Gemma4-31B. Leader-MM denotes the vision-augmented leadership protocol.
D.4
ciency, reflected by higher DOL and AUC. In contrast, the discuss protocol is more effective at reducing conflicts, leading to lower CAR. For weaker
Mix Team: Strong Model as Leader.
We further investigate the role of leader in centralized mode. As shown in Figure 32, using a 23
Table 13: Results under different prior-information settings on parallel tasks. Full prior includes both object appearance and location information. Without appearance removes object appearance description, without location removes prior location information, and no prior removes both. Model
Full Prior SR
CR
Without Appearance Without Location
AUC
SR
CR
AUC
SR
GPT-5-mini 0.917 0.982 0.776 0.865 0.955 Gemma4-31B 0.833 0.945 0.688 0.833 0.937 Qwen3-32B-VL 0.802 0.934 0.694 0.802 0.923
0.730 0.662 0.655
0.344 0.691 0.436 0.271 0.609 0.431 0.229 0.589 0.347 0.302 0.570 0.356 0.260 0.591 0.365 0.344 0.688 0.456
Table 14: Results on parallel tasks without location information. Each cell reports SR/CR/AUC. Qwen-* is short for Qwen3-VL, Inter-* is short for InterVL3.5. 1-Agent
2-Agent
Closed-source MLLMs GPT-5-mini
CR
AUC
0.80 0.70
0.50
Qwen-8B 0.031 / 0.162 / 0.103 0.021 / 0.199 / 0.101 Qwen-32B 0.219 / 0.520 / 0.344 0.260 / 0.591 / 0.365 Qwen-235B 0.115 / 0.389 / 0.264 0.271 / 0.581 / 0.332 Qwen3.5-9B 0.292 / 0.582 / 0.389 0.292 / 0.649 / 0.435 Gemma-31B 0.240 / 0.561 / 0.358 0.229 / 0.589 / 0.347 Intern-38B 0.010 / 0.146 / 0.087 0.073 / 0.253 / 0.144 Intern-241B 0.021 / 0.201 / 0.113 0.062 / 0.217 / 0.114 Llama-4 0.021 / 0.089 / 0.058 0.000 / 0.107 / 0.056
GPT-5-mini Gemma4-31B Qwen3-32B-VL
Full Prior w/o Appearance w/o Location
No Prior
Figure 35: Completion rate under different priorinformation settings on parallel tasks.
ure 34 shows the team size scaling effect under stronger exploration demands, where prior location information is removed. Compared with the full-information setting, Qwen3-32B-VL exhibits a similar but weaker inverted-U-shaped trend, suggesting that moderate team scaling remains helpful. For Qwen3-8B-VL, SR remains low across team sizes, indicating a capability bottleneck. Moreover, when grouping cases by object count, larger teams still show a milder performance drop. More detailed results in Table 15.
(b)
Figure 34: Team size scaling effect without location prior. (a) SR and CR change as the number of agents increases. (b) SR versus the number of target objects, with fitted trend lines for different team sizes.
Table 15: Results on Parallel tasks without location prior for Qwen3-VL models with one to four agents, under the broadcast protocol. The best results of each model is highlighted in bold.
stronger leader improves the mixed team, but does not make it approach the all-strong setting. The mixed team nearly matches the all-strong team in CR, but its SR is only slightly above the average of the all-weak and all-strong teams, and its AUC is even below this average. This indicates that a strong leader mainly improves task coverage, but efficient and overall success still depends heavily on the worker’s capability. D.5
SR
0.90
0.60
0.323 / 0.630 / 0.436 0.344 / 0.691 / 0.436 Open-source MLLMs
(a)
AUC
1.00
Completion Rate (CR)
Model
CR
No Prior
Model
Performance Under Imperfect Information.
#Agents
SR
CR
AUC
Qwen3-8B-VL
1 2 3 4
0.031 0.021 0.094 0.062
0.162 0.199 0.298 0.307
0.103 0.101 0.168 0.163
Qwen3-32B-VL
1 2 3 4
0.219 0.260 0.292 0.219
0.520 0.591 0.657 0.608
0.344 0.365 0.404 0.374
Ablation on prior information To investigate how different types of prior information affect agent performance, we conduct an ablation study by removing object appearance descriptions and prior location information. As shown in fig-
Missing location priors Table 14 reports detailed results under the no-prior-location setting on parallel tasks across different models. Fig24
Strong Models
100
80
80
80
60
60
60
Performance (%)
Parallel Normal
40
2
3
4
5
6
7
0 100
80
80
60
60
40
40 Gemma4-31B GPT-5.4 GPT-5-mini Qwen3.5-27B Gemma4-26B
20
Performance (%)
0
2
3
2
4
5
6
7
0 70
70
60
60
3
20 4
5
6
7
Qwen3.5-27B GPT-5-mini Qwen3.5-9B Gemma4-31B
10
2
3
4
5
6
7 Qwen3-32B-VL Qwen3-235B-VL Qwen3-8B-VL
2
3
4
5
Number of Objects
6
7
5
6
7 Qwen3-8B-VL LLaMA-4 InternVL-241B GLM-4.6V GLM-4.6V-Flash
0
2
3
4
5
6
30
7 InternVL-241B LLaMA-4 InternVL-38B
25 20 15 10 5
10 0
4
10
20
20
3
20 Qwen3.5-9B Gemini-3.1-Pro Qwen3-235B-VL Qwen3-32B-VL InternVL-38B
30
30
2
30
40
40
0
40
50
50
GLM-4.6V Qwen3-32B-VL Qwen3-8B-VL LLaMA-4 GLM-4.6V-Flash
50
20
80
0
InternVL-241B Gemini-3.1-Pro Gemma4-26B InternVL-38B Qwen3-235B-VL
20
100
Performance (%)
Sequential Normal
0
Metric SR CR
Weak Models
40
40 GPT-5-mini GPT-5.4 Qwen3.5-9B Qwen3.5-27B Gemma4-31B
20
Parallel No Location
Middle Models 100
100
2
3
4
5
Number of Objects
6
7
0
2
3
4
5
Number of Objects
6
7
Figure 36: Performance by number of objects across task conditions and model groups.
ure 35, prior location information is critical for task completion, as it helps agents quickly locate target objects and reduces inefficient exploration. Removing location information leads to a substantial performance drop across all models. In contrast, the effect of appearance descriptions is model-dependent: GPT-5-mini benefits from appearance cues, Gemma4-31B shows only minor changes, and Qwen3-32B-VL performs worse with appearance-only priors than with no prior. A possible reason is that models with weaker fine-grained visual grounding are more susceptible to misleading appearance descriptions and may confuse visually similar objects. Detailed results are reported in Table 13.
location information is noisy, comparing singleagent and two-agent settings. D.6
Figure 36 reports performance by the number of target objects across model groups. As each target object corresponds to a subgoal, a larger number of target objects indicates higher task complexity. Overall, both completion rate and success rate decline as task complexity increases.
E
GPT-5-mini Qwen3.5-9B Qwen3-32B-VL
#Agents
SR
CR
AUC
1 2 1 2 1 2
0.760 0.823 0.781 0.812 0.542 0.698
0.902 0.950 0.911 0.903 0.750 0.878
0.719 0.709 0.665 0.663 0.526 0.620
Emergent collaborative behavior
By analyzing the communication and action trajectories under the basic broadcast protocol, we identify several emergent collaborative behavior patterns. In addition to the analysis in Section 4.3, this section provides representative examples and further details.
Table 16: Results on Parallel tasks with noise location prior cross models. Model
Additional Performance Analysis
E.1
Examples
Table 17 presents representative examples of collaborative behaviors, with the relevant text spans highlighted in yellow. E.2
Detailed Analysis
Coordination Queries Indicate Collaboration Challenges. In the main text (Table 2), coordination query correlates negatively with SR and CR.
Noisy location priors. Table 16 reports detailed results on parallel tasks where 30% of the prior 25
Table 17: Examples of collaborative behaviors emerge in the agent interactions in MECoBench under broadcast protocol. Category Information Sharing
Task Coordination
Behavior
Example
Info-Plan Info-Loc
“I will grab the two puddings and check the fridge for the wine. Could you go to the living room to find the cupcake?” “I am in the bathroom and will head to the kitchen to look for the pudding and cupcake.”
Info-Obj
“I see the cupcake on the desk. I will grab it and then head to the living room.”
Act-Del
“Please go to the kitchen to grab the second apple, and then we can meet at the coffee table.”
Task-Assign
“I’ll handle the kitchen and the bathroom.”
Coord-Q.
State confirmation: “Do you have the cupcake and the second apple? If not, please check the fridge.” Location confirmation: “I’m at door 287 holding the cupcake. Are you at door 287 and ready to receive? Please confirm.” Self-correction: “I accidentally grabbed a waterglass instead of the wineglass. I will put it back and then get the wineglass.” Peer-correction: “Agent 0, I see you’re holding milk, but we need the juice. Please drop the milk and grab the correct juice from the kitchen table.”
Alignment & Correction Correct.
Table 18: Coordination queries are associated with harder coordination states. Q. denotes coordination query.
Table 20: Correction effects in parallel tasks stratified by the number of wrong grabs. #WG Group
Task
Group #Case Avg. SR Stuck Neg. #Obj. (%) (%) (%)
Parallel
w/ Q. w/o Q.
370 1727
4.70 4.35
64.6 74.7
8.4 2.0
11.1 4.5
Sequential w/ Q. w/o Q.
591 627
4.51 4.30
32.3 47.8
3.9 1.6
6.6 3.2
Table 19: Overall and wrong-grab-conditional effects of correction. Corr. and WG denote correction and wrong grab. WG-Cond. denotes cases where at least one wrong grab happened. Task
Cond.
SR (%)
CR #WG (%)
w/ Corr. Overall w/o Corr. ∆
67.3 74.1 −6.7
90.1 88.5 +1.6
3.05 1.37 –
w/ Corr. w/o Corr. Raw ∆ Adj. ∆
61.1 66.2 −5.1 +2.4
88.3 87.0 +1.3 +3.5
3.79 2.84 – –
w/ Corr. 29.8 Overall w/o Corr. 41.1 ∆ −11.3
67.4 59.8 +7.7
2.68 0.76 –
w/ Corr. 26.9 67.2 w/o Corr. 21.1 53.4 Raw ∆ +5.8 +13.8 Adj. ∆ +11.8 +15.4
4.33 3.02 – –
Parallel WG
Sequential WG
Group
#Case
SR CR (%) (%)
∆SR/CR
1
Corr. No Corr.
70 82.9 95.2 +2.2/ + 3.7 351 80.6 91.5
2
Corr. No Corr.
53 69.8 89.3 −6.2/ − 1.5 200 76.0 90.8
3
Corr. No Corr.
44 68.2 91.9 +6.9/ + 4.1 80 61.3 87.8
4+
Corr. No Corr.
118 41.5 82.4 +7.4/ + 6.8 208 34.1 75.6
Adjusted effect
∆SR/CR +2.4/ + 3.5
reports failing to locate a target object at a specific location, and stuck, where an agent reports being stuck in middle. As shown in Table 18, cases with coordination query involve more objects on average and exhibit substantially higher stuck and negative report rate. This pattern holds across both task types, suggesting that coordination query serves as a reactive signal to genuine collaboration obstacles. Correction Helps Recover from Errors Section 4.3 shows that correction correlates negatively with SR but positively with CR, suggesting a potential role in error recovery. To examine this effect, we analyze its co-occurrence with wrong grabs (WGs), defined as actions that grasp task-irrelevant objects. Table 19 shows that episodes with correction contain substantially more WGs than those without, indicating that correction is mainly trig-
To further investigate the underlying cause, we introduce two behavioral indicators of collaboration difficulty: negative report (Neg.), where an agent 26
Table 21: Representative examples of multi-agent collaboration failure modes. Failure Mode Agent-induced flict
Con-
Hallucinated Belief Propagation
Subtype
Representative Trace Excerpt
Duplicate Grab
Agent 1: “I found the wine on the coffee table. I will grab it and bring it to the living room.” Agent 4: “I am grabbing the wine from the coffee table. I will take it to the living room.” ⇒ Both agents execute [grab] <wine>.
False Completion
“All items have been successfully placed on the bathroom counter. The task is complete.” ⇒ All agents execute [wait] until the end.
Object Misidentification
Agent 3: “I have found and will now grab the wineglass from the bedroom cabinet.” Agent 3 executes [grab] <dishbowl> (misidentifying dishbowl as wineglass), then [put] <dishbowl> on the kitchentable. Agent 0 & 2: “Agent 3 has already grabbed the wineglass—I will focus on the plate / waterglass.” ⇒ All agents stop searching for the actual wineglass. Task ends with wineglass unsatisfied.
gered by execution errors. Since WG episodes with correction also exhibit greater error severity, direct comparison may underestimate its recovery benefit. We therefore stratify WG episodes by error count and compute a weighted adjusted effect using shared stratum weights for both groups (Table 20). After adjustment, correction consistently improves both SR and CR across task types, confirming its effectiveness in error recovery. Stratum-level results further show larger gains under higher WG counts, suggesting that correction is especially useful when errors are more severe.
F
internvl-3.5-241b glm-4.6v-flash qwen3-32b-vl qwen3-235b-vl qwen3-8b-vl gemma4-26b gpt-5.4 gemma4-31b llama-4 glm-4.6v gpt-5-mini internvl-3.5-38b qwen3.5-9b gemini-3.1-pro qwen3.5-27b 0
Examples
Detailed Analysis
Table 22: Failure mode in collaboration. Freq. denotes the fraction of episodes containing the behavior; lift is the ratio of its rate in failed episodes to that in successful episodes. Behavior
Task
8.3% InternVL GLM GPT Gemini
4.2% 3.1% 5
10
15
Dup-Grab Rate (%)
20
Qwen3-VL Qwen3.5 Gemma Llama
25
30
Table 22 summarizes the occurrence frequency of different failure modes and their impact on success rate (SR), aggregated over all models and team sizes under the broadcast protocol. Figure 37 compares duplicate-grab rates across models in twoagent teams.
Table 21 presents representative examples of different failure modes. F.2
12.5% 11.5% 11.5%
Figure 37: Duplicate grab rate of different models in parallel task under 2-agent team and broadcast protocol.
Failure Mode Analysis
In this section, we provide additional examples and a more detailed analysis of the multi-agent collaboration failure modes presented in Section 6.2. F.1
21.9% 19.8% 18.8% 18.8% 17.7% 17.7% 16.7% 15.6% 14.6%
Freq.
Lift SR w/
Hallucinated Parallel completion Sequential
1.0% 1.5%
∞ ∞
Grab conflict Parallel
22.9% 1.5× 71.4%
∆SR
0.0% −73.7 0.0% −40.9 −9.0
27