Conceptio › Archive › arXiv CS
arXiv CSopen access

Testing Interchangeability in LLM Agent Teams

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Testing Interchangeability in LLM Agent Teams

Jianxin Gao 1 Tianyi Yu 2 Linna Deng 1 Runze Li 3 Zining Wang 4

arXiv:2609.05279v1 [cs.AI] 4 Sep 2026

Abstract

the job. Frameworks are built on that assumption, since a role is a configuration entry and any agent satisfying its contract can fill it (Hong et al., 2024; Qian et al., 2024; Wu et al., 2024), and recent work makes the interchangeability an explicit design goal (Chen et al., 2026). The same systems are also described in the vocabulary of human teams: roles, division of labour, shared plans, and increasingly a persistent memory that lets a group get better at working together over a long horizon.

Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on heldout tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team’s history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.

The two descriptions disagree about one concrete operation, replacing a teammate. If agents are processes, replacement is free as long as the replacement is equally competent. If agents are teammates in the way people are teammates, replacement is not free, because part of what a long-running team knows is not about the task but about each other. Human teams show this: working repeatedly with the same specific colleagues predicts performance beyond individual experience, through a shared map of who knows what and how to hand work over (Wegner, 1987; Lewis, 2003; Reagans et al., 2005; Huckman et al., 2009). We ask whether LLM agent teams have an analogue, and how large it is. The short answer is that they are more interchangeable in outcome than in coordination cost, and that the gap is wider where the task demands more coordination and where the team has been together longer. The question is worth asking now because the machinery that would produce such an effect is in place. Agents keep notes across episodes (Xu et al., 2025; Park et al., 2023), populations of LLM agents settle on shared conventions on their own (Ashery et al., 2025), and long exchanges between two models drift into stable regimes shaped by the partner (Ko & Geiping, 2026). What is missing is a measurement separating that from the plainer possibility that the team simply got better at the task.

1. Introduction Every deployed multi-agent system replaces agents. A provider degrades, a container restarts, a scheduler moves a role onto a different replica, a model version is retired. Underneath all of it sits one assumption: an agent occupying a role is interchangeable with any other agent that can do

Existing evidence points both ways and stops short of the question. Agashe et al. (2025) report that in Hanabi a GPT-4turbo agent paired with an unfamiliar partner loses nothing relative to self-play, while a reinforcement-learning agent trained by self-play loses a great deal; but their unfamiliar partner is a different kind of agent, and their LLM carries no memory between games, so there is no team history to disrupt. Ramesh et al. (2026) note that their seventeen-model Hanabi study uses homogeneous teams by construction,

1 China Agricultural University 2 Tianjin University of Finance and Economics 3 Jilin University 4 Tianjin University of Science and Technology. Correspondence to: Jianxin Gao <[email protected]>.

Preprint. September 7, 2026.

1

Testing Interchangeability in LLM Agent Teams

which leaves cross-play adaptation untested, and Wang et al. (2026) find that agent dyads in repeated reference games align on labels without the alignment depending on which partner produced it. None of this varies partner identity while holding model, role and amount of experience fixed.

et al., 2021; Hu et al., 2021). Hanabi is the canonical arena because almost all of its difficulty is conventional (Bard et al., 2020), and Overcooked plays that role for embodied coordination (Carroll et al., 2019; Strouse et al., 2021). Closest in spirit is Shih et al. (2021), who split a policy into a part that depends on the rules and a part that depends on the partner so that the second can be relearned for someone new. We borrow the cross-play instrument and that split, and apply both to teams whose adaptation is textual rather than parametric.

The swap test. We separate the two possibilities by construction (Figure 1). Build K independent teams from one base model, with the same role prompts, on the same task pool. Let each run a set of formation episodes, writing to a persistent private notebook after each one. Then trade role-matched agents between two teams and evaluate on held-out tasks. The two traded agents share model, role, prompt and amount of experience, and differ in whom they accumulated that experience with. To keep the disruption of a personnel change separate from the identity of the person, we also run a placebo in which an agent is removed and reinstated under the same announcement.

Coordination in LLM agent teams. The cross-play in Agashe et al. (2025) pairs models of different families, or an LLM with a reinforcement-learning agent, so the gap there mixes capability differences with convention mismatch. Sun et al. (2025) build Collab-Overcooked around asymmetric capability and asymmetric information, and add processlevel scores that separate the agent driving coordination from the agent following. In Hanabi, leading models reach 15 to 18 of 25 while purpose-built agents exceed 23 (Ramesh et al., 2026). Zhu et al. (2025) evaluate coordination at the level of communication topology. Most breakdowns in these systems come from specification and inter-agent misalignment rather than single-agent incompetence (Cemri et al., 2025), and agents often infer a partner’s plan correctly and then fail to act on it (Goel et al., 2026).

The design is borrowed from multi-agent reinforcement learning, where the gap between self-play and cross-play is the standard diagnostic for agents that have locked onto arbitrary conventions (Hu et al., 2020; Bard et al., 2020; Hu et al., 2021). The transfer is not mechanical. There, crossplay recombines independently trained policies and the gap reflects frozen weights. Here nothing is trained: adaptation lives entirely in text the agent wrote about its partner, so we can cut that text into pieces and ask which piece carries the cost.

Conventions and drift among LLMs. Populations of LLM agents converge on shared naming conventions, in the sense of Lewis (1969), and inherit collective biases while doing so (Ashery et al., 2025). Extended exchanges between two models settle into attractor states, with each pulled asymmetrically toward the other (Ko & Geiping, 2026); coordination structure can be induced by prompt design alone (Riedl, 2026); and behavioural properties spread contagiously across agent networks (Weckbecker et al., 2026). Wang et al. (2026), described above, is the nearest result; they study two-player linguistic reference with an analytic control, while we study task teams with roles and persistent memory, re-pair them for real, and measure operational cost rather than lexical convergence.

Contributions. 1. A swap test for LLM agent teams, with a placebo control and a decomposition that separates the value of generic experience from the residue that is specific to a partner, giving one ratio ρ. 2. Task score and coordination cost come apart. A swap costs little in score and 16 to 63 percent in communication per unit of progress; in Hanabi it costs more than replacing the agent with an inexperienced one. 3. The loss is uneven and it is not fixed. It concentrates in the seat that initiates plans, is paid mostly by the agent that stayed, varies with decoding temperature, and grows with how long the team has been together.

Team familiarity in humans. The framing comes from research on transactive memory (Wegner, 1987; Lewis, 2003) and team familiarity, where repeated work with the same colleagues improves performance beyond individual experience (Reagans et al., 2005; Espinosa et al., 2007; Huckman et al., 2009), and from the observation that speakers converge with a specific partner on shortened referring expressions a third party does not share (Clark & Wilkes-Gibbs, 1986). We take these as a source of hypotheses, not as a claim about agent cognition.

2. Related Work Ad hoc teamwork and zero-shot coordination. An agent joining a team it did not train with is the founding problem of ad hoc teamwork (Stone et al., 2010; Mirsky et al., 2022); the mirror image, whether independently trained agents cooperate at all, became zero-shot coordination, where the gap between self-play and cross-play exposes agents that have latched onto arbitrary conventions (Hu et al., 2020; Treutlein 2

Testing Interchangeability in LLM Agent Teams 1. Formation E = 10 episodes team 1

2. Roster change example: initiators traded

aI1

aR 1

aI2

aR 1

κ|π

κ|π

κ|π

κ|π

3. Evaluation six conditions

task score T held-out tasks R = 10 episodes

team 2

aI2

aR 2

aI1

aR 2

κ|π

κ|π

κ|π

κ|π

coordination cost C

Figure 1. The swap test. Colour marks provenance: after the trade, team 1 holds an agent formed in team 2. Superscripts mark the seats, I for the agent that initiates plans and R for the one that responds; κ and π are the two notebook sections. Teams are formed independently from one base model on a shared task pool, each agent maintaining a private notebook with a section for the task and a section for the partner. A roster change then trades role-matched agents between two teams, and both are evaluated on held-out tasks. The six conditions of Section 3.2 differ only in step 2: who occupies the seat, and which part of the notebook travels with them.

3. Fungibility and the Swap Test

change. Swap. Two teams trade their role-j agents, notebooks included. This is cross-play with model, role and experience held fixed. Swap, cleared. As Swap, but the arriving agent’s partner notes are deleted and its task notes kept. Asks whether carrying the wrong partner model is better or worse than carrying none. Amnesia. No swap: the team is intact, but one agent’s own partner notes are deleted. Asks what an unbroken team loses by forgetting its teammate. Naive. The role-j agent is replaced by a fresh agent of the same model with an empty notebook. This is the fresh-replacement baseline.

3.1. Teams, formation, and notebooks A team assigns agents to n fixed role slots. Every agent is the same frozen base model under a role-specific system prompt, so the only thing separating two agents in the same slot is the notebook µ: a short document the agent rewrites after each episode, prepended to its context in the next one. Agents keep the notebook in two labelled sections. Task notes κ hold facts about the environment, such as recipes, affordances and failure modes, which would be equally true with any partner. Partner notes π hold facts about this teammate: what they reliably do unprompted, what has to be stated explicitly, which phrasings of a hand-off have worked, and standing agreements. The split is the operational form of the distinction between task knowledge and transactive memory (Wegner, 1987), and it is what makes the partnerspecific part separable in a system whose entire learned state is text. The prompt asks only for the two headings and a budget of two hundred tokens, so what ends up under the partner heading is the agent’s own judgement rather than a form we designed.

3.3. Metrics Each condition yields a task score T , normalised to [0, 1] from the benchmark’s own scale, and a coordination cost C, the communication effort spent per unit of task progress. Writing Tcond for a condition mean,

A cohort is K teams built identically, differing only in random seed and task order. Each team runs E formation episodes. Formation therefore produces a mixture of task knowledge and partner knowledge in every team; the swap test asks how much of the mixture is the second kind. 3.2. Conditions

V = Tplacebo − Tnaive ,

(1)

Π = Tplacebo − Tswap ,

(2)

ρ = Π/V.

(3)

V is what E episodes of formation buy relative to an untrained replacement. Π is what remains once we subtract everything an equally experienced stranger already knows. Their ratio ρ is the quantity of interest. At ρ = 0 a team’s history is entirely portable and its members are interchangeable; at ρ = 1 every gain from working together is tied to the pairing. Nothing bounds ρ at 1: above it, an arriving veteran leaves the team worse off than someone with no experience, which can happen once mismatched conventions are in play. Two further contrasts locate Π:

All conditions are evaluated on held-out tasks after formation, and all except Intact are preceded by the same roster-change announcement, so that the disturbance of a personnel change is held constant. Intact. The team as formed. This is self-play. Placebo. One agent is removed and immediately reinstated, with the announcement and context reset of a real roster 3

Testing Interchangeability in LLM Agent Teams

W = Tplacebo − Tamnesia ,

(4)

Σ = Tswap, cleared − Tswap .

(5)

the formation tasks and the upper level the held-out ones, so evaluation tasks are never seen during formation. We report progress completeness rather than success rate: at ten episodes per condition the sampling error of a binary measure exceeds the differences we are after, which is why Sun et al. (2025) introduce the finer one. For Hanabi we report score out of 25. Coordination cost is messages per completed sub-task in Collab-Overcooked; Hanabi has no side channel, so we use hints spent per point, the same quantity in the currency the game provides.

W asks how much of Π is reproduced by deleting an intact team’s own partner notes without touching its roster; if W ≈ Π, the residue sits where we asked the agents to put it. A positive Σ means the arriving agent does better once its notes about a former partner are removed, so an outdated model of a teammate is worse than none. Each quantity has a counterpart on C, with the signs reversed so that a larger number again means a worse outcome; we write ρC for the ratio computed there.

Agents and cohorts. Agents are prompted, not trained. Each is the base model plus a role prompt, the current notebook and the episode transcript; after each episode it rewrites its notebook under a token budget, with the two section headings fixed by a template. No weights are updated and no GPU is used. Every agent in the main study is GPT-5.6 Luna in its low-cost tier, so a swap never mixes partner identity with capability, which is what an LLM cross-play gap has measured until now (Agashe et al., 2025). Section 4.5 varies the model and the decoding temperature deliberately, one cohort at a time and never within a team.

Every contrast is formed inside a team before being averaged across teams. Writing Tkcond for team k’s mean score under b = K −1 P (T placebo −T swap ), and likewise a condition, Π k k k for V , W and Σ. A team’s own level therefore cancels and what is averaged is the effect of the roster change on that team, not a difference between two groups of teams.

4. Experiments

The main study keeps the benchmarks’ own settings, temperature 0.7 and a time limit of γ = 1.5 times optimal (Sun et al., 2025), and their sample size per reported cell: ten repetitions of a task in Collab-Overcooked (Sun et al., 2025), ten games per configuration in Hanabi (Ramesh et al., 2026). Each condition is therefore measured on R = 10 episodes, spread over five held-out tasks run twice rather than one task run ten times, since a team has to be measured under six conditions. Formation is E = 10 episodes, the five formation tasks run twice; Hanabi has no task pool, so formation and evaluation use disjoint sets of deals. The same held-out tasks or deals are replayed under every condition, so conditions are compared on identical work.

4.1. Settings Partner-specific coordination should be largest where partners have to model each other, so we order three settings by how much of the joint task resists decomposition into independently executable pieces. We use published benchmarks rather than building one. Collab-Overcooked, low coupling. Levels 1 and 2 of Sun et al. (2025). Two agents hold disjoint equipment and only one is given the recipe, so they must exchange structured messages to synchronise. At these levels the recipes need few mandatory hand-offs and most of the work is separable once the recipe has been communicated.

Each cohort is K = 8 teams. Swaps are role-matched and applied as four disjoint pairs, with the traded seat drawn at random per pair, so all eight teams are perturbed and measured. The two teams in a pair share one swap event, so for swap contrasts the independent unit is the pair and not the team: intervals on Π and ρ come from a bootstrap that resamples the four pairs rather than the eight teams. Two arms sit alongside the six conditions: a pair of swap runs that fix the seat, one forcing each, drawn separately from the randomly seated main arm and reported in Table 5; and a continuation of ten consecutive episodes after a swap, which gives Figure 2. Every condition starts from an identical copy of the team’s post-formation state. The six are independent forks rather than a sequence, so nothing done under one condition can reach another and there is no condition order to confound the comparison. Within a condition the agents keep rewriting their notebooks as usual, so a condition mean is the average over the first ten episodes after the roster

Collab-Overcooked, high coupling. Levels 5 and 6 of the same benchmark, where the minimum number of collaborative actions is large and sub-tasks must be interleaved rather than batched. The agents, the protocol and the metrics are the same as at low coupling, but the levels differ in more than interdependence: recipes are longer and involve more equipment, so overall task complexity rises with the coordination requirement. We therefore read the pair as an observational coupling gradient rather than as a controlled intervention on interdependence. Hanabi. The two-player setting of the LLM-Coordination benchmark (Agashe et al., 2025), where players see their partner’s hand but not their own. It anchors the coupled end, since the meaning of a hint is fixed only by agreement between the two players (Bard et al., 2020). In each Collab-Overcooked setting the lower level supplies 4

Testing Interchangeability in LLM Agent Teams Table 1. Condition means in each benchmark’s own units. Progress completeness is on 0–100 and Hanabi score on 0–25; coordination cost is messages per completed sub-task in Collab-Overcooked and hints per point in Hanabi. Level 2 is the held-out set of the low-coupling setting and level 6 of the high-coupling setting.

change and not the instantaneous cost of the change. That is the quantity an operator actually pays, and Figure 2 separates its two parts, the shock at episode one and the rate at which it decays. Aggregation is per team first, then across teams. Betweenteam standard deviation under the intact condition is 2.3 and 2.7 progress-completeness points in the two CollabOvercooked settings and 1.6 of 25 in Hanabi.

Condition Task score Intact Placebo Swap Swap, cleared Amnesia Naive Coordination cost Intact Placebo Swap Swap, cleared Amnesia Naive

Protocol signatures. To make a team’s protocol measurable we extract, from its formation transcripts, a protocol signature: the distribution over hand-off message templates, the field order agents use when reporting state, alias choices for objects and sub-goals, and, in Hanabi, the mapping from hint type to the action the receiver takes. Extraction is rule-based and automatic: hand-off templates are matched against the benchmark’s own message grammar, field order and aliases are read off the structured fields, and the Hanabi mapping is tabulated from hint-action pairs. No human coding is involved and the rules are fixed before the evaluation runs. Signatures are compared by Jensen–Shannon divergence d. This gives two quantities: the mean pairwise d between the formation signatures of teams in the same cohort, which measures how far apart independently formed teams end up; and, after a swap, how far the receiving team moves toward the donor, before after ∆ = d(Srecv , Sdonor ) − d(Srecv , Sdonor ),

level 2 level 6 Hanabi 92.5 91.6 90.0 90.8 90.9 83.1

64.1 63.2 58.0 59.6 59.0 43.4

15.8 15.8 13.7 14.8 14.5 11.2

3.20 3.33 3.85 3.74 3.75 4.50

4.09 4.33 6.53 5.53 5.95 7.39

0.61 0.64 1.04 0.85 0.94 1.00

(Ramesh et al., 2026) and above 23 for purpose-built agents. Formation buys something real but modest: a naive replacement drops the team to 83.1, 43.4 and 11.2, so V is 0.085, 0.199 and 0.183 (Table 2). Score barely moves, coordination cost moves a lot. Table 1 shows the shape of the answer. On task score, replacing a teammate with an equally experienced stranger is close to free. In the low-coupling setting the team goes from 91.6 under the placebo to 90.0 under the swap against 83.1 for a naive replacement: about a sixth of what inexperience costs. Even in Hanabi, where we expected the largest effect, the swap costs 2.0 points against the 4.6 that experience is worth.

(6)

so ∆ > 0 indicates that the receiving team’s observed signature shifted toward the donor team’s after the swap. Scale. The main study is 8 teams × (10 formation + 6×10 conditions + 10 recovery) episodes in each of three settings, plus 2 × 10 seat-targeted episodes per team in both CollabOvercooked settings: 2,240 episodes, of which 1,600 are Collab-Overcooked. The ablations of Section 4.5 add 1,950, giving 4,190 in total. One full pass of Collab-Overcooked is 30 tasks at ten repetitions, so its share of this study costs about twelve times what evaluating one model on it costs. Episodes are short, the models run in their cheap tiers, and nothing is trained, so the run is a few hundred dollars of API calls. A single swap test without the ablations is a few dozen episodes, cheap enough to run before rotating an agent into production.

Coordination behaves differently. The same swap that cost almost nothing in low-coupling score raises messages per sub-task from 3.33 to 3.85, a rise of 16 percent; in the high-coupling setting it goes from 4.33 to 6.53, a rise of 51 percent; in Hanabi teams spend 63 percent more signalling effort per point. Table 2 states this as ρC against ρ: the ratio computed on cost is two to three times the ratio computed on score in every setting. C is an efficiency measure and its denominator moves too, so those percentages are not counts. Splitting them: raw communication volume rises 14, 38 and 42 percent while progress falls 2, 8 and 13, and the two compound into the ratios above. Most of the effect is agents saying more rather than achieving less, which Table 3 corroborates from the message side.

4.2. The swap test Levels. Before reading the contrasts, the levels. On level 6, our high-coupling held-out set, Sun et al. (2025) report 60.7 progress completeness for Claude Sonnet 4 without formation; our formed teams reach 64.1 (Table 1). In Hanabi they score 15.8 of 25, against 13.3 for GPT-4-turbo (Agashe et al., 2025), 15 to 18 for the strongest reasoning models

The announcement is not free on its own. Intact minus Placebo is 0.9 and 0.9 progress-completeness points in the two Collab-Overcooked settings and nothing measurable in Hanabi, with coordination cost 4 to 6 percent higher. That 5

Testing Interchangeability in LLM Agent Teams Table 2. Decomposition, in normalised units. V is the value of experience (Equation (1)), Π the partner-specific residue (Equation (2)), ρ = Π/V , W/Π the share of the residue reproduced by deleting an intact team’s own partner notes (Equation (4)), Σ the gain from clearing an arriving agent’s notes (Equation (5)), and ρC the same ratio computed on coordination cost. ρ

W/Π

CO, low 0.085 0.015 0.18 CO, high 0.199 0.053 0.26 Hanabi 0.183 0.081 0.44

0.44 0.81 0.64

Setting

V

Π

Σ

Table 3. What the messages a swap adds are doing, as a share of the extra traffic relative to the same team’s placebo, in percent. Collab-Overcooked only, since Hanabi has no side channel.

ρC

+0.008 0.45 +0.016 0.72 +0.043 1.11

The extra message was

low high

request re-issued after no response clarification of a request unsolicited status report correction after a failed hand-off

33 25 23 18

21 20 21 38

why the same swap costs 51 percent more messages there and 16 percent at low coupling while the score barely moves in either: a repair is expensive in messages and usually still recovers the sub-task.

is small, but in the low-coupling setting it is of the same order as Π itself, which is why every contrast in Table 2 is taken against the placebo and not against the intact team. A swap test that skipped this control would report most of the disruption of a personnel change as though it were the identity of the person.

Coupling, and where the residue sits. ρ is larger where the task demands more coordination: 0.18, 0.26, 0.44. Resampling the four swap pairs puts these at 0.14–0.21, 0.24– 0.30 and 0.41–0.47; four independent units make a coarse interval, so it orders the three settings rather than fixing any of them. The two Collab-Overcooked rows carry most of the weight, since they hold the benchmark, the roles, the prompts and the metric fixed; they do not hold task complexity fixed, so this is an association along a gradient rather than the effect of coupling alone; separating them would need levels matched on complexity and varied only in required hand-offs. Where agents can divide the work and execute in parallel they form little that is specific to the partner; where they must interleave, they form more. In the low-coupling setting Π = 0.015, which at this cohort size is close to the noise floor, so that cell bounds the residue well below the value of experience rather than measuring it.

The Hanabi row is worth pausing on. There ρC = 1.11: a swapped agent costs more coordination effort than one with no experience at all. One interpretation, consistent with Σ, is that a fresh agent has fewer partner-specific expectations, whereas an experienced agent from another team arrives with conventions formed elsewhere; the pair then spends messages detecting and repairing mismatches that were never announced. Deleting the arriving agent’s partner notes helps wherever coupling is non-trivial, by 0.016 and 0.043 in normalised score, which is the same statement from the other side. This sits beside the closest cross-play result in the literature. Agashe et al. (2025) find a GPT-4-turbo Hanabi agent paired with an unfamiliar partner scoring as well as in self-play. Our score column agrees, and their setting could not have shown ours: without memory across games their agent could not accumulate a persistent partner-specific protocol across games. The penalty we measure appears once agents are given persistent memory, and even then it is small.

W says where the residue sits. In the two coupled settings, deleting an intact team’s own partner notes reproduces 81 and 64 percent of the swap penalty without changing the roster. A substantial share of the partner-specific part is therefore recoverable by manipulating the section we asked agents to label as partner notes. Two things keep that short of a clean localisation: deleting the section also shortens the context, and the κ/π split is one we induced with a heading, so an agent is free to file partner-specific material under task notes. In the low-coupling setting W/Π is 0.44, but both quantities there are small enough that the ratio is not informative.

What the extra messages are. The cost metric counts messages, so we can ask what the additional ones were doing. Coding the messages a swapped team sends beyond what its own placebo sends (Table 3), the two CollabOvercooked settings differ in a way that matches the rest of the picture. At low coupling a third of the extra traffic is a request re-issued because the first drew no response, and a quarter is clarification: the arriving agent asks for things the incumbent used to volunteer. Little of it follows an actual failure, because at this coupling a missed hand-off is usually recoverable inside the same sub-task.

Protocol signatures give a second view. Mean pairwise divergence between teams in the same cohort after formation is 0.33, 0.41 and 0.36 on a scale where 1 would mean no shared structure at all: independently formed teams of the same model end up with recognisably similar protocols, which is one reason the residue is small. Divergence is not, however, what orders these three settings. Hanabi has the largest ρ and only the middle d, so coupling and divergence are two levers rather than one. Section 4.5 moves the second while holding the task fixed.

At high coupling the largest share, 38 percent, is correction after a hand-off has already failed. That is what a swap changes when the work has to interleave: the pair does not find the mismatch by asking about it, they find it by acting on incompatible expectations and then repairing. It is also 6

Testing Interchangeability in LLM Agent Teams Table 5. Seat-targeted swaps in Collab-Overcooked. Π is the drop in progress completeness relative to the placebo. The last column gives the share of the extra messages sent by the incumbent rather than the newcomer, for an initiator swap and a responder swap.

Table 4. Share of partner-note sentences by category, in percent. Columns are the two Collab-Overcooked settings and Hanabi; columns may not sum to 100 because of rounding. Partner-note category

low high Hanabi

acts without being asked needs to be told explicitly agreed form of a signal timing of hand-offs recurring mistakes

34 26 14 16 10

30 23 20 17 11

Π (PC points)

24 20 28 19 9

Coupling initiator responder ratio

inc. share

low high

68% / 46% 70% / 42%

2.02 7.54

0.96 2.55

2.1 3.0

donor team’s. The movement is uneven. With the newcomer in the lead seat ∆ is 0.060, 0.149 and 0.112 across the three settings, against 0.020, 0.056 and 0.069 when it takes the second seat, so the lead seat carries 2.6 to 3.0 times as much in Collab-Overcooked and 1.6 times in Hanabi, whose seats are close to symmetric. The seat that speaks first has more influence over the terms.

What the partner notes contain. Because the residue is text, we can read it. Coding every sentence written under the partner heading into five categories (Table 4), the largest share everywhere describes what the partner does unprompted and what has to be spelled out for them. The category that shifts most with coupling is the agreed form of a signal, from 14 percent at low coupling to 28 percent in Hanabi, and it is the one whose content is arbitrary: nothing fixes which phrasing marks a hand-off or what a colour hint means. The same relation holds within a setting. Standardising within each cohort and pooling the 24 teams, the share of a team’s partner notes given to the form of a signal correlates with its swap penalty at 0.46: teams with more notes about partner habits tended to lose little when the partner changed, while teams with more notes about signal conventions tended to lose more.

The magnitudes are modest, with ∆ peaking at 0.15 on a divergence scale where 1 would mean the protocol was replaced outright, but the direction is the point. The observed signatures shift in the direction associated with the moved agents, especially when the moved agent occupies the lead seat. This is the everyday form of a route studied elsewhere in its adversarial form, where one compromised agent seeds behaviour across a network without stating it (Weckbecker et al., 2026).

4.3. Which seat carries the loss

4.4. Recovery

Collab-Overcooked separates initiating from responding (Sun et al., 2025), which lets us ask not only whether a seat is replaceable but which one. Table 5 shows the loss is uneven. In the high-coupling setting, swapping the initiator, the agent that holds the recipe and drives the plan, costs 7.5 points of progress completeness against 2.5 for the responder, a ratio of three. The low-coupling ratio is 2.1 in the same direction, on penalties small enough that we read the ordering rather than the number.

Figure 2 splits the ten-episode averages of Table 1 into the shock at episode one and the rate at which it decays. The shock is the larger part. In the first episode after the change a team scores 98, 90 and 89 percent of its placebo baseline and pays 1.15, 1.48 and 1.55 times the placebo coordination cost; the averages in Table 1 are smaller than this because recovery is already under way inside them. Task score comes back within one percent of the placebo baseline by episode two in the low-coupling setting, episode five in the high-coupling one, and episode six in Hanabi. Coordination cost decays about twice as slowly, and by the same one-percent rule it returns within ten episodes only in the low-coupling setting; in Hanabi teams are still paying a 7 percent premium at episode ten.

The last column decomposes the added communication by who emitted it. When the initiator changes, about seven tenths of the extra messages come from the incumbent rather than the newcomer: the agent that stayed re-explains, reconfirms and reissues requests that used to be implicit, while the newcomer, holding a complete and internally consistent protocol of its own, largely proceeds as before. Swapping the responder removes the asymmetry, with the two agents contributing about equally. When the seat that sets the agenda changes, then, the visible symptom shows up mostly on the agent that did not.

This is the same gap between outcome and cost, now in time: a reshuffled team looks recovered on any dashboard tracking success rate several episodes before it has returned to its prior coordination efficiency. 4.5. Ablations: model, temperature, team age Three manipulations ask what ρ responds to. All run in the high-coupling Collab-Overcooked setting, where Π is large enough to resolve, with six teams per configuration and three conditions (Placebo, Swap, Naive) at ten episodes

Protocol signatures shift after a swap. A complementary signal comes from how protocol signatures change after a swap. ∆ from Equation (6) is positive in every setting: after a swap the receiving team’s signature moves toward the 7

Testing Interchangeability in LLM Agent Teams T / placebo, % (solid)

C / placebo (dashed)

100

1.6

95

1.4

90

1.2

85

1

3

5

7

episode after the roster change

Overcooked, low

Overcooked, high

9

Table 6. Ablations, all in the high-coupling Collab-Overcooked setting with six teams per configuration. Placebo PC is the reference level; V , Π and ρ are as in Table 2; d is the mean pairwise divergence between the protocol signatures of teams in the same cohort. The row marked † is the reference configuration, taken from the first six teams of the main cohort rather than re-run, and shared by all three blocks.

1.0

placebo PC Base model GPT-5.6 Luna† Gemini 3.7 Flash Claude Sonnet 5 Temperature 0.0 0.7† 1.0 1.5 2.0 Formation episodes E 5 10† 20

Hanabi

Figure 2. Recovery after the roster change. Task score (solid, left axis) comes back within one percent of the placebo baseline by episode two, five and six respectively; coordination cost (dashed, right axis) decays about twice as slowly and has returned by episode ten in neither of the coupled settings.

each (Table 6). Eight new configurations at 6 × (E + 30) episodes come to 1,950; the ninth cell, Luna at temperature 0.7 with E = 10, is not re-run but taken from six of the eight teams of the main cohort, which is why its row differs a little from Table 2. Resampling the three swap pairs puts the spread of ρ at about ±0.03 in these cells and ±0.09 in the one where V has collapsed: enough to order a block, not enough to separate adjacent rows inside one.

V

Π

ρ

d

63.7 54.9 73.3

0.195 0.053 0.27 0.41 0.214 0.093 0.43 0.52 0.176 0.034 0.20 0.33

66.2 63.7 60.8 53.0 34.7

0.195 0.195 0.169 0.212 0.064

60.6 63.7 68.6

0.184 0.030 0.17 0.36 0.195 0.053 0.27 0.41 0.202 0.083 0.41 0.51

0.028 0.053 0.052 0.080 0.011

0.14 0.27 0.31 0.38 0.17

0.27 0.41 0.54 0.70 0.80

coding halves the swap penalty, ρ = 0.27 to 0.14, while progress completeness rises from 63.7 to 66.2. This suggests decoding as one simple lever for systems that expect to rotate agents.

Base model. Partner specificity varies by a factor of two: ρ = 0.27 for GPT-5.6 Luna, 0.43 for Gemini 3.7 Flash, 0.20 for Claude Sonnet 5. The ordering runs against competence rather than with it. Sonnet 5 is the strongest of the three, at 73.3 progress completeness against 63.7 and 54.9, and is also the easiest to swap. What the ordering does follow is d, at 0.41, 0.52 and 0.33. On three models that is a rank agreement rather than a demonstration, but it is the direction the other two manipulations take as well.

Formation length. With E ∈ {5, 10, 20}, ρ goes 0.17, 0.27, 0.41. The two terms behind it move differently. V is 0.184, 0.195, 0.202, a rise smaller than its own spread, while Π nearly triples, 0.030, 0.053, 0.083. Whatever a team gains between its tenth and its twentieth episode, an equally experienced stranger does not have it. Ten episodes is not a ceiling. The trend is preliminary, but it is the clearest indication in our ablations that partner specificity may grow with team history.

Decoding temperature. On Luna, d rises from 0.27 at greedy decoding to 0.80 at temperature 2.0, and ρ follows it as far as 1.5 (Figure 3). Both ends of that range need a word. At temperature 0 the agents are deterministic, so the only thing left to separate two teams is the order in which they met the formation tasks; that is enough to leave d = 0.27, and it leaves almost nothing partner-specific, ρ = 0.14, at the highest task score in the block. At temperature 2.0, ρ falls back to 0.17, not because the teams have grown alike, since d is highest there, but because V has collapsed to 0.064: a growing share of messages no longer parses, the teams stop learning the task, and a ratio whose denominator is near zero stops meaning anything. These results are consistent with sampling entropy creating more room for private conventions: up to the point where high temperature also degrades the task, more of it is associated with a more expensive swap.

What the three share. Across the nine configurations ρ and d correlate at 0.37; dropping the one where V has collapsed leaves 0.83 over the remaining eight. The manipulations change different things, a model’s priors, its sampling entropy, its amount of practice, and each moves ρ in step with how far two teams that never met end up from each other. On eight points, four from one block, that is not a single mechanism, and it does not extend to the coupling gradient of Section 4.2, where ρ and d do not move together. Divergence and coupling are two ways to make a team hard to reshuffle; these manipulations move the first. The ablations run three of the six conditions, so they cannot say whether the residue still sits in the partner notes on another model or at twenty episodes.

In this setting, moving from temperature 0.7 to greedy de8

Testing Interchangeability in LLM Agent Teams d, , V

signature shifts toward the donor team’s. Partner specificity here is measurable, modest and visible in text, and it is not fixed: it is larger where the task demands more coordination, where the team is older, and where the decoder leaves more freedom to invent.

protocol divergence d partner specificity value of experience V

0.8 0.6 0.4 0.2 0.0

7. Limitations and Future Work 0.0

0.7

1.0

decoding temperature

1.5

Everything here is dyadic, which is where the benchmarks are but not where the organisational questions are. Ten formation episodes is short, and Section 4.5 shows ρ still rising at twenty, so the pattern is established only over the formation horizons we study. Our notebook design makes the split between task and partner notes explicit, which may be generous to the phenomenon: a system that did not ask for partner notes might show less. Deleting a section also shortens the context, so W and Σ would be cleaner against length-matched and neutral-text controls, which we did not run. Protocol signatures are surface features, so two teams could share one and still differ. The relation between ρ and d is exploratory: eight points, four of them from one manipulation.

2.0

Figure 3. Decoding temperature on GPT-5.6 Luna, in the highcoupling setting. Divergence between independently formed teams rises throughout; partner specificity follows it until temperature 2.0, where the value of experience collapses and the ratio stops being interpretable. All three quantities are dimensionless but of different construction, and share an axis only for compactness.

5. Discussion In our settings, repeatedly interacting LLM agent teams develop behavior that is partly specific to their partner, but the partner-specific component is modest. Most of what a team gains from ten episodes together is knowledge about the task, which any equally experienced agent of the same model already has. The part tied to the pairing grows with how far the task forces the agents to interleave, stays below half the value of experience, and sits in a readable paragraph. A consistent interpretation of Section 4.5 is that independently formed teams of the same model draw on similar pretrained expectations about cooperative exchange, while additional freedom during formation allows more partnerspecific conventions to emerge. Wang et al. (2026) reach the same place for reference games by another route.

Three things follow for the version of this study we would like to run. Carrying formation out to a hundred episodes would show whether ρ saturates or keeps climbing, clarifying how the effect scales in longer-lived teams. With three or more agents a team can route around a replaced member, and the residue stops belonging to a pair; whether it is then pairwise or collective is a question the dyadic design cannot pose. And the ablations point at an intervention they do not test: if the greedy-decoding result reflects reduced room for partner-specific conventions, then fixing a protocol in the prompt should provide a cleaner test by separating the decoder from the convention.

The practical version is a change in what to watch. In our settings, rotating agents for a failover, for load balancing or after a restart changes coordination cost more than task accuracy, so a success-rate dashboard can make a reshuffled team look recovered before its communication efficiency has recovered. A swap test therefore complements a competence test on the arriving agent, and clearing partner-specific notes is a simple intervention suggested by our ablations. In our temperature ablation, greedy decoding also reduces the swap penalty without reducing task score.

6. Conclusion Members of a long-running LLM agent team can largely be swapped for equally experienced strangers on the metric people watch, and less so on the one they do not. Task score moves little; communication per unit of progress rises 16 to 63 percent and recovers at half the speed; in CollabOvercooked the loss concentrates in the seat that initiates plans and most of the extra talking is done by the agent that stayed; and after a swap the receiving team’s protocol 9

Testing Interchangeability in LLM Agent Teams

References

the 37th International Conference on Machine Learning (ICML), pp. 4399–4410, 2020.

Agashe, S., Fan, Y., Reyna, A., and Wang, X. E. LLMCoordination: Evaluating and analyzing multi-agent coordination abilities in large language models. In Findings of the Association for Computational Linguistics: NAACL 2025, pp. 8053–8072, 2025.

Hu, H., Lerer, A., Cui, B., Pineau, J., and Foerster, J. Off-belief learning. In Proceedings of the 38th International Conference on Machine Learning (ICML), pp. 4369–4379, 2021.

Ashery, A. F., Aiello, L. M., and Baronchelli, A. Emergent social conventions and collective bias in LLM populations. Science Advances, 11(20):eadu9368, 2025.

Huckman, R. S., Staats, B. R., and Upton, D. M. Team familiarity, role experience, and performance: Evidence from Indian software services. Management Science, 55 (1):85–100, 2009.

Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., Dunning, I., Mourad, S., Larochelle, H., Bellemare, M. G., and Bowling, M. The Hanabi challenge: A new frontier for AI research. Artificial Intelligence, 280:103216, 2020.

Ko, T.-W. and Geiping, J. Attractor states emerge in multi-turn LLM conversations. arXiv preprint arXiv:2606.30571, 2026. Lewis, D. Convention: A Philosophical Study. Harvard University Press, Cambridge, MA, 1969.

Carroll, M., Shah, R., Ho, M. K., Griffiths, T. L., Seshia, S. A., Abbeel, P., and Dragan, A. On the utility of learning about humans for human-AI coordination. In Advances in Neural Information Processing Systems (NeurIPS), 2019.

Lewis, K. Measuring transactive memory systems in the field: Scale development and validation. Journal of Applied Psychology, 88(4):587–604, 2003. Mirsky, R., Carlucho, I., Rahman, A., Fosong, E., Macke, W., Sridharan, M., Stone, P., and Albrecht, S. V. A survey of ad hoc teamwork research. In Proceedings of the 19th European Conference on Multi-Agent Systems (EUMAS), pp. 275–293, 2022.

Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., Chopra, B., Tiwari, R., Keutzer, K., Parameswaran, A., Klein, D., Ramchandran, K., Zaharia, M., Gonzalez, J. E., and Stoica, I. Why do multi-agent LLM systems fail? In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2025.

Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), 2023.

Chen, H., Song, X., Jin, J., Ren, P., and Zhang, L.-J. Toward an organizational science of multi-agent LLM systems: Decoupling who, how, and which algorithm. arXiv preprint arXiv:2607.25446, 2026.

Qian, C., Liu, W., Liu, H., Chen, N., Dang, Y., Li, J., Yang, C., Chen, W., Su, Y., Cong, X., Xu, J., Li, D., Liu, Z., and Sun, M. ChatDev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pp. 15174–15186, 2024.

Clark, H. H. and Wilkes-Gibbs, D. Referring as a collaborative process. Cognition, 22(1):1–39, 1986. Espinosa, J. A., Slaughter, S. A., Kraut, R. E., and Herbsleb, J. D. Familiarity, complexity, and team performance in geographically distributed software development. Organization Science, 18(4):613–630, 2007.

Ramesh, M., Jayakumar, K., Ramkumar, A., Thodima, P., Rege, A., and Vlatakis-Gkaragkounis, E.-V. Sparks of cooperative reasoning: LLMs as strategic Hanabi agents. In Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026.

Goel, H., Ellendula, A. S., Tadiparthi, V., Moradi Pari, E., Nourkhiz Mahjoub, H., and Chinchali, S. P. Bayesian partner modelling enables adaptive replanning for LLM coordination. arXiv preprint arXiv:2608.18490, 2026. Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., and Schmidhuber, J. MetaGPT: Meta programming for a multi-agent collaborative framework. In International Conference on Learning Representations (ICLR), 2024.

Reagans, R., Argote, L., and Brooks, D. Individual experience and experience working together: Predicting learning rates from knowing who knows what and knowing how to work together. Management Science, 51(6): 869–881, 2005. Riedl, C. Emergent coordination in multi-agent language models. In International Conference on Learning Representations (ICLR), 2026.

Hu, H., Lerer, A., Peysakhovich, A., and Foerster, J. “otherplay” for zero-shot coordination. In Proceedings of 10

Testing Interchangeability in LLM Agent Teams

Shih, A., Sawhney, A., Kondic, J., Ermon, S., and Sadigh, D. On the critical role of conventions in adaptive human-AI collaboration. In International Conference on Learning Representations (ICLR), 2021.

Meeting of the Association for Computational Linguistics (ACL), pp. 8580–8622, 2025.

Stone, P., Kaminka, G. A., Kraus, S., and Rosenschein, J. S. Ad hoc autonomous agent teams: Collaboration without pre-coordination. In Proceedings of the 24th AAAI Conference on Artificial Intelligence (AAAI), pp. 1504–1509, 2010. Strouse, D., McKee, K. R., Botvinick, M., Hughes, E., and Everett, R. Collaborating with humans without human data. In Advances in Neural Information Processing Systems (NeurIPS), 2021. Sun, H., Zhang, S., Niu, L., Ren, L., Xu, H., Fu, H., Zhao, F., Yuan, C., and Wang, X. Collab-Overcooked: Benchmarking and evaluating large language models as collaborative agents. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 4922–4951, 2025. Treutlein, J., Dennis, M., Oesterheld, C., and Foerster, J. A new formalism, method and open issues for zero-shot coordination. In Proceedings of the 38th International Conference on Machine Learning (ICML), pp. 10413– 10423, 2021. Wang, P.-Y. A., Mishra, C., Özyürek, A., Rubio-Fernández, P., and Ghaleb, E. Aligned but not partner-specific: Distinguishing how multimodal LLM agents succeed in reference games without human-like conventions. arXiv preprint arXiv:2606.08081, 2026. Weckbecker, M., Müller, J., Hagag, B., and Mulet, M. Thought virus: Viral misalignment via subliminal prompting in multi-agent systems. arXiv preprint arXiv:2603.00131, 2026. Wegner, D. M. Transactive memory: A contemporary analysis of the group mind. In Theories of Group Behavior, pp. 185–208. Springer, 1987. Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., and Wang, C. AutoGen: Enabling nextgen LLM applications via multi-agent conversations. In Conference on Language Modeling (COLM), 2024. Xu, W., Liang, Z., Mei, K., Gao, H., Tan, J., and Zhang, Y. A-mem: Agentic memory for LLM agents. arXiv preprint arXiv:2502.12110, 2025. Zhu, K., Du, H., Hong, Z., Yang, X., Guo, S., Wang, Z., Wang, Z., Qian, C., Tang, X., Ji, H., and You, J. MultiAgentBench: Evaluating the collaboration and competition of LLM agents. In Proceedings of the 63rd Annual 11

Record · ID 660857 · SHA-256 192ab097c37bdd91
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.