Conceptio › Archive › arXiv CS
arXiv CSopen access

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence Eshika Pathak∗† and Leela Krishna†

arXiv:2609.21942v1 [cs.RO] 18 Sep 2026

∗ University of Illinois Urbana-Champaign

Abstract—A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. We study when initiating dialogue is necessary relative to the evidence the acting model actually has. Choosing well requires two things current systems lack: knowing how much the robot’s sensors reveal about the cause, and knowing how reliable the robot’s own diagnosis is. We build a simulated benchmark in which every failure’s true cause is known, because we injected it, and measure what each sensor reveals, with explicit checks against data leakage. Some failures are diagnosable from camera images; others only from the robot’s force data (0.99 from force data, no image method above 0.55). We then test six open vision-language models. Their behavior tracks the surface of the prompt, not the evidence: moving the refusal option from last to first in the answer list collapses refusal rates from 78–100% to 0–6% in three of the six swept modeland-family pairs. Both Cosmos generations’ refusal on grasp failures survives every variant. Accuracy from frames stays at or below a majority-class baseline under every prompt variant, with or without worked examples, and stated confidence carries no information about correctness. Handing the same models the force data as ten lines of text produces the study’s first abovebaseline diagnoses, in four of the six models: much of the failure reflects missing sensor data, not missing ability. We therefore pose the choice as a three-action decision problem, act, consult your own sensors, or ask a human, whose optimal policy follows from measured accuracy. The models do not follow it, and their ask rates ignore a fourfold change in question cost. One question to a human still lifts them from that baseline to roughly the answerer’s own reliability (0.70–0.81 when they ask). Today, the decision to ask should be wired to measured accuracy and stated costs, not to the model’s confidence.

I. I NTRODUCTION Corrective human-robot dialogue begins before either party has said anything. When a robot working alongside a person fails [1], its first communicative decision is whether to keep acting on its own diagnosis or interrupt the task and ask. A question costs attention, delays the task, and may change whether the physical state is still recoverable. Detecting that something went wrong is well studied [2], and repair is straightforward once the cause is known; the unexamined step sits between them, and it is a dialogue decision: when should the robot stop acting and start talking? By dialogue we mean situated corrective guidance after a failed attempt, in which the robot may ask what went ∗ Work done while at Centific. Email: [email protected]

† Centific

wrong, confirm a suspected cause, or request a correction, and the person may answer, demonstrate, or take the arm. Our experiments isolate the first move of that exchange, whether to initiate it at all, and the comprehension step that follows. The literature holds two opposite assumptions about this diagnosis stage. One line of work names the cause of a failure from the robot’s observations, either by reasoning over textual summaries of an episode without training [3] or by training on failures synthesized from successful demonstrations [4], [5], [6], with real-world failure benchmarks following [7]. What these share is the assumption that the cause is recoverable from the observations given; none measures whether it is. A second line builds mechanisms for asking humans for help [8], [9], [10], assuming that sometimes it is not; neither line checks which assumption holds for a given failure. We built a benchmark to check; the answer is that diagnosability depends on which failure it is and which sensor you consult. The study answers three questions in order. First, what can be known? We create failures with known causes and measure how well each sensor reveals the cause, with classifiers screened for data leakage (Section IV-B). Second, do visionlanguage models know it? We test whether six open models recover the recoverable causes and whether their confidence reflects what can be known (Section VI-B). Third, do they choose well? We give the models the act-or-ask choice under explicit costs and compare their choices to the mathematically best policy (Section VI-D). One quantity links the three: whether to ask depends on p, the probability that the robot’s own diagnosis is correct (Section III-C). The audit bounds p by what the sensor data contain, the diagnosis test measures what each model extracts, and with both measured the best choice is computable and each model is scored against it. The sensor measurements were more specific than we expected. Detection failures (the robot cannot find the object it was asked to fetch) can be diagnosed from images, and the score survives changes of viewpoint and scene appearance. Grasp failures (the object slips out during a lift) can be diagnosed almost perfectly from the robot’s force data (wrist force-torque and gripper measurements), at 0.99, but not from its camera: a sequence of increasingly capable image classifiers, ending with a pretrained vision backbone and a video model trained end to end, levels off at 0.55. The cause of a grasp failure is recorded inside the robot and largely invisible

from outside, so a frames-only diagnoser cannot reliably tell what happened even though the robot’s own sensors could settle it; whether human observers do better is untested here. A third family, placement failures, shows the pitfall of this kind of measurement. Its image classifier scores 0.998, but the score does not come from evidence about the failure: one injected cause moves the container, so the classifier succeeds by learning where the container sits in the frame, and from camera angles held out of training its accuracy falls to 0.605. We keep this family as the concrete case for why no score is believed here until it survives such tests (Section IV-B). We also state plainly which part of the grasp-versus-detection contrast follows from our own construction: grasp causes are physical parameters the renderer never draws (grip force, object mass, surface friction), while detection causes are visible properties of the scene (an occluder, an absent object, dim lighting), so images were always going to reveal more about detection failures than about grasp failures. What the construction does not determine, the experiments do: the actual amount of information each sensor carries, and whether a given high score is genuine, which the detection score proves by surviving the held-out tests and the placement score fails. The six models fail in a way we did not anticipate. At first sight they split into two camps: three refuse to diagnose nearly everything, three commit to a cause on nearly every episode, and both camps score at or below a majority-class baseline, which always guesses the most common cause. A robustness sweep dissolves the split. Moving the “cannot be determined” option from the last position in the list to the first collapses the refusals almost entirely in three of the six model-and-family pairs we swept, and one model then picks whatever option sits second on 33 of 35 episodes regardless of content. The models are not cautious or reckless; they are answering the layout of the prompt. Accuracy from frames stays at or below the majority-class baseline under every variant, and stated confidence carries no information in any of them. One behavior survives the sweep: the refusal on grasp failures persists across all six variants for both Cosmos models. What no model does, under any prompt we tried, is behave differently between failures whose cause is demonstrably recoverable from its input and failures whose cause is not. Worked examples in the prompt move no model toward what the images verifiably contain, so the failure is not a zero-shot artifact. Handing the models the robot’s force data as a few lines of text does move most of them: four of the six then diagnose above the majority-class baseline, the first above-baseline results in this study. And when the models ask a human, the single answer does almost all the work: the four models with decision-task ask data rise from at or below the majority-class baseline to 0.70–0.81 when they ask, roughly the answerer’s 0.80 reliability passing through. This paper contributes: 1) Audited measurements of which failures are diagnosable from which sensors, including one family where the camera is nearly useless and force data nearly sufficient (Section IV-B).

2) Evidence from six open 7B–16B vision-language models that diagnosis behavior follows prompt layout rather than evidence, with accuracy from frames at or below a majority-class baseline under every variant and no informative confidence channel but one (Section VI-B). 3) A way to decide among acting, sensing, and asking: explicit costs plus measured accuracy give an optimal policy, and a model’s distance from it is reported as wasted cost (Section III-C). 4) The measured value and misuse of asking: one question lifts models from at or below the majority-class baseline to roughly the answerer’s reliability (0.70–0.81 when they ask), yet no ask rate tracks the cost of asking (Section VI-D). II. R ELATED W ORK Diagnosing failures. Systems name the cause of a failure from robot observations, by prompting a language model over textual episode summaries [3], by training on synthesized failures [4], [5], [6], [11], or from real-world trajectories [7]. All assume the cause is recoverable from the observations given, and none verifies that assumption against data leakage; our third family shows what verification catches. Corrective dialogue and asking for help. A robot’s turn in a shared task updates both the conversation and the physical state, so managing it involves grounding, clarification, and repair [12], [13]. Prior work generates help requests by modeling how a listener will interpret them [9], [14], recovers from faults through dialogue [15], and learns from natural-language corrections [16], [17], [18]; in the latter the human initiates, so the robot’s problem is comprehension rather than initiation. Ours inverts that, and we then measure whether the reply can be used. When to ask, and at what cost. KnowNo [8] decides when to ask using conformal prediction, giving a coverage guarantee rather than a cost-optimal policy; Ask When It Pays [19] prices questions in goal navigation from an information-gain analysis; ESearch-R1 [20] unifies asking, memory retrieval, and navigation into one cost-aware process and learns the policy by reinforcement learning. We instead compute the reference policy from audited sensor evidence, measured model accuracy, and stated costs, which lets us score off-the-shelf models against it rather than train one. Related benchmarks examine when software agents defer to a human [21], decouple question quality from navigation [22], and document models that rarely refuse to answer [23], [24]; because option order alone can produce apparent refusal [25], we test that directly. Our selective-asking comparison follows selective classification [26]. Appendix D expands on the closest systems. III. P ROBLEM S ETUP A. Failures with known causes For real failures the true cause is unobservable, so no ground truth exists to score a diagnosis against; we therefore generate failures whose causes are known by construction. Each episode samples a cause c from a family-specific set C = {c1 , . . . , ck },

samples a severity from a cause-specific range, and runs a scripted pick-and-place policy in a physics simulator [27] under that perturbation; failing episodes are retained, each labeled with its injected cause. Two observation channels are recorded per episode: the camera frames that the evaluated vision-language models receive, and the robot’s proprioceptive and force measurements (wrist force-torque, gripper aperture, kinematic status), referred to below as the force data, which the models never see in the frames-only conditions but which the classifiers of Section IV-B use. B. Measuring what a sensor reveals To measure how much a sensor’s data reveals about the cause, we train a small classifier to predict the injected cause from that data and treat its accuracy as a lower bound on the information present: if a simple classifier recovers the cause at 0.91, the information is demonstrably there. A high score, however, can reflect label leakage or shortcut features rather than genuine evidence about the failure. To prevent crediting a sensor with information it does not carry, we accept a score only after three checks, which together we call the audit. Shuffle check: retrain with the cause labels randomly scrambled; accuracy must fall to chance, otherwise the label is leaking into the data through some side channel. Transfer check: test on camera viewpoints and scene appearances held out of training; accuracy must survive, judged by a proportional rule (a drop larger than 25% of the margin above chance fails). A classifier that scores 0.99 but collapses from a camera angle it never saw has learned the scene, not the failure; Section IV-B shows precisely this in our placement family. Separation check: training and test episodes must come from different simulation runs, so near-duplicate episodes cannot straddle the split. A score that passes all three is a reliable lower bound, established by exhibiting the classifier, and no future model can lower it; we call it that sensor’s certificate for the family, and it appears as Cert. in the tables. The reverse claim, that a sensor does not contain the answer, is harder, because it quantifies over every possible classifier. We support such claims only by running a sequence of checked classifiers of increasingly capable designs until their accuracies level off, and we state the result as the plateau of that sequence, never as a proven ceiling. C. The act-or-ask decision After a failure, an agent either acts, executing the repair matched to its diagnosis, or asks a human one question and then acts. Each cause has one scripted repair (for a weak grip, squeeze harder and retry; for a heavy object, change the lift), so the diagnosis selects among prepared repairs and a wrong diagnosis triggers the wrong repair. Costs are stated in units of one retry: a correct repair costs Cr =1, a wrong repair costs Cw =10 (execute it, undo it, start over), and a question costs Ca =3 (a person’s interrupted attention). The human’s answer is imperfect, for reasons given in Section IV-C: correct with probability qc =0.80, uninformative with qv =0.15, wrong with qw =0.05. Let p be the probability that the agent’s diagnosis

on an episode is correct. The expected costs of the two choices are EVact (p) = p Cr + (1 − p)(Cw + Cr ),

(1)

EVask (p) = Ca + qc Cr + qv EVact (p) + qw (Cw + Cr ), (2) and setting them equal gives the threshold p⋆ =

(1 − qv )(Cw + Cr ) − Ca − qc Cr − qw (Cw + Cr ) , (3) (1 − qv ) Cw

below which asking is the better choice. At the default costs p⋆ ≈ 0.588. A robot whose diagnosis is right 40% of the time expects cost 11 − 10(0.4) = 7.0 by acting and 6−1.5(0.4) = 5.4 by asking, so it should ask; at 80% accuracy the numbers are 3.0 and 4.8 and it should act. Sweeping Ca over {1, 3, 6, 10} moves p⋆ over {0.82, 0.59, 0.24, never ask}, so a rational agent’s ask rate must move with the cost. A third action, consulting an onboard sensor at cost Csense (Section V), is evaluated by the same comparison. The threshold itself is classical, the value of information in the sense of Howard [28] and the metareasoning tradition [29]. Our contribution is not the formula but its two inputs, which no prior failure benchmark supplies: audited measurements of how much each sensor reveals about each failure type (Section IV-B), and measured diagnosis accuracies for each evaluated model (Section VI-B), so that the optimal policy is computable and each model’s behavior can be scored against it. Note that p is a property of the agent, not of the world: on a diagnosable failure a strong classifier exists, yet an agent whose own accuracy is low should still ask, because the decision depends on the agent’s accuracy, not on the best achievable one. We therefore measure each model’s accuracy a per failure type on a held-out calibration set and compare four policies: the best fixed policy, which knows only a and asks on every episode exactly when a < p⋆ ; a confidence policy, which asks whenever the model’s confidence falls below a swept threshold [26] and can win only if confidence predicts the model’s errors; an evidence policy (grasp failures only), which asks based on per-episode decidability (Section IV-B) and so separates knowing the evidence from knowing oneself; and an oracle policy, which asks exactly on the model’s errors, uses ground truth, and serves only as an upper bound on the value of self-knowledge. A model’s regret, its realized cost minus that of the best fixed policy (hereafter the reference policy), weights mistakes by their consequences and compares across models, failure types, and costs. The apparent circularity of a decision rule that uses measured accuracy is addressed in Appendix A. IV. B ENCHMARK AND AUDIT A. Three kinds of failure, and why these three Manipulation failures arise at several stages of a task: perception (the target is not found or is misidentified), physical interaction (the grasp or transport fails), and task preconditions (the goal state cannot be reached), alongside classes out of scope here (planning errors, hardware faults, mis-specified

instructions). We construct one family per covered stage, chosen so that the plausible location of diagnostic information differs across them: pre-action scene appearance for perception, contact-time forces for interaction, and scene state for preconditions. This choice lets the benchmark identify which sensor carries the diagnostic information, not only whether a model recovers it. All episodes are tabletop pick-and-place scenes in the LIBERO simulator [27]: a simulated seven-degree-of-freedom arm must grasp a named object from the table and carry it to a goal location, observed by an external camera (example frames: Fig. 4). The families are referred to as A, B, and C in tables. Each is a mixture of physical causes with sampled severities; the families were constructed under the requirement that the scripted policy succeed on at least 95% of unperturbed episodes, so failures are attributable to the injected cause. Each family contains 2,200 failure episodes (2,000 in the audit pool and 200 held out for model evaluation), with table texture, lighting, and camera angle randomized and logged so the transfer check can hold them out. Grasp failures (A): the object leaves the gripper during a lift, because grip force was scaled down by U (0.2, 0.8), or object mass was scaled up by U (1.5, 4.0), or surface friction was reduced to U (0.05, 0.30). Detection failures (B): the robot cannot find the object it was asked to fetch, because the target is occluded 40–95%, or absent, or the instruction uses a name missing from the robot’s label set, or the lighting is dimmed to U (0.1, 0.4). Placement failures (C): placing the held object fails, because the container is closed, or 70–100% full, or moved out of reach.

discontinuity. No feature reads simulator object pose or contact flags, and restricting to force-torque and gripper alone gives the same 0.986, so the score rides on the physics. It passes every check. That includes severity-band transfer, in which the classifier is trained only on mild perturbations and tested only on severe ones, and the reverse, scoring 0.95 in the worse direction; and a stricter variant in which evaluation groups are matched on perturbation size, so that severity alone cannot reveal the cause: within these groups the causes remain separable at 0.97, against a 0.35 chance level. The cause of a grasp failure is recorded in the force data and mostly missing from the camera. The image classifier also provides a per-episode measure of ambiguity: the probability it assigns to its most likely cause, which we call the episode’s decidability. Raw classifier probabilities are typically overconfident, so we rescale them by temperature scaling [30] on held-out episodes, until the stated probability matches the observed frequency of being correct. After this correction, decidability averages 0.574, close to the classifier’s 0.545 accuracy; for a calibrated measure, the average stated probability must match the frequency of being right. Placement failures show the pitfall. Their classifier scores 0.998, yet accuracy falls to 0.605 from held-out camera angles: one cause moves the container, so the classifier succeeds by learning where the container sits in the frame, which is exactly what a viewpoint change alters. We report it with its held-out numbers and exclude it from the claims above; why we accept the force-data score while rejecting this one is discussed in Appendix A.

B. What each sensor reveals

C. The human’s answers

The image classifiers range from gradient-boosted trees over PCA-compressed frames to frozen DINOv2 embeddings under a boosted head and a small 3D-convolutional network; the force-data classifier uses summary statistics of the force-torque and gripper traces. Detection failures can be diagnosed from images. A classifier over pixel features reaches 0.910 and passes every check: shuffled labels fall to chance; accuracy holds at 0.844 on held-out appearances and 0.813 on held-out viewpoints (Table I). A single pre-attempt frame already gives 0.912: an occluder, an empty table, and dim lighting are visible before the robot moves. Detection failure is a scene-understanding problem that ends in a failure. Grasp failures can be diagnosed from force data, not from images. From images, the classifier sequence of Section III-B plateaus (Table VI, appendix): snapshots reach 0.442 but fail the viewpoint check (0.345, chance); adding motion gives 0.503; a frozen pretrained backbone reaches 0.545 [0.522, 0.567]; a video model trained end to end adds nothing (0.531). All but the snapshot pass the checks, and a pre-attempt frame scores at chance (0.366), as it must: grip force, weight, and friction are invisible until contact. From the force data, a classifier reaches 0.986 using only signals a physical robot publishes: wrist force-torque, gripper aperture and force, and slip onset computed from the force

When a model asks a question, a scripted stand-in for the human replies. It knows the injected cause, rolls a die seeded per episode (all models face identical draws), and returns one short sentence: with probability 0.80, a template naming the true cause (“that one is heavier than it looks”); with probability 0.15, an unhelpful reply (“hard to say from here”); with probability 0.05, the template for a wrong cause. Replies do not depend on question wording, which is recorded and described in Section VI-E but has no effect in this version. The imperfection is deliberate: a perfect answerer would build the pro-dialogue conclusion into the apparatus and assume away the possibility of inaccurate human feedback. With a noisy answerer, asking carries real risk, and reaching the answerer’s own 0.80 reliability means the human’s knowledge passed through intact. V. E XPERIMENTAL P ROTOCOL Models. Six open vision-language models: Qwen2.5-VL7B-Instruct [31], Cosmos-Reason2-8B, Cosmos3-Nano [32], InternVL3-8B, Pixtral-12B, and Llama-3.2-11B-VisionInstruct. Cosmos3-Nano’s reasoning tower is built on a Qwen3-VL backbone, so it is not architecturally independent of the Qwen entry. Coverage differs by model (the tests are defined below). Qwen2.5-VL, Cosmos-Reason2, and Cosmos3-Nano ran every test, including all three prompt

versions of the decision test and the option-position sweep; InternVL3, Pixtral, and Llama ran both diagnosis rounds and the decision test in its costs-given (swept over question cost) and rule-given versions. Qwen2.5-VL and Cosmos-Reason2, the two primary models, additionally carry the detailed confidence and question-form analyses. A smaller relative, Qwen2-VL-2B, ran the same protocol and is reported in Appendix H. Llama required an isolated serving environment and two footnoted accommodations: four input frames where eight exceeded memory, and no few-shot condition (architectureinfeasible). All models received identical prompts verbatim (Appendix G); unparseable responses are scored as errors. Diagnosis test. Each model sees eight frames from the episode, the task instruction, and a statement that the attempt failed, then answers a multiple-choice question over that failure type’s causes plus an explicit “cannot be determined” option. The models receive frames at least as informative as those the classifiers used (same episodes, equal or higher resolution), so the certificates are a fair yardstick for what the models’ input contains; we do not expect the models to match them, only their confidence to reflect the gap. Confidence is recorded three ways: a stated 0–100 rating; the probability the model itself assigned to its chosen option, read from its output distribution; and how often the same answer recurs across five runs sampled at temperature 0.7, the parameter that scales the output distribution before sampling (the committed answer itself is generated deterministically, and the repeated-run readout requires a positive temperature). Because the first round produced constant confidence (Section VI-B), we ran a second round with the “cannot be determined” option removed, which we call the forced-answer round. This variant was registered in writing before it was run and is reported alongside the original, never instead of it. Every table reports three numbers: accuracy counting refusals as errors, accuracy on answered episodes only, and the answer rate; any one of these alone can mislead. Throughout, accuracies and scores are on the 0–1 scale, while percentages denote shares of episodes or of a gap. Two further input conditions probe where the bottleneck lies. In the telemetry conditions, the admissible force-torque and gripper channels are serialized as a fixed ten-line text summary (peak and mean forces, aperture trajectory, slip onset), the telemetry text, and given to the model in place of, or alongside, the frames. In the few-shot condition, three worked examples (frames plus correct answer), drawn only from calibration episodes, precede the question. Accuracy measurement. Evaluation episodes are split, balanced by cause and separated by simulation run, into a calibration set (25 per failure type) used only to measure each model’s accuracy a, and a test set (35–36) on which all decisions are scored. Decision test. On test episodes the model chooses ACT: <cause> or ASK: <one question>, and after any answer it must commit to a cause, whose scripted repair then runs. Three prompt versions separate different failures: plain (no costs mentioned), costs-given (costs stated, conclusion left

to the model), and rule-given (the model is told to ask exactly when it estimates its chance of being right is below the computed p⋆ ). The cost sweep Ca ∈ {1, 3, 6, 10} runs under the costs-given version. Policies and scoring follow Section III-C; the evidence policy uses per-episode decidability on grasp failures. Question texts are recorded and classified by form, with no scoring of quality. Every mean carries a bootstrap confidence interval (10,000 resamples), and paired comparisons use McNemar’s test [33], which compares paired conditions using only the episodes where they disagree. A three-action variant adds SENSE at cost Csense =0.5: the model receives the telemetry text, then must commit. Sensing is asking an instrument, with answer reliability equal to that model’s measured frames-plus-telemetry accuracy; the three-way thresholds follow from the same expected-cost comparison as Eq. (3). VI. R ESULTS A. What can be known: sensor audit Table I compares what the camera and the force data reveal, family by family, with placement failures shown alongside the held-out numbers that disqualify them; the full imageclassifier sequence appears in Table VI. The audit itself is in Section IV-B. B. Do the models know it: frames The models neither recover the recoverable causes nor reflect this in their confidence. No model clears the majorityclass baseline on grasp failures from frames in any round or prompt variant (best 0.343 against a 0.400 baseline), and none approaches the 0.910 certificate on detection failures (best overall 0.17, Table II). At the default prompt the models appear to split into refusers and overcommitters, but the split does not survive a robustness sweep. Moving the “cannot be determined” option across three phrasings and two positions collapses refusal from 100% to 3–6% and 78% to 0% for Qwen (grasp, detection) and 94% to 3% for CosmosReason2 (detection), while Cosmos3-Nano’s detection refusal only dips to 61–92%; with the refusal option first, Qwen selects whatever option sits second on 33 of 35 grasp episodes regardless of its content, consistent with the option-position bias documented in multiple-choice evaluation of language models [25]. Token-probability margins are decisive in both directions (medians +0.30 to +0.63 for refusing when the option is last, equally decisive avoidance when it is first): the models are confidently answering the position, not the question. One model-and-family pair is prompt-robust: both Cosmos models refuse on grasp failures at or near 100% across all six variants (Table VII), while their detection refusals differ sharply, the 8B collapsing and the smaller Nano barely moving (Appendix H). Accuracy is unchanged by any of this, at or below the majority-class baseline everywhere, so no prompt arrangement was concealing competence. The surviving claim is narrower and sharper: under every arrangement tried, no model’s willingness to answer or stated confidence responds to whether the cause is recoverable from

TABLE I W HAT EACH SENSOR REVEALS , BY FAMILY: GRASP CAUSES ARE RECOVERABLE FROM TELEMETRY, NOT FROM FRAMES . Fam.

Observation set

Acc.

Base.

Shuffle

Tint

Azim.

Sever.

A A A B C C

F/T + gripper + kinematics F/T + gripper only camera (2-frame snapshot) camera (2-frame snapshot) camera (2-frame snapshot) camera, contrast-normalized

0.986 0.986 0.442 0.910 0.998 1.000

0.347 0.347 0.347 0.315 0.397 0.397

0.339 0.337 0.338 0.267 0.360 0.352

0.974 0.974 0.445 0.844 0.926 0.881

– – 0.345 0.813 0.605 0.683

0.950 0.961 0.382 0.904 0.976 –

Families: A grasp, B detection, C placement. Base.: majority-class baseline. Slip onset in the telemetry sets is computed from the force discontinuity, not from object pose. Sever.: train on one severity band, test on the other (worst direction). For telemetry rows, Tint is a negative control, since table appearance cannot affect force traces, and camera angle does not apply. Family C’s near-perfect score fails the camera-angle check and is a cautionary example, not a result.

its input, and for most pairs the apparent epistemic stance is an artifact of option order. The two primary models ran the full protocol, so we examine them in detail. When Qwen answers on detection failures it scores 0.250, chance among four causes; Cosmos’s answered-only 1.000 rests on two episodes at a 6% answer rate. Stated confidence is a constant: Qwen says 50 on every grasp-failure episode and Cosmos says 100, including when its answer is “cannot be determined.” Spearman’s rank correlation [34] between confidence and per-episode decidability is ρ = −0.15 and −0.13 for the two models: no relationship. Two analyses close the last loophole, that refusing models knew the answer and refused to say it. In the forced-answer round, every refuser scores at or below the majority-class baseline (grasp: 0.31, baseline 0.400). And on the refused episodes themselves, the model’s top-ranked non-refusal option, read from the token probabilities it assigned while refusing, is no better than its forced-answer accuracy in any of the four model-and-family pairs (largest difference 0.07). What the models refused to say was at or below chance. The refusals were hiding nothing. Confidence in the forced-answer round is no longer constant, but it takes only a few distinct values and does not correlate with correctness (ρ ≈ 0). Figure 2 (appendix) plots all three confidence readouts against per-episode decidability for the primaries and Cosmos3-Nano: no readout is positively informative. On the paired episodes, McNemar’s test confirms the forcedanswer round differs from the first round on grasp failures for both primaries (eleven episodes changed from wrong to right and none the other way, p = 0.001) and for Cosmos on detection failures (p = 0.004); the primaries are statistically indistinguishable from each other in both rounds (p > 0.3), so we make no claim that either is better. Worked examples do not help. With three in-context exemplars, no model moves toward the certificate on either family; several drop, the best few-shot result anywhere is 0.333, majority-class picking, and refusal behavior is untouched (Qwen still refuses on every grasp episode). The failure on frames is a ceiling on what these models extract from images, not an artifact of zero-shot prompting. The answer format is not the obstacle either: given the scene fact as one sentence of text, the same models map it to the correct cause at 0.90–1.00, while free-form diagnosis without options falls to 0.08 and 0.00 (Appendix B). The gap is in perceiving the manipulation, not in reasoning from it.

C. Do the models know it: force data If frames do not carry the cause of a grasp failure, can the models use the sensor that does? The certificate of Section IV-B establishes that the cause of a grasp failure is recoverable from the force channels at 0.986. Serializing those channels as ten lines of text and giving them to the models produces the first above-baseline diagnoses in this study (Table III): four of the six models clear both the majority-class baseline and their own frames-only score from telemetry text alone (InternVL3 0.514, Pixtral 0.486, Llama 0.543, Cosmos3-Nano 0.457, against a 0.400 baseline and 0.31 frames-only). Qwen recovers only when the text accompanies the frames (0.286 alone, 0.514 fused). Cosmos-Reason2 does not recover under either condition (0.286 in both), a notable result for a model built for physical reasoning, while Cosmos3Nano does recover from the same text (0.457 alone, 0.486 fused). The two differ in backbone as well as in release, so we do not read this contrast as a generational or scaling effect. Adding frames to the telemetry text helps some models and hurts others: it is the difference between failing and recovering for Qwen, helps Cosmos3-Nano slightly, costs InternVL3 and Pixtral a few points, and collapses Llama’s best-in-study telemetry-only 0.543 back to its frames-only level. No model approaches the 0.986 certificate, the text summary being a lossy instrument, but the direction is unambiguous: much of what looked like inability to diagnose was information sitting in a sensor the models were never given, though one model cannot use that information even when it is given. This changes the dialogue reading of a grasp failure. It is not automatically a reason to interrupt someone: for four of the six models the cheaper move is to expose the robot’s own force data in a form the model can read, and a person is the right source only when no usable robot channel carries the answer. Partial competence also makes the confidence question testable for the first time: at zero accuracy, no confidence signal could have shown itself. At 0.46–0.54 accuracy, stated confidence remains uninformative and is mildly anti-calibrated (Pixtral verbalized ρ = −0.44; Llama samplefrequency ρ = −0.55): the models are more confident when wrong. The one informative channel anywhere is Pixtral’s token log-probability (ρ = +0.50), which also recovers 56% of the oracle selective-asking value, while its stated confidence recovers none (InternVL3’s token probabilities are excluded; see Appendix H).

TABLE II D IAGNOSIS ACCURACY OF THE SIX MODELS AGAINST THE CERTIFICATE . Model

Fam.

Cert.

Overall

Commit

Rate

Forced

Gap

Qwen2.5-VL-7B Qwen2.5-VL-7B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos3-Nano Cosmos3-Nano Llama-3.2-11B-V Llama-3.2-11B-V InternVL3-8B InternVL3-8B Pixtral-12B Pixtral-12B

A B A B A B A B A B A B

0.545 0.910 0.545 0.910 0.545 0.910 0.545 0.910 0.545 0.910 0.545 0.910

0.000 0.056 0.000 0.056 0.000 0.000 0.286 0.194 0.314 0.167 0.257 0.139

– 0.250 – 1.000 – – 0.286 0.194 0.314 0.176 0.290 0.143

0% 22% 0% 6% 0% 0% 100% 100% 100% 94% 89% 97%

0.314 0.167 0.314 0.306 0.314 0.389 0.314 0.278 0.314 0.194 0.343 0.417

0.545 0.854 0.545 0.854 0.545 0.910 0.259 0.716 0.231 0.743 0.288 0.771

Cert.: the certificate from Table I. Overall counts refusals as errors; Commit: accuracy on answered episodes, at the Rate shown; Forced: the pre-registered round with the refusal option removed; Gap = Cert. − Overall. Majority-class baselines on the test split: 0.400 (A), 0.306 (B).

TABLE III FAMILY-A DIAGNOSIS ACCURACY BY INPUT CONDITION ( FORCED - ANSWER ROUND ). F OUR OF THE SIX MODELS RECOVER THE CAUSE FROM TELEMETRY TEXT ALONE ; C OSMOS -R EASON 2 RECOVERS UNDER NEITHER CONDITION , WHILE ITS SMALLER , NEWER - GENERATION RELATIVE DOES . Model

Frames

Telemetry

Frames+telem.

Qwen2.5-VL-7B Cosmos-Reason2-8B Cosmos3-Nano InternVL3-8B Pixtral-12B Llama-3.2-11B-V

0.314 0.314 0.314 0.314 0.343 0.314

0.286 0.286 0.457 0.514 0.486 0.543

0.514 0.286 0.486 0.457 0.457 0.314a

Majority-class baseline 0.400; telemetry certificate 0.986. a Four input frames instead of eight (memory limit).

D. Do the models choose well: act or ask We first report what a question is worth, then compare the models’ choices with the reference policy of Section III-C. For the four models with decision-task ask data, one question brings final task accuracy to 0.70–0.81 when they ask (Qwen 0.76, Cosmos-Reason2 0.70, Llama 0.76, Cosmos3Nano 0.81), while unaided all four sit at or below the majorityclass baseline. The reply names the true cause 80% of the time, the model maps it onto a listed option, and the matching repair runs; the shortfall from 0.80 is the vague and wrong replies. Realized competence is almost entirely the answerer’s reliability passing through the model, which raises the question of whether the models deploy so valuable a resource sensibly. With forced-answer accuracy a ≈ 0.32 on grasp failures for the primaries (Wilson interval roughly ±0.10 at n=25), the reference policy asks at costs 1 and 3 and acts at 6 and 10; the break-even sits near 5.3 (4.6–6.6 across the estimate’s interval). Neither primary tracks the cost (Fig. 1). Qwen asks on 100% of episodes at every cost. While the reference policy asks (costs 1 and 3), Qwen’s regret is still slightly positive (0 to +1.3), because its commitments after unhelpful or wrong answers fall short of the reference policy’s; once the reference policy switches to acting Qwen keeps asking, at growing expense (+0.6 to +2.1 at cost 6, +4.6 to +6.3 at cost 10). Its constant policy fails twice over: the wrong action when questions are dear, imperfect execution while they are cheap. Cosmos’s ask rate swings with prompt wording (11% rule-given to 89% costs-given) but not with the cost number;

the rule-given version feeds the rule its constant stated confidence of 100, so Cosmos asks least exactly where asking is optimal. Qwen ignores every instruction, including the rule. The other four models are equally insensitive to cost, each in its own direction. InternVL3’s ask rate drifts (97% to 77%) but never approaches the act-always policy required beyond the break-even; Pixtral asks on every episode at every cost; Cosmos3-Nano likewise always asks (350 of 350 under stated costs, 94–100% in every prompt version); Llama under-asks throughout (19–29%), which is costliest exactly where questions are cheap and the reference policy asks, with regret reaching +5.05 per episode at Ca =1. Table VIII gives regret by cost and prompt version. Confidence could in principle compensate by selecting which episodes to ask about, but ranked by its best readout it recovers almost none of the fixed-to-oracle gap: 0% for Qwen, InternVL3 and Llama, 8% for Cosmos-Reason2, and 56% only for Pixtral’s token log-probability (Fig. 3, appendix). Two simple triggers, disagreement across repeated runs and more than one cause being listed, recover 0%. The three-action variant behaves as the measured accuracies predict. Once questions grow expensive (Ca ≥ 6), consulting the robot’s own sensors at Csense =0.5 becomes the reference policy’s choice for four of the six models (Qwen, InternVL3, Pixtral, Cosmos3-Nano), and Qwen’s predicted flip from asking to sensing lands exactly at Ca =6. For the other two, sensing is dominated for opposite reasons: Llama cannot fuse the text with its frames, and Cosmos cannot read the text at all. Reading the instrument and profiting from it are different abilities, and a deployment should measure which its model has. What the answer is worth depends on how it is phrased. The values above come from an answerer that names the cause in the words of the listed options. As a first probe of how much that matters, we replay only the commit turn with the same information phrased as a person would put it, implying the cause through an observation (“it kept sliding right out of the fingers”) or describing what a bystander saw. It changes the value of a question sharply for half the models (Table IV, with the frozen reply bank and per-register detail in Appendix F): Qwen falls from 0.788 to 0.538 and Cosmos3-Nano from 0.771 to 0.480 (p < 10−4 , McNemar on paired episodes), while Cosmos-Reason2 and Llama hold within noise. The

family A 1.0

ask rate

0.8 0.6 0.4 0.2 0.0

family B 1.0

ask rate

0.8 0.6 0.4 0.2 0.0 1

3

6

stated cost of asking Ca

break-even band (across models / CI) optimal fixed policy Qwen2.5-VL-7B Cosmos-Reason2-8B

10

Llama-3.2-11B-V InternVL3-8B Pixtral-12B Cosmos3-Nano

Fig. 1. How often each model asks as the stated cost of a question rises from 1 to 10. The dashed step is the reference policy at each model’s measured forced-answer accuracy: ask until the break-even cost near 5.3, then act; the shaded band is the break-even across the accuracy estimate’s interval. No model tracks the cost.

two that fall were near-ceiling under menu phrasing on these episodes (0.92 and 0.98), so the highest ask values here are also the most inflated by an accommodating answerer. On the episodes each model got right under menu phrasing, the robust pair retains 0.79–0.86 of them against 0.59–0.64 for the other two, so the difference is not only headroom. The failure is lexical: with the option words removed, “only half of it is visible” is committed as the wrong-name cause and “it skidded out” as the object being absent. Asking is worth much less than the headline numbers suggest unless the person answers in the robot’s vocabulary, and how much less is a property of the model. This is a single-reply probe against a scripted answerer, not an evaluation of comprehension in conversation; measuring that properly is future work (Section VII). E. What the models ask The value of asking also depends on what is asked. The recorded questions show how the models frame a request for help. Not one took the open form: neither primary ever asked “what went wrong?”, even on episodes it had just refused to diagnose. Qwen’s grasp questions split roughly 60/40 between confirming a suspected cause and ruling one out; on detection failures 97% are confirmations, as is every Cosmos question. They rarely target the model’s own uncertainty, matching the cause it is least sure of in 14%, 3%, and 0% of cases: the models ask about the cause they already favor. The wording is stereotyped too, seven distinct questions across 35 grasp episodes for Qwen and seven across all 139 asks for Cosmos. Put next to the diagnosis results this is a direct contradiction: a model that has just said the cause cannot be determined then asks a question presupposing one. Different question forms buy different information for the same cost, the modern form of the listener-modeling problem [9]; scoring that choice needs a protocol like [22] and richer ambiguity than ours. VII. L IMITATIONS AND F UTURE W ORK We evaluated six open 7B–16B vision-language models; frontier systems are the natural next subjects, and whether

they sit above the majority-class baseline from frames or at the overconfident pole that prior work documents [23], [24] is directly testable with our protocol. Test sets are small (35– 36 episodes per failure type), though the reported effects are extreme enough to survive this. The sensor findings are properties of these families in simulation with a scripted policy: the 0.55 image plateau belongs to our classifier sequence and this renderer, whose images omit cues real cameras carry, such as the look of a heavy object or the sag of a loaded arm. Whether the split holds under a different simulator, a learned policy, or a physical robot is open; repeating the audit on real hardware is the test we would run first. Part of the modality contrast is built in, since injected grasp causes are unrendered physics while detection causes are rendered scene properties; a family whose causes span modalities by design is in progress. The decision layer invites extension. Costs are stipulated, swept fourfold, and a prompt variable, so the decision results measure instruction-conditioned behavior; a cause-by-repair matrix in place of the scalar Cw would tie the policy to each model’s confusion structure. The answerer ignores question wording, so the ask values are conditional on a question having been asked rather than on its quality. Comprehension is the clearest opening: people paraphrase, hedge, answer a different question than the one asked, and correct themselves across turns. Evaluating that needs replies conditioned on the question, multi-turn repair, and human speakers rather than templates, with comprehension scored separately from the decision to ask, as navigation benchmarks now do for question quality [22]. A deployed system would also not stop at one question but would ask, interpret, confirm, act, and re-check until the task is recovered or safely abandoned (Appendix E). VIII. C ONCLUSION Whether a failure hides its cause depends on the sensor and on the model that reads it: detection failures are diagnosable from images, grasp failures almost perfectly from force data but not from frames, and placement failures only appeared diagnosable until the viewpoint check. The six models cannot tell these situations apart. Refusal versus commitment is set largely by prompt layout, accuracy from frames never clears the majority-class baseline, confidence carries no information, and ask rates ignore the price of a question; yet one human reply lifts them to roughly the answerer’s reliability when phrased in their terms, and when the missing information sits in the robot’s own sensors, consulting them first is often rational. For human-robot dialogue, the first turn is itself a cost-sensitive information decision: act when measured competence suffices, consult an onboard sensor when that is cheaper, and ask a person when the missing information lies outside any usable robot channel. That question should open a stateful exchange, and its value depends on whether the robot understands the answer. Both are missing here, so the policy should be tied to measured accuracy and costs, not self-reported confidence (Appendix C).

R EFERENCES [1] S. Honig and T. Oron-Gilad, “Understanding and resolving failures in human-robot interaction: Literature review and model development,” Frontiers in Psychology, vol. 9, p. 861, 2018. [2] E. Zhou, Q. Su, C. Chi, Z. Zhang, Z. Wang, T. Huang, L. Sheng, and H. Wang, “Code-as-monitor: Constraint-aware visual programming for reactive and proactive robotic failure detection,” in 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2025, pp. 6919–6929. [3] Z. Liu, A. Bahety, and S. Song, “Reflect: Summarizing robot experiences for failure explanation and correction,” in Conference on Robot Learning (CoRL), 2023. [4] J. Duan, W. Pumacay, N. Kumar, Y. R. Wang, S. Tian, W. Yuan, R. Krishna, D. Fox, A. Mandlekar, and Y. Guo, “Aha: A vision-languagemodel for detecting and reasoning over failures in robotic manipulation,” arXiv preprint arXiv:2410.00371, 2024. [5] P. Pacaud, R. Garcia, S. Chen, and C. Schmid, “Scaling crossenvironment failure reasoning data for vision-language robotic manipulation,” arXiv preprint arXiv:2512.01946, 2025. [6] D. Li, J. Lei, H. Wang, L. Liu, Y. Yang, Z. Wang, B. Liu, M. Zheng, and Z. Fan, “Learning actionable manipulation recovery via counterfactual failure synthesis,” arXiv preprint arXiv:2603.13528, 2026. [7] X. Zeng, X. Zhou, Y. Li, J. Shi, T. Li, L. Chen, L. Ren, and Y.-L. Li, “Diagnose, correct, and learn from manipulation failures via visual symbols,” arXiv preprint arXiv:2512.02787, 2025. [8] A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, Z. Xu, D. Sadigh, A. Zeng, and A. Majumdar, “Robots that ask for help: Uncertainty alignment for large language model planners,” in Conference on Robot Learning (CoRL), 2023. [9] S. Tellex, R. A. Knepper, A. Li, D. Rus, and N. Roy, “Asking for help using inverse semantics,” in Robotics: Science and Systems (RSS), 2014. [10] K. T. Ly, K. Lu, and I. Havoutis, “Inteliplan: An interactive lightweight llm-based planner for domestic robot autonomy,” IEEE Robotics and Automation Letters, 2026. [11] Y. Dai, J. Lee, N. Fazeli, and J. Chai, “Racer: Rich languageguided failure recovery policies for imitation learning,” arXiv preprint arXiv:2409.14674, 2024. [12] M. M. Reimann, F. A. Kunneman, C. Oertel, and K. V. Hindriks, “A survey on dialogue management in human-robot interaction,” ACM Transactions on Human-Robot Interaction, vol. 13, no. 2, pp. 1–22, 2024. [13] S. M. Lukin, C. Bonial, M. Marge, T. Hudson, C. J. Hayes, K. A. Pollard, A. Baker, A. N. Foots, R. Artstein, F. Gervits, M. Abrams, C. Henry, L. Donatelli, A. Leuski, S. G. Hill, D. Traum, and C. R. Voss, “SCOUT: A situated and multi-modal human-robot dialogue corpus,” in Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING), 2024, pp. 14 445–14 458. [14] R. A. Knepper, S. Tellex, A. Li, N. Roy, and D. Rus, “Recovering from failure by asking for help,” Autonomous Robots, vol. 39, no. 3, 2015. [15] J. Blankenburg, M. Zagainova, S. M. Simmons, G. Talavera, M. Nicolescu, and D. Feil-Seifer, “Human-robot collaboration and dialogue for fault recovery on hierarchical tasks,” in International Conference on Social Robotics. Springer, 2020, pp. 144–156. [16] L. X. Shi, Z. Hu, T. Z. Zhao, A. Sharma, K. Pertsch, J. Luo, S. Levine, and C. Finn, “Yell at your robot: Improving on-the-fly from language corrections,” in Robotics: Science and Systems (RSS), 2024. [17] L. Zha, Y. Cui, L.-H. Lin, M. Kwon, M. G. Arenas, A. Zeng, F. Xia, and D. Sadigh, “Distilling and retrieving generalizable knowledge for robot manipulation via language corrections,” in IEEE International Conference on Robotics and Automation (ICRA), 2024. [18] P. Sharma, B. Sundaralingam, V. Blukis, C. Paxton, T. Hermans, A. Torralba, J. Andreas, and D. Fox, “Correcting robot plans with natural language feedback,” in Robotics: Science and Systems (RSS), 2022. [19] X. Zhao, S. Lin, G. Zhou, Z. Li, S. Li, W. Tao, J. Liu, and Q. Wu, “Ask when it pays: Cost-aware open-ended interaction for instance goal navigation,” arXiv preprint arXiv:2606.03175, 2026. [20] W. Zhou, X. Xiong, Y. Tian, L. Yue, X. Wu, W. Li, C. Zhao, H. Dong, M. Tang, J. Wang, and Z. Zhang, “Esearch-r1: Learning cost-aware mllm agents for interactive embodied search via reinforcement learning,” arXiv preprint arXiv:2512.18571, 2025.

[21] T. Trinh, M. Elfeki, G. Luo, K. Luu, N. Hunt, E. Hernandez, N. Marwaha, Y. Y. He, C. Wang, F. Carabedo et al., “Hil-bench (human-in-loop benchmark): Do agents know when to ask for help?” arXiv preprint arXiv:2604.09408, 2026. [22] E. Zorzi, F. Taioli, Y. Wang, M. Cristani, A. Farinelli, A. Castellini, and L. Bazzani, “Benchmarking interaction, beyond policy: A reproducible benchmark for collaborative instance object navigation,” arXiv preprint arXiv:2604.00265, 2026. [23] T. Wu, C. Zhou, G. Zhao, H. Cao, Y. Pu, and J. Yang, “When robots should say "i don’t know": Benchmarking abstention in embodied question answering,” arXiv preprint arXiv:2512.04597, 2025. [24] D. Yeke, E. S. Temirel, A. Shreekumar, B. Lee, D. Xu, and Z. B. Celik, “The yes-man syndrome: Benchmarking abstention in embodied robotic agents,” arXiv preprint arXiv:2605.20544, 2026. [25] C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang, “Large language models are not robust multiple choice selectors,” in International Conference on Learning Representations (ICLR), 2024. [26] R. El-Yaniv and Y. Wiener, “On the foundations of noise-free selective classification,” Journal of Machine Learning Research, vol. 11, pp. 1605–1641, 2010. [27] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” in Advances in Neural Information Processing Systems (NeurIPS), 2023. [28] R. A. Howard, “Information value theory,” IEEE Transactions on Systems Science and Cybernetics, vol. 2, no. 1, pp. 22–26, 1966. [29] S. Russell and E. Wefald, “Principles of metareasoning,” Artificial Intelligence, vol. 49, no. 1-3, pp. 361–395, 1991. [30] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in International Conference on Machine Learning (ICML), 2017. [31] S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin, “Qwen2.5-VL technical report,” arXiv preprint arXiv:2502.13923, 2025. [32] NVIDIA, “Cosmos 3: Omnimodal world models for physical ai,” arXiv preprint arXiv:2606.02800, 2026. [33] Q. McNemar, “Note on the sampling error of the difference between correlated proportions or percentages,” Psychometrika, vol. 12, no. 2, pp. 153–157, 1947. [34] C. Spearman, “The proof and measurement of association between two things,” The American Journal of Psychology, vol. 15, no. 1, pp. 72–101, 1904. [35] F. Gervits, R. Thielstrom, A. Roque, and M. Scheutz, “It’s about time: Turn-entry timing for situated human-robot dialogue,” in Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), 2020, pp. 86–96. [36] R. Roy, J. Raiman, S.-g. Lee, T.-D. Ene, R. Kirby, S. Kim, J. Kim, and B. Catanzaro, “Personaplex: Voice and role control for full-duplex conversational speech models,” arXiv preprint arXiv:2602.06053, 2026. [37] S. Dass, K. Pertsch, H. Zhang, Y. Lee, J. J. Lim, and S. Nikolaidis, “PATO: Policy assisted teleoperation for scalable robot data collection,” in Robotics: Science and Systems (RSS), 2023.

A PPENDIX A AUDIT D ISCUSSION A fair question is why we accept the force-data score for grasp failures while rejecting the image score for placement failures, when both classifiers profit from the injected cause. The distinction is not what the feature measures but what survives scrutiny. The placement classifier read the container’s position in the image, a reading that collapsed under a heldout camera angle and, when suppressed by normalization, resurfaced through a different channel. The force-data classifier reads the mechanical consequences of the failure through streams any real wrist sensor produces, its score is unchanged when every ground-truth-adjacent feature is removed, and it separates causes even when perturbation magnitude is matched. In an injected-cause benchmark every honest

diagnosis ultimately traces the injected cause; the standard we apply is that the tracing must run through deployable sensors and survive every transfer we can construct. Two conventions follow from this. No filter may depend on any statistic of the number being reported, and claims that information is absent are reported as the plateau of a classifier sequence, never as a proven limit. A related detail from the placement family: normalizing image contrast does not remove its shortcut; it moves it between the appearance and viewpoint checks (0.926 falling to 0.881 on one axis, 0.605 rising only to 0.683 on the other). Who measures the accuracy, and when. A natural objection is that deciding whether to ask requires knowing the model’s accuracy in advance. The objection conflates evaluation with deployment. In evaluation, the measured accuracy defines the reference policy that regret is scored against; the gap between a model’s behavior and that reference policy is the reported result, so no circularity arises. In deployment, accuracy would be measured once during commissioning, on held-out failures staged before the system is fielded, in the same way conformal methods require a calibration set before their guarantees hold [8]. Only a model whose confidence tracked its correctness episode by episode could skip this step and adapt from the inside; Section VI-B shows that ability is missing in the models we test, which is why one externally measured number outperforms their own judgment. A PPENDIX B F REE -F ORM D IAGNOSIS P ROBE The multiple-choice format could in principle mask competence that free elicitation would reveal. To test this, two models with different backbones, Pixtral-12B and Cosmos3Nano, were asked on every detection-failure test episode for a free-form diagnosis with no options: “state what you think caused the failure.” Accuracy falls to 0.08 and 0.00, below the multiple-choice 0.417 and 0.389, so the option list surfaces perception hypotheses the models do not generate on their own; for contrast, handing either model the ground-truth scene fact as one sentence of text yields 0.90 and 1.00 through the same multiple-choice question. Free-form answers are scored by a per-cause keyword rubric that counts any mention of the injected mechanism as correct, lenient by construction, with non-matching answers reviewed by hand. The answers themselves show what fills the gap. On episodes with the target 62–94% occluded, Pixtral diagnoses in grasp vocabulary (“incorrect gripper positioning,” “did not correctly grasp”), and Cosmos3-Nano narrates motion (“moved away instead of approaching”) while asserting it sees the hidden object (“the milk carton, which was located on the table”). On dimmed episodes, neither model ever mentions lighting. The models do not produce wrong perception hypotheses; they produce no perception hypotheses, substituting a generic manipulation story for the scene in front of them. Given the fact, the models find the cause; given the pixels, they do not find the fact.

A PPENDIX C I MPLICATIONS FOR P RACTICE Our results suggest a concrete procedure for deployments. Diagnosis accuracy can be measured during commissioning by staging failures of each type the robot will encounter, in the same way conformal methods require a calibration set before their guarantees hold. With deployment-specific costs assigned to a repair, a wrong repair, a person’s interrupted attention, and a read of the robot’s own sensors, the threshold of Eq. (3) then identifies, per failure type, whether acting or consulting a cheaper information source is the better default: the robot’s own sensors where they carry the answer, a person where they do not. Because any externally measured accuracy goes stale, re-measurement is warranted whenever the model, the environment, or the costs change. What the results argue against is tying this decision to the model’s confidence: on our benchmark, a single externally measured number outperformed the models’ episode-by-episode judgment, which carried no information. For model builders, the results suggest two priorities. The first is identifying which sensor carries the diagnostic information before scaling the vision pipeline; for our grasp failures it is the force data, and models that could not recover the cause from frames could recover it from ten lines of text. The second is calibrated self-assessment, confidence that tracks correctness: the missing capability whose value our oracle policy bounds, and the one that would let a robot set its own threshold from the inside. A PPENDIX D E XTENDED R ELATED W ORK Failure diagnosis. REFLECT [3] converts an episode into text and prompts a language model to explain the failure, without training. AHA [4] fine-tunes a VLM on synthetically perturbed demonstrations, and further systems scale related data-generation approaches [5], [6]; RACER [11] guides imitation-learned recovery with rich language annotations; recent benchmarks draw on real-world trajectories [7]. Where such a system reports imperfect accuracy, the shortfall cannot be split between a weak model and undeterminable evidence without a measurement like ours. Interactive planners and help-seeking. InteLiPlan [10] incorporates human interaction into a lightweight LLM-based planner for domestic robots; we supply the evaluation that says when interrupting a person is justified. Ask When It Pays [19] derives question costs from an information-gain analysis and penalizes each query in its success metric, so its decision reduces uncertainty about a navigation goal, while ours is a threshold from measured diagnosis accuracy. ESearch-R1 [20] is closest in structure: its Ask, GetMemory, and Navigate actions parallel our ask, sense, and act, and it trains a multimodal model to trade information gain against per-action costs. Our three actions are analogous, but the policy is computed rather than learned, which is what lets us ask whether off-the-shelf models match the optimum. QAskNav [22] is complementary on the evaluation side, decoupling

TABLE IV F INAL TASK ACCURACY WHEN THE MODEL ASKS , BY HOW THE HUMAN PHRASES THE ANSWER . T HE SAME INFORMATION IS DELIVERED IN EACH CONDITION ; ONLY THE WORDING CHANGES . Model Qwen2.5-VL-7B Cosmos3-Nano Cosmos-Reason2-8B Llama-3.2-11B-V

Menu-phrased

Indirect

Narrative

0.788 0.771 0.707 0.753

0.538 0.480 0.680 0.691

0.599 0.542 0.627 0.654

Accuracy on the episodes where each model chose to ask. Menu-phrased is the oracle used elsewhere in the paper, which names the cause in the words of the answer options; Indirect implies the cause through an observation; Narrative describes what a bystander saw without diagnosing. Correctness of the reply is held fixed across conditions; only the wording of the informative replies changes.

question asking from navigation so that question quality can be scored on its own. A PPENDIX E F ROM I NITIATION TO C ONTINUING D IALOGUE A deployed system would not stop at one question. It would maintain a loop of asking, interpreting, confirming, acting, and re-checking until the task is recovered or safely abandoned, which raises three questions this benchmark only opens. Spoken interaction adds turn-entry timing decisions of its own [35], along with recognition and grounding errors, and full-duplex speech models make continuous listening and speaking plausible [36]; these enter our framework as extra noise in the answer channel and lower the value of a question below what Table IV measures. Asking also interrupts a physical process: a robot holding an unstable object may need to reach a safe checkpoint before speaking, and the reply must then be bound to the state from which execution resumes. And corrective feedback need not be verbal at all, since a person may demonstrate, guide the arm, or teleoperate briefly [37]; treating such an intervention as a dialogue turn raises the further question of whether it is a one-time rescue or a reusable correction. A PPENDIX F A NSWER -P HRASING BANK The phrasing conditions of Table IV deliver identical information in three registers. Replies were generated once per cause, screened so that no indirect or narrative reply contains a word from any cause’s vocabulary, then frozen; the correct/uninformative/wrong draw and its per-episode seed are unchanged across conditions, so only the wording of the informative replies differs. Two of the fourteen sets follow. Grasp failure, object heavier than expected. Menu-phrased: “that one is heavier than it looks.” Indirect: “The object barely moved when the robot lifted it, it seemed like it was too much for the arm.” / “The arm was trembling while trying to lift the object, it didn’t go very far.” / “I noticed the arm was straining a lot while attempting to pick up the object.” Narrative: “That thing barely lifted an inch despite the robot’s best efforts.” / “The robot’s attempt to lift that object didn’t get it far off the table.” / “The arm seemed to struggle and the object hardly left the table’s surface.”

Detection failure, target occluded. Menu-phrased: “it’s tucked behind the other thing.” Indirect: “There seems to be something blocking the view of what you’re trying to pick up.” / “Only half of the thing you’re aiming for is visible.” / “It looks like another object is in front of the one you’re reaching for.” Narrative: “Another item was in the way, blocking part of the target object.” / “There was a second object right in front of the one the robot tried to pick up.” / “I could see the target object was there, but another item was in front of it.” The misroutings are lexical. Given the second indirect reply above, one model commits to the instruction naming an object the robot cannot recognize, and given a slippage description (“it skidded out”) another commits to the object being absent: with the option words removed, surface overlap with the wrong option outweighs the meaning of the sentence. A PPENDIX G D IAGNOSIS P ROMPT Both rounds of the diagnosis test use the templates below, identical for every model. {options} is the lettered cause list for the family; in the refusal-allowed round the refusal line is appended to it, with the option-position sweep varying its wording across “cannot be determined from the available observations,” “not determinable from these frames,” and “insufficient evidence to identify the cause,” and its position between last and first. The forced-answer round omits the refusal option entirely. Templates for the few-shot, telemetry, and decision conditions follow the same structure and are described in Section V. Diagnosis, refusal allowed. You are a robot's failure analyst. The robot attempted: "{instruction}". The attempt failed. You are shown frames from the episode in order. Which of the following is the cause of the failure? {options} Answer with the single letter, then on a new line your confidence 0-100 that this letter is correct.

Diagnosis, forced answer. You are a robot's failure analyst. The robot attempted: "{instruction}". The attempt failed. You are shown frames from the episode in order. Which of the following is the most likely cause of the failure? {options} You must choose one option even if you are uncertain . Do not decline to answer. Answer with the single letter, then on a new line your confidence 0-100 that this letter is correct.

A PPENDIX H A DDITIONAL R ESULTS Table V reports the remaining smaller relative, Qwen2-VL2B, with the same three accuracy numbers as Table II. In the decision task it rarely produces a well-formed choice, echoing both options on nearly every episode, so we do not read its decision behavior as a policy. Table VI gives the full imageclassifier sequence behind the 0.55 plateau of Section IV-B.

TABLE V T HE REMAINING SMALLER RELATIVE , WITH THE SAME THREE ACCURACY NUMBERS AS TABLE II: AT OR BELOW THE MAJORITY- CLASS BASELINE ON BOTH FAMILIES , CONSISTENT WITH THE MAIN RESULTS .

Model

Fam.

Qwen2-VL-2B A Qwen2-VL-2B B

Cert. Overall Commit 0.545 0.910

0.000 0.167

Rate Forced

– 0% 0.167 100%

Gap

0.314 0.545 0.167 0.743

TABLE VI I MAGE CLASSIFIERS FOR GRASP FAILURES ( FAMILY A): ACCURACY PLATEAUS NEAR 0.55.

Image classifier

Acc. Shuffle

Tint Azimuth Retained

Final-state snapshot (2×24×24 px) Motion features (8×32×32 + diffs) DINOv2 frozen, linear head (4×224 px) DINOv2 frozen, boosted head (4×224 px) 3D-conv net, end to end (8×48×48)

0.442 0.503 0.523 0.545 0.531

0.445 0.498 0.531 0.545 0.537

0.338 0.327 0.345 0.341 0.338

0.345 0.413 0.460 0.506 0.494

−0.02 0.42 0.64 0.80 ✓ 0.80 ✓

Majority-class baseline 0.347. Shuffle: retrain on scrambled labels; must sit at the baseline. Tint / Azimuth: held-out scene appearance / camera angle. Retained: share of the above-baseline margin surviving the held-out angle; pass (✓) above 0.75; negative means held-out accuracy fell to the baseline.

TABLE VII R EFUSAL RATES FOR THE TWO C OSMOS MODELS UNDER THE OPTION - POSITION SWEEP : THREE PHRASINGS ( V 1– V 3) OF THE “ CANNOT BE DETERMINED ” OPTION , EACH IN LAST AND FIRST LIST POSITION . G RASP REFUSAL IS PROMPT- ROBUST IN BOTH GENERATIONS ; DETECTION REFUSAL COLLAPSES FOR THE 8B BUT BARELY MOVES FOR THE NANO .

Model

Fam. Last (v1/v2/v3) First (v1/v2/v3)

Reason2-8B Cosmos3-Nano Reason2-8B Cosmos3-Nano

A A B B

100 / 100 / 100 100 / 100 / 100 100 / 100 / 100 97 / 100 / 100 94 / 89 / 97 3 / 6 / 25 100 / 100 / 97 92 / 61 / 83

Values are refusal rates in % with the refusal option in last or first list position. Families: A grasp, B detection.

Figure 2 plots the three confidence readouts against decidability. Table VIII reports regret by model, family, prompt version, and question cost. Figure 3 shows the selectiveasking comparison at Ca =3. Table VII reports the optionposition sweep for both Cosmos models in full. Figure 4 shows example episodes from each failure family. In the confidence analysis of Section VI-C, InternVL3’s token probabilities are saturated (all values within 8 × 10−5 of 1.0) and carry no usable ranking; they are excluded rather than reported as a correlation.

Qwen2.5-VL-7B stated confidence

1.0 0.8

Cosmos-Reason2-8B stated confidence

token logprob

5-sample modal frequency

= 0.07

= 0.20

= 0.07

= 0.10

= 0.46

= n/a

= +0.23

= 0.55

= 0.42

0.6 0.4 0.2 0.0 1.0 0.8 0.6 0.4 0.2 0.0 1.0

Cosmos3-Nano stated confidence

verbalized 0-100

0.8 0.6 0.4 0.2 0.0

0.4

0.6

0.8

decidability ptop of the episode

1.0

0.4

0.6

0.8

decidability ptop of the episode correct

1.0

0.4

0.6

0.8

decidability ptop of the episode

1.0

wrong

Fig. 2. Stated confidence does not track how much the images determine the cause. Grasp failures, forced-answer round; each column is one confidence readout; the dashed line marks the model’s own accuracy (0.314 for all three). No readout is positively informative for any of the three models.

TABLE VIII R EGRET AGAINST THE BEST FIXED POLICY, BY PROMPT VERSION ( PLAIN , COSTS - GIVEN , RULE - GIVEN ). T HE REFERENCE ASKS - ALL WHILE THE MODEL’ S FORCED - ANSWER ACCURACY a IS BELOW p⋆ (Ca ) AND ACTS - ALL OTHERWISE ; WITH a ≈ 0.32 ON GRASP FAILURES THE BREAK - EVEN IS Ca ≈ 5.3, SO Ca =6 IS ACT- OPTIMAL . R EGRET IS MEAN REALIZED COST MINUS THE REFERENCE ’ S , ± A HALF - WIDTH FROM THE 10 K - RESAMPLE BOOTSTRAP. C ELLS MARKED † ARE INDETERMINATE : THE SIGN OF REGRET CHANGES WITHIN THE ACCURACY ESTIMATE ’ S CONFIDENCE INTERVAL . A LL NEGATIVE CELLS ARE †- MARKED AND WITHIN THEIR INTERVALS : THE REFERENCE ’ S ACCURACY INPUT COMES FROM 25 CALIBRATION EPISODES , SO IT IS OPTIMAL IN EXPECTATION GIVEN THAT ESTIMATE , NOT AN ORACLE , AND CAN BE BEATEN BY SAMPLING NOISE . OVERLAPPING INTERVALS MEAN CONDITION DIFFERENCES AT n=35 ARE MOSTLY NOT SEPARABLE . R EGIMES Ca ∈ {3, 10} SIDE BY SIDE ; THE FULL SWEEP OVER Ca ∈ {1, 3, 6, 10} SHOWS THE SAME PATTERN . Ca = 10

Ca = 3 Model

Fam.

Cond.

Ask

Cost

Reg.

Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano InternVL3-8B InternVL3-8B InternVL3-8B InternVL3-8B Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Pixtral-12B Pixtral-12B Pixtral-12B Pixtral-12B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B

A A A B B B A A A B B B A A B B A A A B B B A A B B A A A B B B

costs plain rule costs plain rule costs plain rule costs plain rule costs rule costs rule costs plain rule costs plain rule costs rule costs rule costs plain rule costs plain rule

86% 46% 11% 61% 3% 14% 100% 94% 100% 100% 100% 94% 86% 43% 6% 3% 23% 9% 3% 28% 19% 8% 100% 100% 100% 100% 100% 100% 100% 100% 100% 100%

6.71 7.23 7.63 7.56 8.31 8.64 5.43 5.83 6.00 5.67 5.67 6.61 8.14 8.00 7.00 6.92 7.40 8.40 7.94 8.50 7.42 8.47 9.14 7.14 5.94 7.89 6.86 5.43 6.86 6.22 5.94 6.22

+1.19 ±1.5 +1.71 ±1.6 +2.11 ±1.6 +2.16 ±1.5 +2.91 ±1.4 +3.24 ±1.3 −0.09 ±1.1† +0.31 ±1.4 +0.48 ±1.3 +0.27 ±1.2 +0.27 ±1.2 +1.21 ±1.4 +2.62 ±1.5 +2.48 ±1.6 +1.30 ±1.6 +1.22 ±1.6 +1.88 ±1.5 +2.88 ±1.5 +2.42 ±1.5 +2.86 ±1.5 +1.78 ±1.4 +2.83 ±1.3 +3.56 ±1.7 +1.56 ±1.4 +0.18 ±1.2 +2.13 ±1.7 +1.34 ±1.4 −0.09 ±1.1† +1.34 ±1.4 +0.46 ±1.4 +0.18 ±1.2 +0.46 ±1.4

Model

Fam.

Cond.

Ask

Cost

Reg.

Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos-Reason2-8B Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano Cosmos3-Nano InternVL3-8B InternVL3-8B InternVL3-8B InternVL3-8B Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Llama-3.2-11B-V Pixtral-12B Pixtral-12B Pixtral-12B Pixtral-12B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B Qwen2.5-VL-7B

A A A B B B A A A B B B A A B B A A A B B B A A B B A A A B B B

costs plain rule costs plain rule costs plain rule costs plain rule costs rule costs rule costs plain rule costs plain rule costs rule costs rule costs plain rule costs plain rule

86% 46% 11% 58% 3% 14% 100% 94% 100% 100% 100% 94% 77% 43% 0% 3% 29% 9% 3% 22% 19% 8% 100% 100% 100% 100% 100% 100% 100% 100% 100% 100%

13.00 10.43 8.43 11.83 8.50 9.61 13.00 12.43 13.00 12.94 12.67 13.22 12.71 11.00 7.39 7.11 9.00 9.00 8.14 10.44 8.78 9.06 16.71 14.14 12.94 14.89 14.14 12.43 13.86 13.22 12.94 13.22

+5.20 ±1.7 +2.63 ±2.0 +0.63 ±1.7† +4.83 ±1.7 +1.50 ±1.4† +2.61 ±1.1 +5.20 ±1.3 +4.63 ±1.7 +5.20 ±1.3 +5.94 ±1.2 +5.67 ±1.2 +6.22 ±1.4 +4.91 ±1.7 +3.20 ±2.0 −1.61 ±1.7† −1.89 ±1.7† +1.20 ±1.6† +1.20 ±1.6† +0.34 ±1.7† +1.84 ±1.4 +0.18 ±1.4† +0.46 ±1.2† +8.51 ±1.7 +5.94 ±1.4 +3.54 ±1.2 +5.49 ±1.7 +6.34 ±1.6 +4.63 ±1.1 +6.06 ±1.4 +3.82 ±1.4 +3.54 ±1.2 +3.82 ±1.4

mean realized cost

mean realized cost

family A (Ca = 3) 8 7 6 5 0.0

0.2

0.0

0.2

0.4

0.6

0.8

1.0

0.6

0.8

1.0

family B (Ca = 3)

8 6 4 2 0.4

ask rate Qwen2.5-VL-7B Cosmos-Reason2-8B Llama-3.2-11B-V InternVL3-8B

Pixtral-12B Cosmos3-Nano verbalized logprob

sample freq. evidence policy oracle skyline constant tier

Fig. 3. Realized cost against ask rate at Ca = 3, the two family panels stacked. Curves: asking whenever a confidence readout falls below a swept threshold; the evidence policy ranks episodes by decidability instead. Dashed line: the best fixed policy (the reference policy). Star: the oracle policy that asks exactly on the model’s errors (uses ground truth; an upper bound, not achievable). On the grasp panel the model curves coincide with the evidence curve: with accuracy near zero, every episode is an error, so no ranking has anything to reorder. The overlap is the result, not a rendering artifact.

A: grasp failure B: detection failure C: placement failure Fig. 4. Example episodes from the three failure families, four of the eight frames the models receive. Top (A): a grasp failure. The frames look unremarkable because the injected causes, grip force, object mass, and surface friction, are never rendered; this is why image classifiers plateau near 0.55 while the force-data classifier reaches 0.986. Middle (B): a detection failure; the occluder hiding the target is plainly visible, matching the 0.910 image certificate. Bottom (C): a placement failure; the container’s state is the injected cause, and its position in the frame is the shortcut behind the disqualified 0.998 score.

Record · ID 1006888 · SHA-256 7ad8bbfe0e4d4e7b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.