Conceptio › Archive › arXiv CS
arXiv CSopen access

What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

W HAT S HOULD W E A SK N EXT ? R ETRIEVAL -AWARE Q UESTION L EARNING UNDER PARTIAL E VIDENCE Lyucheng Qian1 , John Yuehan Zhang2 , Pingyu Wang1∗

arXiv:2609.21924v1 [cs.AI] 18 Sep 2026

A BSTRACT Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs, although question value is determined by the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification. RAVEL initializes from supervised question generation, observes the current Top-4 candidates directly, and optimizes the question policy with rank feedback from the full question–answer–retrieval loop. Experiments on Interactive-PEDES show that RAVEL delivers progressively stronger retrieval performance across five interaction rounds. Further analysis shows that RAVEL reallocates the questioning budget toward localized open-ended attributes, which provide more useful retrieval evidence and yield the largest gains on initially difficult queries.

1

I NTRODUCTION

Interactive retrieval is a general information-acquisition problem: when the initial evidence is incomplete, an agent must decide what to ask before it can make a reliable search decision. This setting appears in chat-based search (Levy et al., 2023), visual search with relative feedback (Kovashka et al., 2015), and question-driven video retrieval (Madasu et al., 2022; Liang & Albanie, 2023). We study it through Interactive ReID, where an initial person description is refined through a questioner–answerer–retriever loop over a gallery. The difficulty is that a question has no fixed value in isolation. Its usefulness depends on the answer it elicits, the dialogue state, the text-processing pipeline, and the resulting gallery ranking, making question selection a sequential decision problem under a retrieval objective. Figure 1 illustrates the consequence: from the same Top-4 state, a question that appears highly discriminative can lower the target rank, while a retrieval-aware question raises it substantially. Effective question learning must therefore optimize the complete interaction loop. Existing interactive ReID pipelines largely use offline behavior cloning. Fine-grained descriptions are decomposed into candidate QA pairs, and a greedy look-forward strategy selects the order with the largest immediate ranking improvement (Lu et al., 2025). The resulting questioner imitates this fixed order. This formulation motivates three questions: Although these labels are derived from ranking changes, the supervision remains offline. Each candidate QA pair is scored before deployment under a fixed state, and the questioner is trained only to reproduce the selected order; it does not observe the answer elicited by its own wording, the cleaned evidence entering the retriever, or the subsequent re-ranking. Behavior cloning can therefore learn static candidate-separation cues as a proxy for rank-improving questions, without learning their conditional value as the dialogue and candidate state evolve. 1. Does candidate-level discriminativeness reliably define a good question? 2. Does the questioner actually use the candidate images it receives? ∗

Corresponding author

1

Sichuan University

2

University of California, Berkeley

1

3. If offline supervision is unreliable, can the question strategy be learned directly from online retrieval feedback?

Our analysis establishes three findings. First, question discriminativeness and human-perceived usefulness show almost no association with downstream rank change, and model judgments agree poorly with human judgments. Second, the questioner responds to visual tokens while showing limited sensitivity to candidate order and candidate composition. Third, online reinforcement learning raises retrieval performance by reallocating questioning behaviors across rounds, with the largest conditional gains on initially difficult queries. We introduce RAVEL (Retrieval-Aware Verbal Evidence Learning), a retrieval-guided online policy-optimization framework for question learning under partial evidence. RAVEL initializes the questioner with supervised LLaVA-ReID training, presents the current Top-4 retrieval results directly, and learns in the actual multi-turn retrieval loop. The learning signal couples each question with the answer it elicits and the resulting ranking improvement; a validity gate keeps exploration inside the interaction protocol. On Interactive-PEDES, RAVEL reaches 73.73 Rank-1 after five rounds, surpassing LLaVA-ReID by 5.79 points and every interactive baseline in our comparison. Our contributions are:

2. Retrieval-aware learning and analysis. RAVEL learns questions from online rank feedback over the complete question–answer–retrieval loop. We analyze how its trained policy schedules question types, targets localized evidence, and accumulates retrieval gains.

R ELATED W ORK

2.1

I NTERACTIVE R ETRIEVAL

Memory: "I saw a man wearing all black standing by a door."

Same pre-question Top-4 Rank before: 31

cand 1

cand 2

I should ask what separates the four candidates.

Is the man wearing a short jacket?

Yes, the man is wearing a fitted black jacket, which can be considered a short jacket.

1. Question-value diagnosis. We find that static question labels weakly predict retrieval outcomes, exposing the limits of offline QA ordering and candidate-level sensitivity.

2

A good-looking question can still fail retrieval

cand 3

cand 4

I should ask what helps the retriever move the target.

What kind of shoes was the man wearing?

The man was wearing black shoes with white soles.

GT

GT

Rank 31 Rank 38

(a) Discriminativeness-driven

Rank 31 Rank 2

(b) Retrieval-aware

Figure 1: A discriminative question can still hurt retrieval. From the same Top-4 state, it moves the target from Rank 31 to 38, whereas a retrieval-aware question raises it to Rank 2.

Interactive retrieval systems iteratively collect evidence and revise a ranked result. Earlier systems solicit relative-attribute feedback (Kovashka et al., 2015), generate dialogue questions for video retrieval (Madasu et al., 2022; Liang & Albanie, 2023), or rewrite image-search queries with language and vision-language models (Zhu et al., 2024). Dialog-based image retrieval has also optimized rank improvement directly with reinforcement learning (Guo et al., 2018). Interactive ReID instantiates this setting with a questioner, an answerer, and a visual gallery. LLaVA-ReID constructs Interactive-PEDES and trains the questioner using offline QA ordering (Lu et al., 2025). ChatIR uses chat-based image retrieval (Levy et al., 2023), while PlugIR constructs questions from candidate descriptions (Lee et al., 2024). Human-centered interaction similarly uses multimodal large language model (MLLM) question answering to refine difficult text-based ReID queries at test time (Qin et al., 2025). Complementary person-search and pedestrian attribute-recognition work improves robustness to noisy correspondences, ambiguous image–text pairs, and unreliable fine-grained attribute evidence (Qin et al., 2024; Sun et al., 2026; Lou et al., 2026). Recent multimodal embedding and reranking models further strengthen static cross-modal relevance estimation (Li et al., 2026). RAVEL advances this setting by treating question generation as a retrieval-conditioned sequential decision problem. 2

Figure 2: Question-quality judgments and downstream retrieval gain. Each point is one of the 2,000 observed test-set interaction states. The y-axis is reciprocal-rank gain, ∆RR = 1/rafter − 1/rbefore , so positive values indicate rank improvement; the dashed line marks ∆RR = 0. Grey points show all states with small horizontal jitter, red points mark positive judgments followed by a negative gain, and green markers show bin means with 95% confidence intervals. The panels report Spearman and point-biserial correlations with their p-values.

2.2

Q UESTION L EARNING AND R EINFORCEMENT L EARNING

Multimodal question generation, active perception, and adaptive sensing all select actions that expose missing task-relevant information. Recent work trains MLLMs for active perception and finegrained visual reasoning through online or reinforcement-learning (RL) objectives (Zhu et al., 2026; He et al., 2026; Zhang et al., 2026). Group-based policy optimization provides an efficient mechanism for language-model reasoning (Shao et al., 2024). RAVEL brings this perspective to interactive retrieval: the action is an open-ended natural-language question, the environment returns a response, and retrieval converts the resulting evidence into rank feedback. This formulation directly optimizes the information-acquisition decisions that shape the final ranked result. Guo et al. (Guo et al., 2018) also formulate dialog-based image retrieval as RL, using a learned user simulator to provide natural-language feedback and rewarding rank improvement at each dialog turn. RAVEL shares the use of retrieval feedback but differs in the decision interface and optimization target: its action is a free-form question generated by a multimodal questioner conditioned on the current Top-4 candidates, the response comes from a frozen multimodal answerer, and validity filtering and answer cleaning determine what reaches a frozen ReID retriever. Token-level clipped group-relative updates then compare alternative questions at a fixed partial observation. Retrievalaware RL has also been used to optimize query generation directly against retrieval outcomes (Jiang et al., 2025); process-supervised retrieval reasoning addresses credit assignment in long multi-hop trajectories with intermediate reward models or step-level exploration (Wang et al., 2026; Samarinas et al., 2026). Value-of-information methods likewise formalize selective information acquisition under resource constraints (Bhope et al., 2026); RAVEL instead allocates a fixed interaction budget to questions that create useful ranking evidence.

3

O FFLINE S UPERVISION D IAGNOSIS

3.1

P ROBLEM S ETUP

Let M denote the private witness memory available only to the answerer, and let X denote the initial description visible to the questioner. At round t, Ct is the ordered Top-4 candidate set, Rt is the accumulated retrieval text, and rt is the target rank maintained by the environment. The environment state is st = (M, X, Q<t , A<t , Ct , Rt , rt , t), whereas the questioner receives the partial observation ot = (X, Q<t , A<t , Ct , Rt , t) and never observes M or rt . It generates qt , the answerer produces at , and the retriever updates the ranking. The retrieved text is updated by applying the original answer-cleaning rules to the witness response and appending the result to the previous text. We use the same gallery, retriever, answer-length limit, and text preprocessing protocol across comparisons. 3

v1

v2

v3

v4

t1

t2

t3

···

tL

q1

Layer 1

q2

Layer 2

q3

Layer 3 · · · Layer N

q4 q5 · · ·

qL

Figure 3: Layer-wise visual reliance during teacher-forced question scoring. Darker cells indicate larger attention weights in the visual-token region. 3.2

O FFLINE S UPERVISION FAILURE

A natural hypothesis is that a question separating the current Top-4 candidates should be useful. We test it on 2,000 observed test-set interaction states, balanced across rounds, using Qwen3-VL-32B judgments of discriminativeness and perceived usefulness; 101 cases also receive human usefulness labels. Using reciprocal-rank gain, the full-sample Spearman correlations are 0.051 and 0.048, while the point-biserial correlations with the rank-improvement event are 0.107 and 0.099. The 101 human labels show no reliable association with reciprocal-rank gain (ρ = −0.016, p = 0.875) or the improvement event (rpb = −0.041, p = 0.687). Human–model agreement is 65.35% with Cohen’s κ = 0.312. Figure 2 visualizes the automatic analysis; label mappings and per-round statistics are in Supplementary Material A. These results do not support treating static question labels as reliable supervision, while the human subset also suggests that the weak association is not solely an artifact of VLM labeling noise. The required supervision is sequential: each question must be evaluated through the answer it elicits and the retrieval evidence that answer creates. This result directly motivates online policy optimization. 3.3

V ISUAL R ELIANCE

We also examine visual reliance in candidate-conditioned question generation, following recent analyses of underused decisive regions in MLLM perception (Peng et al., 2026; Wei et al., 2026; Yuan et al., 2026). We use three diagnostics: Visual Reliance Ratio (VRR) for attention assigned to visual tokens, Image Ablation Sensitivity (IAS) for the change in question likelihood after image removal, and Candidate Sensitivity (CS) for one-candidate ablations. For an observed question q = (q1 , . . . , qL ), let o denote the questioner-visible non-image context (initial description, dialogue history, accumulated retrieval text, and round index), and let C denote the ordered candidate images. Teacher forcing evaluates pθ (q | o, C) =

L Y

pθ (qℓ | q<ℓ , o, C).

(1)

ℓ=1 L

1X log pθ (qℓ | q<ℓ , o, C). ℓTF (q; o, C) = − L

(2)

ℓ=1

(n)

Let αℓ,u denote the attention weight averaged over all heads in layer n from input position u while scoring qℓ , with V the visual-token positions, U all attended input positions, and N the number of layers included. The diagnostics are PN PL P (n) n=1 ℓ=1 v∈V αℓ,v VRR(q) = PN PL P . (3) (n) n=1 ℓ=1 u∈U αℓ,u IAS(q) = ℓTF (q; o, ∅) − ℓTF (q; o, C). 4

(4)

K

1 X CS(q) = [ℓTF (q; o, C \ ck ) − ℓTF (q; o, C)] . K

(5)

∆R1 = ℓTF (q; o, CR1 ) − ℓTF (q; o, C).

(6)

∆Shuffle = ℓTF (q; o, π(C)) − ℓTF (q; o, C).

(7)

k=1

Here K = 4, ∅ denotes image ablation, C \ck zeros the visual input for candidate ck , CR1 retains the Rank-1 image, and π(C) deterministically permutes the same Top-4 set. All likelihood differences use the token-averaged ℓTF values above, excluding prompt, dialogue, and image-token positions. The five complementary diagnostics are reported in Supplementary Material A. Positive IAS and CS indicate stronger visual conditioning, while the two candidate-control values measure sensitivity to candidate composition and order. The attention and likelihood masks follow the definitions above. The results reveal a clear separation between visual conditioning and candidate-level discrimination: IAS and VRR increase across rounds, while CS remains small, and candidate retention changes the likelihood of the observed question very little. These findings motivate RAVEL, which optimizes each question by its downstream retrieval utility after answer processing and re-retrieval rather than imitating a fixed offline question order.

4

O NLINE Q UESTION O PTIMIZATION

4.1

F RAMEWORK

RAVEL follows a two-stage training procedure. It first obtains a supervised cold-start questioner by fine-tuning LLaVA-OneVision-Qwen2-7B-ov (Li et al., 2024a) for one epoch with QLoRA (Dettmers et al., 2023) on the official question-generation supervision. RAVEL then initializes online policy optimization from this checkpoint and updates the language adapter in the retrieval environment using Group Relative Policy Optimization (GRPO) with a group-relative clipped objective (Shao et al., 2024). At each interaction state, the policy receives the corresponding partial observation, including the current Top-4 retrieval results, and samples G candidate questions. Each question is answered, cleaned, appended to the retrieval text, and scored by the frozen retriever through the updated target rank. This trajectory supplies the reward for the question-policy update. In words, the online loop maps each state to a group of questions, evaluates the answer and the resulting rank change for each question, and uses the group-normalized feedback to update the question policy. The answerer and retriever remain fixed during RL, concentrating learning capacity on the question strategy. This modular design makes every policy update attributable to the retrieval evidence elicited by the questioner. It also enables direct cross-retriever evaluation: the learned policy can be paired with a separately trained retrieval backbone without further questioner optimization. 4.2

S TATE AND E NVIRONMENT

The policy action is an open-ended natural-language question sampled from πθ (qt | ot ). The questioner observes the initial description, all previous questions and answers, the current accumulated retrieval text, the current round, and the candidate context, but not the private witness memory or the target rank. The action space is not restricted to a fixed list of QA pairs. The answerer receives the witness memory and the generated question. Its output is processed with the same answer-length limit and original cleaning rules used during evaluation. The processed answer is appended to the retrieval text. The retriever then computes the target identity rank over the fixed gallery. This creates an online feedback loop in which a question is valued by the evidence it causes the witness to contribute and by how that evidence interacts with the retriever. 4.3

R EWARD D ESIGN

The reward preserves reciprocal-rank improvement while emphasizing successful entry into Rank-1. Let rbefore and rafter denote the target ranks before and after adding the answer. The reciprocal-rank 5

Stage 2: Supervised cold start

Stage 1: Retriever training Contrastive Learning

image embedding

Coarse description: I saw a man wearing all black standing by a door ...

Retriever

text embedding

Top-4 Candidates

Question Supervision

Dialogue history

SFT Questioner

Instruction

Stage 3: Retrieval-aware question learning

Reference policy

Trainable Module

Frozen Module

Top-4 Candidates

Questioner

Question A: ... Question B: ... ...... Question G: ...

Validity gate

Witness

Retriever

Answer A: ... Answer B: ... ...... Answer G: ...

Re-retrieval Rank-before Rank-after

Invalid question

Optimization

Retrieval-guided reward

Group compute

Figure 4: Overview of RAVEL’s three-stage training pipeline. A retriever is trained first, a questioner is initialized with supervised cold-start data, and retrieval-aware online question learning uses the current Top-4 candidates, a validity gate, witness answers, re-retrieval, and rank-guided policy updates.

component is ∆RR =

1 rafter

−

1 rbefore

.

(8)

Because target ranks are positive integers, no numerical rank offset is needed. With V denoting the set of valid questions, the reward is  R(q) =

∆RR + 0.351[rbefore > 1, rafter = 1] − 0.021[repeat], q ∈ V, −0.2, q∈ / V.

(9)

The 0.35 Rank-1 entry bonus rewards successful completion of the retrieval objective, while the repetition and invalid-action terms preserve question diversity and the interaction protocol. The remaining policy-normalization and optimization equations are given in Supplementary Material B. The gate is applied during both training and test-time rollouts. It enforces one answerable witnessinteraction protocol for every method and prevents candidate-selection, ranking, self-filled-answer, or memory-dump language from entering the retrieval text. Such output would expose information unavailable in a fair interaction and create a shortcut that distorts the ranking comparison. Empty questions are additionally excluded from the policy-loss update, whereas other invalid questions remain in the group for advantage normalization but receive no retrieval credit. The complete invalidaction taxonomy and trigger rules are given in Supplementary Material B. The main experiment uses the curated 3K-state configuration. The design principle is to keep the reward tied to retrieval utility. We do not directly reward a predefined question category, question length, or explicit visual terminology. This is important because encouraging a single manually chosen “good” template can reduce exploration and may not generalize across rounds. For each state, the policy samples a group of G questions and normalizes their rewards within the group. The resulting advantage, importance ratio, reference-policy penalty, and clipped objective are specified in Supplementary Material B and are applied only to generated question tokens. The questioner is optimized with a QLoRA (Dettmers et al., 2023) adapter while the vision tower, answerer, and retriever remain frozen; the selector is removed from the RAVEL pipeline. The SFT checkpoint serves as the reference policy, and the loss is masked to generated question tokens. Supplementary Material B lists the complete implementation settings and invalid-action rules. 6

Table 1: Interactive retrieval performance on Interactive-PEDES. RAVEL achieves the strongest retrieval results after both three and five rounds and the lowest BRI. “-” denotes an unavailable or inapplicable metric. Method

R@1

Round 3 R@5 R@10

mAP

R@1

Round 5 R@5 R@10

mAP

Initial PlugIR (Lee et al., 2024) ChatIR (Levy et al., 2023) GPT-5.6 Luna (OpenAI, 2026) SimRV (Liang & Albanie, 2023) LLaVA-ReID (Lu et al., 2025)

37.61 43.15 45.02 46.37 62.91 60.11

61.74 67.17 69.24 67.73 83.26 81.41

72.01 76.78 78.60 77.54 89.18 88.31

28.23 30.92 33.07 35.43 38.65 40.49

37.61 47.15 49.06 47.88 63.42 67.94

61.74 70.83 72.94 69.27 84.21 86.82

72.01 79.73 81.57 78.19 89.34 92.47

28.23 33.80 35.85 36.56 39.46 45.16

1.012 0.997 1.042 0.720 0.703

RAVEL (ours)

64.13

84.28

90.40

42.81

73.73

90.53

94.98

47.89

0.642

5

E XPERIMENTS

5.1

E XPERIMENTAL S ETUP

BRI ↓

We use CLIP (Radford et al., 2021) with IRRA (Jiang & Ye, 2023) as the frozen retriever, initialize the questioner from LLaVA-OneVision-Qwen2-7B-ov with QLoRA (Dettmers et al., 2023), and use Qwen2.5-7B-Instruct (Yang et al., 2024) as the frozen answerer. RAVEL directly receives the current Top-4 results and learns from 3,000 training states sampled across rounds and source datasets. We evaluate five interaction rounds on the 7,373-query Interactive-PEDES test split and report Rank-1, Rank-5, Rank-10, mAP, and BRI (Lee et al., 2024). BRI (Best log Rank Integral) summarizes retrieval quality over the interaction trajectory, with lower values indicating better cumulative retrieval. All methods share the same answer and text-cleaning protocol; implementation details are in Supplementary Material B. 5.2

I NTERACTIVE R ETRIEVAL AND Q UESTIONING B EHAVIOR

We compare RAVEL with PlugIR (Lee et al., 2024), ChatIR (Levy et al., 2023), SimRV (Liang & Albanie, 2023), LLaVA-ReID (Lu et al., 2025), GPT-5.6 Luna (OpenAI, 2026), and the nointeraction Initial setting under the same evaluation protocol. Table 1 shows that RAVEL achieves the strongest retrieval results at both evaluation rounds. At Round 5, it reaches 73.73 R@1 and reduces BRI to 0.642, improving the standard selector-based LLaVA-ReID pipeline by 5.79 Rank-1 points. The matched input and retriever comparisons are reported in Table 4. RAVEL improves over Initial as dialogue accumulates. The following analysis examines the question allocation and retrieval evidence learned by the policy. Closed-source vision-language model. GPT-5.6 Luna (OpenAI, 2026) follows the same candidate, answerer, text-cleaning, and evaluation protocol. After five rounds, it reaches 47.88 Rank-1, 36.56 mAP, and 1.042 BRI; prompt and decoding details are provided in Supplementary Material B. Questioning behavior and retrieval utility. RAVEL improves retrieval by learning a statedependent allocation of the five question turns. Supplementary Figure 6 shows a consistent shift toward localized attributes, especially hair and head cues, while Figure 5 quantifies the corresponding question utility. The question-type distribution in Figure 5(a) and the upper-left retrieval-gain panel in Figure 5(b) jointly explain the improvement: the share of local WH/open questions declines across rounds for both methods, but remains consistently higher for RAVEL, ending at 54% versus 31% for LLaVA-ReID. This relative advantage matters because local WH/open questions yield larger mean retrieval gains than local yes/no questions for both methods: 6.90 vs. 2.25 for RAVEL and 6.48 vs. 2.64 for LLaVA-ReID. RAVEL therefore improves by allocating more turns to the question type that produces stronger retrieval evidence, then using targeted verification to refine it. The resulting retrieval text becomes shorter while retaining more useful evidence. Among 963 cases that finish outside Rank-1 with LLaVA-ReID but reach Rank-1 with RAVEL, the five-round 7

RAVEL

LLaVA-ReID Mean retrieval gain

RAVEL

8.0

Mean Δrank

85 75

4.0

45

Five-round query length 96.9 90.8

80 60 40 20 0

Local WH/open

Final query

Local yes/no

Negative-word count 5

1

2

3

4

Mean count

35 25

2.64 2.25

2.0 0.0

55

100

6.0

5

Interaction round

Attribute coverage 82 73

100

4

80

3.23

3

1.91

2

%

Ratio (%)

65

6.48

6.90

Words

LLaVA-ReID

1

60 40 20

0

0

Final round

(a)

Final round

(b)

Figure 5: Behavioral comparison of LLaVA-ReID and RAVEL. (a) The ratio of local WH/open questions across interaction rounds. (b) Five-round retrieval gain, query length, negative-word count, and attribute coverage. description decreases from 96.9 to 90.8 words, while the mean negative-word count drops from 3.23 to 1.91. Attribute coverage increases from 73% to 82%; the largest additions concern shoes (+10.8 points), environment (+9.6), lower-body clothing (+5.5), and hair (+4.6). Fewer negative and other noisy words reduce interference and concentrate the retrieval text on positive, visually grounded attributes, helping explain how RAVEL converts the interaction budget into more effective evidence. 5.3

T RANSFER TO T EXT- BASED R E ID

Table 2: Comparison of text-based ReID retrievers and interactive questioners on three benchmarks. Underlined values indicate the second-best result in each column. “-” denotes an unavailable or inapplicable metric. Method

R@1

CUHK-PEDES R@5 R@10

mAP

R@1

ICFG-PEDES R@5 R@10

mAP

R@1

RSTPReid R@5 R@10

mAP

CFine (Yan et al., 2023) RaSa (Bai et al., 2023) APTM (Yang et al., 2023) AUL (Li et al., 2024b) RDE (Qin et al., 2024) IRRA (Jiang & Ye, 2023)

69.57 76.51 76.53 77.23 76.20 73.44

85.93 76.51 90.04 90.43 90.53 89.36

91.15 94.25 94.15 94.41 94.20 93.34

69.38 66.91 67.85 66.09

60.83 65.28 68.51 69.16 67.83 63.57

76.55 80.40 82.99 83.32 82.50 80.36

82.42 85.12 87.56 88.37 87.28 85.78

41.29 41.22 40.79 38.17

50.55 66.90 67.50 71.65 67.40 59.30

72.50 86.50 85.70 87.55 85.65 81.50

81.60 91.35 91.45 92.05 90.60 88.50

52.31 52.56 52.01 47.69

IRRA (Jiang & Ye, 2023) + LLaVA-ReID (Lu et al., 2025) IRRA (Jiang & Ye, 2023) + RAVEL

78.51 80.65

92.43 94.35

95.67 97.86

70.61 72.49

67.44 69.51

82.91 84.75

87.69 89.80

40.60 42.55

69.85 72.03

88.10 90.01

92.55 94.68

54.92 56.85

We integrate RAVEL with existing text-based ReID frameworks and evaluate transferability on CUHK-PEDES (Li et al., 2017), ICFG-PEDES (Ding et al., 2021), and RSTPReid (Zhu et al., 2021). The comparison covers CLIP-driven fine-grained matching (CFine) (Yan et al., 2023), relation and sensitivity-aware representation learning (RaSa) (Bai et al., 2023), large-scale multi-attribute pretraining (APTM) (Yang et al., 2023), and adaptive uncertainty-based learning (AUL) (Li et al., 2024b). Dataset annotations provide the initial description, and the questioner conducts five rounds of interaction to refine the retrieval text. The base T-ReID model encodes the initial description, while the interactive retriever encodes the accumulated dialogue; their matching scores are averaged for final re-ranking. Table 2 reports the cross-dataset transfer results, placing RDE (Qin et al., 2024) and IRRA (Jiang & Ye, 2023) alongside other conventional text-based ReID methods and comparing IRRA (Jiang & Ye, 2023) with its interactive counterparts. The RDE (Qin et al., 2024) retriever substitution is evaluated 8

Table 3: Ablation of the training procedure. Model

SFT RL Gate R@1 R@5 R@10 mAP

LLaVA-OV LLaVA-OV ✓ LLaVA-OV LLaVA-OV ✓ LLaVA-OV ✓

✓ ✓ ✓

✓ ✓

Table 4: Ablation of candidate input and retriever choice after five interaction rounds.

BRI

44.12 68.24 77.73 32.55 1.075 69.44 87.78 92.80 45.48 0.698 65.66 84.66 90.21 40.30 0.698 Collapse 73.73 90.53 94.98 47.89 0.642

Configuration

R@1

Candidate input LLaVA-ReID (selector) LLaVA-ReID (direct Top-4) RAVEL (direct Top-4)

R@5 R@10

mAP BRI ↓

67.94 86.82 67.84 87.54 73.73 90.53

92.47 45.16 92.72 45.33 94.98 47.89

0.703 0.701 0.642

Retriever RAVEL + IRRA (Jiang & Ye, 2023) 73.73 90.53 RAVEL + RDE (Qin et al., 2024) 77.53 92.08

94.98 47.89 95.56 49.06

0.642 0.614

separately in Table 4. Using IRRA (Jiang & Ye, 2023) as the retriever, RAVEL raises Rank-1 by 7.21 points on CUHK-PEDES, 5.94 points on ICFG-PEDES, and 12.73 points on RSTPReid. Its mAP gains over the IRRA (Jiang & Ye, 2023) backbone are 6.40, 4.38, and 9.16 points, respectively. The consistent gains across three benchmarks indicate that retrieval-aware question learning complements standard cross-modal retrieval models beyond Interactive-PEDES. 5.4

A BLATION S TUDY

We study the training recipe and input/retriever configuration under the same four-candidate, answer-length, and cleaning settings. The RL state-pool scaling study is reported in Supplementary Material B. Training-procedure ablation. Table 3 separates the native questioner, SFT-only, RL-only, and SFT-to-RL paths. Rank-1 rises from 44.12 for the native model to 69.44 with SFT, 65.66 with RL from the native initialization, and 73.73 with SFT followed by gated RL. Removing the gate causes a rapid invalid-action collapse: the policy starts emitting candidate-person references that prompt broad “candidate person” descriptions, together with too-short or non-question outputs, and these patterns dominate the rollouts as training proceeds. This behavior exposes a rank-only reward loophole, whereas the gate keeps the policy focused on answerable questions. Candidate input and retriever ablation. Table 4 isolates candidate access and retriever choice. The RDE (Qin et al., 2024) row replaces frozen IRRA (Jiang & Ye, 2023) after RAVEL training, while the questioner and answerer remain unchanged. Direct Top-4 input alone changes little for LLaVA-ReID, whereas pairing it with RAVEL’s online learning yields stronger gains, showing that the improvement comes from learning retrieval-useful questions. Replacing IRRA with RDE after training further improves retrieval without updating RAVEL, indicating transferability across retriever backbones. 5.5

Q UALITATIVE A NALYSIS

We select two cases by baseline final rank: one near the gallery front and one long-tail recovery case. Supplementary Material C gives the five-round records. RAVEL replaces repeated verification with localized attribute questions and improves the target rank through sequential evidence acquisition. At inference, RAVEL receives no ground-truth rank, target image, or oracle attribute list. The cases show how question timing and local attribute choice shape the final ranking.

6

C ONCLUSION

RAVEL treats interactive person ReID as sequential information acquisition and learns retrievalaware questions from online rank feedback. Our analysis shows that localized open-ended questions provide more useful evidence on difficult queries. RAVEL reaches 73.73 Rank-1 after five rounds and improves IRRA (Jiang & Ye, 2023) across three text-based ReID benchmarks. 9

R EFERENCES Yang Bai, Min Cao, Daming Gao, Ziqiang Cao, Chen Chen, Zhenfeng Fan, Liqiang Nie, and Min Zhang. Rasa: Relation and sensitivity aware representation learning for text-based person search. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, pp. 555–563, 2023. Rahul Atul Bhope, K. R. Jayaram, Vinod Muthusamy, Ritesh Kumar, Vatche Isahagian, and Nalini Venkatasubramanian. Voila: Value-of-information guided fidelity selection for cost-aware multimodal question answering. arXiv preprint arXiv:2602.03007, 2026. Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. In Advances in Neural Information Processing Systems, volume 36, 2023. Zefeng Ding, Changxing Ding, Zhiyin Shao, and Dacheng Tao. Semantically self-aligned network for text-to-image part-aware person re-identification. arXiv preprint arXiv:2107.12666, 2021. Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Schmidt Feris. Dialog-based interactive image retrieval. arXiv preprint arXiv:1805.00145, 2018. Hulingxiao He, Zijun Geng, and Yuxin Peng. Fine-r1: Make multi-modal llms excel in fine-grained visual recognition by chain-of-thought reasoning. In International Conference on Learning Representations, 2026. Ding Jiang and Mang Ye. Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2787–2797, 2023. Pengcheng Jiang, Jiacheng Lin, Lang Cao, Runchu Tian, SeongKu Kang, Zifeng Wang, Jimeng Sun, and Jiawei Han. Deepretrieval: Hacking real search engines and retrievers with large language models via reinforcement learning. In Proceedings of the Conference on Language Modeling, 2025. Adriana Kovashka, Devi Parikh, and Kristen Grauman. Whittlesearch: Interactive image search with relative attribute feedback. International Journal of Computer Vision, 115(2):185–210, 2015. Saehyung Lee, Sangwon Yu, Junsung Park, Jihun Yi, and Sungroh Yoon. Interactive text-to-image retrieval with large language models: A plug-and-play approach. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pp. 791–809, 2024. Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischinski. Chatting makes perfect: Chat-based image retrieval. In Advances in Neural Information Processing Systems, volume 36, 2023. Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024a. Mingxin Li, Yanzhao Zhang, Dingkun Long, Keqin Chen, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, Jingren Zhou, and Junyang Lin. Qwen3-vl-embedding and qwen3-vl-reranker: A unified framework for state-of-the-art multimodal retrieval and ranking. arXiv preprint arXiv:2601.04720, 2026. Shenshen Li, Chen He, Xing Xu, Fumin Shen, Yang Yang, and Heng Tao Shen. Adaptive uncertainty-based learning for text-based person retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 3172–3180, 2024b. Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang. Person search with natural language description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1970–1979, 2017. Kaiqu Liang and Samuel Albanie. Simple baselines for interactive video retrieval with questions and answers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 11091–11101, 2023. 10

Zhuofan Lou, Shihang Zhang, Fangle Zhu, Shengjie Ye, and Pingyu Wang. Uncertainty-aware pedestrian attribute recognition via evidential deep learning. arXiv preprint arXiv:2604.26873, 2026. Yiding Lu, Mouxing Yang, Dezhong Peng, Peng Hu, Yijie Lin, and Xi Peng. Llava-reid: Selective multi-image questioner for interactive person re-identification. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pp. 40868–40887, 2025. Avinash Madasu, Jun Oliva, and Gedas Bertasius. Learning to retrieve videos by asking questions. In Proceedings of the 30th ACM International Conference on Multimedia, pp. 356–365, 2022. OpenAI. Gpt-5.6 luna model. OpenAI API Documentation, 2026. URL https:// developers.openai.com/api/docs/models/gpt-5.6-luna. Ruiying Peng, Xueyu Wu, Jing Lei, Lu Hou, Yuanzheng Ma, and Xiaohui Li. Deeper thought, weaker aim: Understanding and mitigating perceptual impairment during reasoning in multimodal large language models. arXiv preprint arXiv:2603.14184, 2026. Yang Qin, Yingke Chen, Dezhong Peng, Xi Peng, Joey Tianyi Zhou, and Peng Hu. Noisycorrespondence learning for text-to-image person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 27197–27206, 2024. Yang Qin, Chao Chen, Zhihang Fu, Dezhong Peng, Xi Peng, and Peng Hu. Human-centered interactive learning via mllms for text-to-image person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14390–14399, 2025. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, pp. 8748–8763, 2021. Chris Samarinas, Haw-Shiuan Chang, and Hamed Zamani. Truncated step-level sampling with process rewards for retrieval-augmented reasoning. arXiv preprint arXiv:2602.23440, 2026. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, Daya Guo, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. Jintao Sun, Zhedong Zheng, and Gangyi Ding. Harnessing weak pair uncertainty for text-based person search. arXiv preprint arXiv:2604.08877, 2026. Zhao Wang, Ziliang Zhao, and Zhicheng Dou. Prorag: Process-supervised reinforcement learning for retrieval-augmented generation. arXiv preprint arXiv:2601.21912, 2026. Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong, Yutong Cai, Siyuan Li, Huijia Zhu, Weiqiang Wang, Linghe Kong, Yue Wang, Zhuosheng Zhang, and Weiran Huang. Zooming without zooming: Region-to-image distillation for fine-grained multimodal perception. arXiv preprint arXiv:2602.11858, 2026. Shuanglin Yan, Neng Dong, Liyan Zhang, and Jinhui Tang. Clip-driven fine-grained text-image person re-identification. IEEE Transactions on Image Processing, 32:6032–6046, 2023. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024. Shuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang, Li Zhu, and Yujiao Wu. Towards unified text-based person retrieval: A large-scale multi-attribute and language search benchmark. In Proceedings of the 31st ACM International Conference on Multimedia, pp. 4492–4501, 2023. Qianhao Yuan, Jie Lou, Xing Yu, Hongyu Lin, Le Sun, Xianpei Han, and Yaojie Lu. Vision-opd: Learning to see fine details for multimodal llms via on-policy self-distillation. arXiv preprint arXiv:2605.18740, 2026. 11

Quan Zhang, Jingze Wu, Jialong Wang, Xiaohua Xie, Jianhuang Lai, and Hongbo Chen. Thinking before matching: A reinforcement reasoning paradigm towards general person re-identification. arXiv preprint arXiv:2604.19218, 2026. Aichun Zhu, Zijie Wang, Yifeng Li, Xili Wan, Jing Jin, Tian Wang, et al. Dssl: Deep surroundingsperson separation learning for text-based person retrieval. In Proceedings of the 29th ACM International Conference on Multimedia, pp. 209–217, 2021. Hainan Zhu, Jie Huang, Stevan Rudinac, and Evangelos Kanoulas. Enhancing interactive image retrieval with query rewriting using large language models and vision language models. In Proceedings of the 2024 International Conference on Multimedia Retrieval, pp. 978–987, 2024. Muzhi Zhu, Hao Zhong, Canyu Zhao, Zongze Du, Mingyu Liu, Zheng Huang, Anzhou Li, Hao Chen, Cheng Zou, Jingdong Chen, Ming Yang, and Chunhua Shen. Active-o3: Empowering mllms with active perception via pure reinforcement learning. arXiv preprint arXiv:2505.21457, 2026.

12

A. Q UESTION -VALUE D IAGNOSTICS Section 3.3 of the main paper defines the teacher-forced likelihood and the five internal visualreliance diagnostics. This section records the sampling protocol, statistical estimators, and annotation details used to reproduce the analysis. Table 5: Internal visual-reliance diagnostics of the LLaVA-ReID questioner on InteractivePEDES. IAS, CS, and VRR measure visual conditioning; ∆R1 and ∆Shuffle measure sensitivity to candidate retention and ordering. Definitions are given in Section 3.3. Round IAS ↑ 0 1 2 3 4

CS ↑

Table 6: Per-round question-quality statistics on the 2,000-state model-annotated sample. d+ and u+ are the positive rates of the model’s discriminativeness and usefulness judgments; ρd and ρu are Spearman correlations with reciprocal-rank gain; rd and ru are pointbiserial correlations with rank improvement.

∆R1 ∆Shuffle VRR ↑

0.0272 0.0042 -0.0000 0.0369 0.0066 0.0001 0.0460 0.0080 -0.0011 0.0506 0.0090 -0.0004 0.0542 0.0091 0.0002

Round

-0.0115 0.0631 -0.0168 0.0869 -0.0199 0.0945 -0.0247 0.0965 -0.0251 0.0970

0 1 2 3 4

n Improve 400 400 400 400 400

d+

u+

ρd

ρu

rd

ru

25.00% 41.00% 39.00% 0.034 0.008 0.082 0.047 25.00% 38.00% 37.75% 0.017 0.012 0.059 0.051 25.00% 34.50% 34.25% 0.088 0.080 0.140 0.131 25.00% 32.75% 32.25% 0.080 0.087 0.151 0.157 25.00% 30.50% 30.00% 0.083 0.094 0.107 0.113

The offline-supervision diagnostic contains 2,000 interaction states sampled from observed testset rollouts. The sample is balanced across the five interaction rounds (400 states per round) and stratified by outcome: 500 rank-improvement states and 1,500 non-improvement states. Each state stores the current Top-4 candidates, dialogue history, question, answer, and rank change. Qwen3VL-32B receives the state and returns structured judgments for candidate discrimination, shared attributes, absent or invisible attributes, generic background content, negative or unknown answers, redundancy with history, visual uncertainty, and likely retrieval usefulness. We use 101 humanannotated cases for the human–model agreement analysis. For each state i, let di ∈ {0, 1} denote the model’s discriminative-attribute judgment, ui ∈ {0, 1} its likely-useful judgment, ∆RR,i = 1/riafter − 1/ribefore the reciprocal-rank gain, and zi = 1[∆RR,i > 0] the rank-improvement indicator. We compute Spearman’s rank correlation between each binary judgment and ∆RR by applying average ranks to tied values: ρs (x, ∆RR ) = Corr(rank(x), rank(∆RR )) . For the binary improvement event, we use the point-biserial correlation: r x̄1 − x̄0 n1 n0 rpb (x, z) = , sx n2

(10)

(11)

where x̄1 and x̄0 are the means of x among improved and non-improved states, sx is the sample standard deviation, and n1 , n0 are the corresponding counts. For the human labels, ”useless,” ”partial,” and ”useful” are mapped to 0, 0.5, and 1 for continuous analyses. For agreement, ”partial” and ”useful” are grouped as a positive label. With hi and mi denoting the resulting human and model binary labels, Cohen’s kappa is κ=

po − p e , 1 − pe

po =

1X 1[hi = mi ], n i

(12)

where pe is the chance agreement obtained from the two annotators’ marginal positive and negative rates. On the 2,000-state sample, the overall values are ρs (d, ∆RR ) = 0.051, ρs (u, ∆RR ) = 0.048, rpb (d, z) = 0.107, and rpb (u, z) = 0.099. For the 101 human-annotated cases, mapping “useless,” “partial,” and “useful” to 0, 0.5, and 1 gives ρs (h, ∆RR ) = −0.016 (p = 0.875); grouping “partial” and “useful” as positive gives rpb (h, z) = −0.041 (p = 0.687). The paired human–model audit gives po = 0.654 and κ = 0.312 for 101 completed labels. 13

Table 7: Prompt used by the Qwen3-VL question-quality evaluator. The evaluator judges a question– answer step from the ordered Top-4 candidates and dialogue context without access to the observed rank change. Prompt 1: Qwen3-VL Question-Quality Evaluator You are auditing a question-answer step in an interactive person re-identification system. The four images are ordered candidates for the target person. Image 1 is the current top-ranked candidate; Images 2-4 are similar candidates. Judge the question and the witness answer using the images and context. Focus on whether the answer creates evidence that can distinguish the target from the candidates and improve retrieval. Do not use the known rank change as evidence for your judgment. Initial description: {initial_query} Previous questions: {history_questions} Current question: {question} Witness answer: {answer} Return exactly one JSON object and no markdown. Use boolean values for all boolean fields. Required fields: - discriminative attribute: the question-answer pair distinguishes at least some candidates using a visible attribute. - shared attribute: the answered attribute is visibly shared by most or all candidates. - absent or invisible attribute: the attribute is absent, occluded, too small, or not reliably visible. - generic background: the question mainly asks about generic scene/background information. - negative or unknown answer: the answer is negative, unknown, or does not provide usable visual evidence. - redundant with history: the question repeats an already asked attribute or intent. - uncertain visual evidence: the visual evidence is ambiguous or unreliable. - likely useful for retrieval: the complete question-answer pair is likely to help distinguish the target in retrieval. - main attribute: one short category such as upper clothing, lower clothing, shoes, hair, accessory, bag or carried item, body or posture, background, or other. - reason: one concise sentence grounded in the four images and answer.

14

B. I MPLEMENTATION D ETAILS Retriever. We use CLIP-ViT-B/16 with the IRRA training framework. The retriever is trained for 30 epochs on fine-grained Interactive-PEDES descriptions with batch size 128, and is frozen for questioner training and evaluation. The text encoder uses the extended positional-embedding setting required by the long descriptions. Questioner. The questioner is initialized from LLaVA-OneVision-Qwen2-7B-ov. SFT uses QLoRA with rank 128, alpha 256, dropout 0.05, 4-bit NF4 quantization with double quantization, bf16 computation, learning rate 1 × 10−5 , batch size 4, gradient accumulation 4, one epoch, cosine scheduling, 2% warm-up, weight decay 0, and a maximum sequence length of 4096. The question length is limited to 96 tokens during interaction. Answerer and interaction. The answerer is Qwen2.5-7B-Instruct and remains frozen. Its answer length is limited to 64 tokens. The original answer-cleaning rule and the fixed five-round interaction budget are used in all primary comparisons. The LLaVA-ReID baseline uses its original selector, whereas RAVEL directly passes the current Top-4 candidates to the questioner and does not use a selector. For cross-dataset transfer on CUHK-PEDES, ICFG-PEDES, and RSTPReid, the simulated answerer receives the ground-truth target image as witness memory, but no target caption, identity label, rank, or other target-side metadata. The original dataset caption initializes retrieval, while the trained questioner and retriever are reused without retraining. Reinforcement learning (RL) state construction. The Interactive-PEDES training split contains 47,376 images and 11,543 identities. We construct the RL state pool from existing multi-round interaction records, sampling only training states whose target is not already Rank-1. The pool is stratified across interaction rounds and preserves the original proportions of the two source datasets, so that the policy observes different dialogue-history lengths and retrieval stages without introducing split or source imbalance. The primary experiment uses a curated 3,000-state pool; 1K and 5K pools use the same construction for controlled scaling analysis. Table 8: Primary-run scale, action validity, and compute cost. Category

Value

RL training states Question rollouts Optimizer updates Invalid-gate trigger rate Valid-action rate Mean unique questions / group Mean repetition rate

3,000 24,000 (8 per state) 1,500 2.02% 97.98% 7.93 / 8 9.20%

Training hardware Training wall-clock time Training compute Five-round evaluation time Evaluation compute

2 × NVIDIA A800 13 h 15 min 26.5 GPU-hours 8 h 19 min 16.6 GPU-hours

Table 9: Scaling study of the curated RL state pool. The question policy and reward are fixed across pool sizes. RL states

R@1

R@5 R@10

mAP

BRI

0K (SFT-only) 1K 3K 5K

69.44 69.46 73.73 74.35

87.78 87.97 90.53 90.71

45.48 45.59 47.89 48.09

0.698 0.689 0.642 0.635

92.80 93.03 94.98 94.89

RL optimization and cost. RAVEL starts from the SFT checkpoint and updates only the QLoRA adapter. The primary study uses the curated 3,000-state pool, group size G = 8, token-level clipped group-relative policy updates, the SFT checkpoint as the reference policy, and the K3 token-level KL estimator. We use AdamW with learning rate 1 × 10−6 , constant scheduling (no warm-up or decay), zero weight decay, and gradient-norm clipping at 1.0. Gradients accumulate over two state-level groups, so each optimizer update uses 16 sampled questions from two states; one epoch over 3,000 states therefore gives 1,500 optimizer updates and 24,000 question rollouts. During RL rollouts, generation uses sampling with temperature 1.0, top-p = 0.95, no top-k or beam search, and at most 100 new tokens. During test-time RAVEL evaluation, sampling uses temperature 1.0 (the model default), top-p = 0.5, no top-k or beam search, and at most 100 new tokens. The selector is not used during RAVEL training or inference; the policy receives the current Top-4 candidates directly. 15

The complete reward, objective, and state-construction details are given below, while invalid-action rules are listed in the following subsection. The main paper defines the reciprocal-rank reward and validity-gated reward function. For completeness, the remaining policy quantities are specified below. For completeness, for a sampled group with rewards r1 , . . . , rG , we use the normalized advantage Ai =

ri − µG . max(σG , εadv )

(13)

Here εadv = 10−6 is the advantage-normalization floor. For generated token yi,t , the importance ratio is ρi,t = exp(log πθ (yi,t | xi , yi,<t ) − log πold (yi,t | xi , yi,<t )) , (14) and the K3 reference penalty is dK3 i,t = exp(δi,t ) − δi,t − 1,

δi,t = log πref − log πθ .

The optimized objective is 1 X 1 X K3 1 X 1 X min(ρi,t Ai , clip(ρi,t , 1 − εclip , 1 + εclip )Ai ) + β d . L=− G i Ti t G i Ti t i,t

(15)

(16)

The policy-clipping parameter is εclip = 0.2, and the K3 KL penalty coefficient is β = 0.03. The reciprocal-rank reward in the main paper has no additional rank-offset epsilon. All sums are masked to generated question tokens; prompt, dialogue, and image-token positions do not contribute to the policy loss. Closed-source Luna baseline. GPT-5.6 Luna is used only as an API questioner. It receives the ordered Top-4 candidate images and dialogue context, while the retriever, answerer, invalidquestion gate, and five-round evaluation protocol remain identical to the main comparison. We use reasoning effort=none, temperature 1.0, top-p 0.5, max completion tokens=100, up to 16 concurrent requests, and seed 42. The complete prompt and input format are given in Table 10. Invalid-question rules. The validity gate is applied to a normalized, lower-cased question before retrieval evaluation in both training and test-time rollouts. It enforces a common answerable witness protocol: candidate indices, ranking instructions, self-filled answers, and memory-dump requests cannot enter the retrieval text. The deterministic checks are applied in the following order. Surface form. Empty output, fewer than five words, or no question mark. Selection and self-filling. Candidate/image or ranking references, self-filled “conversation” or “answer” content, and unconditional requests for a complete memory or description. Broad appearance and repetition. Whole-person prompts or broad requests for additional details without a concrete local target such as clothing, bags, shoes, hair, accessories, body regions, colors, environment, setting, or background are rejected; token-set Jaccard similarity of at least 0.70 with a previous dialogue question is also rejected, while repetition within a sampled group is penalized separately. State sampling details. A state is identified by its query and interaction round. We remove states whose target is already Rank-1 before the action, sample across all valid rounds, and prevent duplicate states within a training pool. The source-dataset allocation is fixed to the proportions of Interactive-PEDES before random sampling. The 3K primary pool and 5K scaling pool use the same balanced construction procedure; only the number of sampled states changes. Evaluation. We use the 7,373-query Interactive-PEDES test split and a fixed gallery. We report Rank-1, Rank-5, Rank-10, mAP, and BRI after five rounds. All reported runs use fixed checkpoint paths, decoding parameters, and random seeds; the main comparison uses seed 42 for inference. Question-type classification. Figure 5 uses a deterministic lexical classifier rather than an LLM judge or manual labels. Each generated question is lower-cased, stripped, and flattened across line breaks. The rules are applied in this order: (i) a multiple-choice marker such as “A)”–“E)” yields the 16

multiple-choice class; (ii) a question beginning with “is,” “are,” “was,” “were,” “does,” “do,” “did,” “has,” “have,” or “had” yields the yes/no class; (iii) broad prompts containing phrases such as “any additional details,” “anything else,” or “appearance or surroundings” yield the broad-open class; (iv) a question beginning with “what,” “which,” “where,” or “how,” or a request to describe/provide/tell an attribute such as clothing, shoes, bags, hair, hats, accessories, posture, or environment, yields the local-open class; and (v) all remaining questions yield the other class. Figure 5 aggregates the local-open bucket as local WH/open and the yes/no bucket as local yes/no. The classifier uses only the question text, not the answer, rank change, or reward, so the reported type proportions and conditional gains do not use downstream outcomes as labels. Table 10: Prompt used by GPT-5.6 Luna. Prompt 2: GPT-5.6 Luna Vision Questioner SYSTEM MESSAGE You are an advanced question generator designed to assist in identifying a specific individual from a collection of candidates. USER MESSAGE [Candidate Person] <Rank-1 candidate image> [Similar Person #1] <Rank-2 candidate image> [Similar Person #2] <Rank-3 candidate image> [Similar Person #3] <Rank-4 candidate image> [Conversation] Witness: {initial description} You: {question from round 1} Witness: {answer from round 1} You: {question from round 2} Witness: {answer from round 2} ... Based on these information: [Candidate Person] The closest match to the provided description and answers. [Similar Persons] Individuals who also match but are less similar. [Conversation] The initial description from the witness and the ongoing dialogue. generate a question that could help differentiate the candidate from the similar individuals. Avoid repeating questions that already answered with "I don't know."

Baseline Baseline

RAVEL (RL) RAVEL (RL)

(a) LLaVA-ReID

(b) RAVEL

Figure 6: Attribute shares before and after retrieval-aware training. RAVEL reallocates the interaction budget toward localized attribute questions while reducing broad clothing, posture, and environment questions.

17

C. Q UALITATIVE C ASES To illustrate how retrieval-aware question learning changes the interaction trajectory, we present two representative five-round cases. Both methods receive the same initial description and candidate state; each panel then records the Top-4 candidates, the generated question, the witness answer, and the rank transition. The first case starts at Rank 5 and ends at Rank 3 for LLaVA-ReID versus Rank 1 for RAVEL, while the more difficult second case starts at Rank 52 and ends at Rank 85 versus Rank 1, respectively. Initial Description Initial Description

He is wearing a brown padded jacket with a fur collar and dark jeans.

He is wearing a brown padded jacket with a fur collar and dark jeans.

RAVEL | final rank 1

LLaVA-ReID | final rank 3

Round 1

Round 2 Rank: 6→3

Round 2

Rank: 5→6

Q: Is the man standing with a relaxed posture, slightly leaning forward?

Q: Is the man's hair long and light colored?

A: The man is standing with a relaxed posture, but there is no mention of him leaning forward in my observation.

A: No, the man's hair is dark and short.

A: The man is standing with a relaxed posture, but there is Rank: no mention of him leaning forward in my observation. 3→3

Rank: 2→2

Q: Is the man standing in a quiet, rural area?

Q: Is the man Round 4 walking with a relaxed posture?

A: No, the man is not standing in a busy city street. The

surroundings appear to be an indoor space, likely a public Rank: area such as a mall or an airport. 3→4

Q: Is the man wearing a jacket with a lot of logos and patterns?

Q: Is the man walking with a relaxed posture?

Q: Is the man standing in a busy city

A: The man is standing with a relaxed posture, not walking.

A: No, the man is not standing in a busy surroundings appear to be an indoor sp area such as a mall or an airport.

A: The man is not holding any visible items and has his hands placed in front of him.

Round 5

Round 5 Rank: 4→3

Rank: 1→1

He is wearing a brown padded jacket with a fur collar and dark jeans.

Rank: 5

Q: Is the man wearing a jacket with a lot of logos and patterns?

Q: What can you tell me about the man's accessories or items he is carrying?

A: No, the man is wearing a brown padded jacket with a distinct two-tone appearance but it does not have a lot of logos and patterns.

Target

al rank 3

A: The man is not holding any visible items and has his hands placed in front of him.

RAVEL | final rank 1 Round 2

Round 1 Rank: 6→3

th a relaxed posture, slightly leaning

Q: Is the man's hair long and light colored? A: No, the man's hair is dark and short.

a watch or carrying a bag?

ng any visible items, so he is not wearing g.

Round 2 Rank: 5→3

Rank: 3→2

Q: What kind of jacket is the man wearing?

Q: What can you tell me about the man's hair?

A: The man is wearing a brown padded jacket with a fur collar and a distinct two-tone appearance, featuring a lighter brown on the sleeves and a darker brown body.

Round 4

A: His hair is dark, short, and styled in a casual manner.

Round 3

Rank: 3→3

Rank: 3→4

Q: Is the man standing in a quiet, rural area? A: No, the man is not standing in a quiet, rural area; the surroundings appear to be an indoor space, likely a public area, such as a mall or an airport.

Round 4 Rank: 2→2

Rank: 2→1

Q: Is the man walking with a relaxed posture?

Q: Is the man standing in a busy city street?

A: The man is standing with a relaxed posture, not walking.

A: No, the man is not standing in a busy city street. The surroundings appear to be an indoor space, likely a public area such as a mall or an airport.

Round 5 Rank: 4→3

Round 4 Rank: 2→2

Rank:

1 → 1 area? Q: Is the man standing in a quiet, rural

A: No, the man is not standing in a quiet, rural area; the surroundings appear to be an indoor space, likely a Q: What can you tell me about the man's accessories or public area, such as a mall or an airport. items he is carrying?

A: The man is not holding any visible items, so he is not wearing a watch or carrying a bag.

A: No, the man is wearing a brown padded jacket with a distinct two-tone appearance but it does not have a lot of logos and patterns.

Initial Description

h a relaxed posture, but there is orward in my observation.

A: His hair is dark, short, and styled in a

Round 5 Rank:

Rank: 5→6

Q: What can you tell me about the m

A: The man is wearing a brown padded jacket with a fur collar and a distinct two-tone appearance, featuring a lighter brown Rank: 2 → 1a darker brown body. on the sleeves and

Q: Is the man standing in a Round busy city street? 3

A: The man is standing with a relaxed posture, not walking.

3→3

4→3 Q: Is the man wearing a watch or carrying a bag?

with a lot of logos and

Q: What kind of jacket is the man wearing? Round 4

A: No, the man's hair is dark and short.

A: No, the man is not standing in a quiet, rural area; the surroundings appear to be an indoor space, likely a Rank: public area, such as a mall or an airport.

Round 2 Rank: 5→3

A: His hair is dark, short, and styled in a casual manner.

Q: Is the man's hair long and light colored?

Rank: 3→4

Rank: 3→2

Q: What can you tell me about the man's hair?

Round 3

Round 5

padded jacket with a distinct ot have a lot of logos and

Round 1 Rank: 6→3

A: The man is wearing a brown padded jacket with a fur collar and a distinct two-tone appearance, featuring a lighter brown on the sleeves and a darker brown body.

Round 4

A: The man is not holding any visible items, so he is not wearing a watch or carrying a bag.

Round 2 Rank: 5→3

Q: What kind of jacket is the man wearing?

Q: Is the man standing with a relaxed posture, slightly leaning forward?

Q: Is the man wearing a watch or carrying a bag? Round 3

RAVEL | final rank 1

Round 1

Rank: Round 1 5→6

Round 3

Rank: 5

Target

Target

LLaVA-ReID | final rank 3

Rank: 5

Rank: 1→1

Q: What can you tell me about the man's accessories or items he is carrying? A: The man is not holding any visible items and has his hands placed in front of him.

Figure 7: Qualitative comparison on a case where the target starts at Rank 5. The initial description is followed by the LLaVA-ReID and RAVEL interaction trajectories.

18

Initial Description Initial Description

A man is wearing a dark jacket, jeans, and red sneakers while walking.

A man is wearing a dark jacket, jeans, and red sneakers while walking.

RAVEL | final rank 1

LLaVA-ReID | final rank 85

Round 1

Round 2 Rank: 90 → 124

Round 2

Rank: 52 → 90

Q: Is the man walking with a relaxed posture?

Q: Is the man's hair long and light colored?

A: Yes, the man is walking with a relaxed posture.

A: The man's hair is dark and medium-length, not long and light colored.

Round 1

Q: What color are the young man's jeans?

A: The man is not standing in a well-lit area with clear details of his surroundings, as the description provided indicates he is in an indoor setting with tiled flooring and informational signage in the background, but no specific mention of lighting conditions.

Round 4

Rank: 124 → 143

Q: Is the man standing in a well-lit area with clear details of his surroundings?

Round 4

3→2

Q: Is the man wearing a dark-colored shirt?

Q: What is the Round 4 man carrying on his shoulder?

A: Yes, the man is wearing a dark-colored jacket.

A: The man is carrying a backpack slung over one shoulder.

Rank: 124 → 143

Round 5

A: The man is not standing in a well-lit area with clear details of his surroundings, as the description provided indicates he is in an indoor Rank: and informational signage in the background, setting with tiled flooring 2 → 1 of lighting conditions. but no specific mention

Q: What can you tell me about the man's Round 3 hair? A: The man's hair is dark, medium-length, and slightly tousled.

Rank: 143 → 122

Initial Description

Q: What is the man carrying on his shoulder?

Q: What can you tell me about the m

A: Yes, the man is wearing a dark-colored jacket.

A: The man is carrying a backpack slung over one shoulder.

A: The man's hair is dark, medium-leng

A: Yes, the man's posture is relaxed as he walks.

Round 5

Round 5 Rank: 122 → 85

Rank: 1→1

A man is wearing a dark jacket, jeans, and red sneakers while walking.

Rank: 52

Q: Is the young man wearing blue jeans that fit him well?

Q: Is the man's posture relaxed as he walks?

A: The man is wearing medium blue jeans that have a casual fit, with some visible wear, including a distressed detail.

Target

al rank 85

A: Yes, the man's posture is relaxed as he walks.

RAVEL | final rank 1 Round 2

Round 1 Rank: 90 → 124

Q: Is the man's hair long and light colored? A: The man's hair is dark and medium-length, not long and light colored.

Round 2 Rank: 52 → 4

Rank: 4→3

Q: Is the man standing in a well-lit area with clear details of his surroundings?

A: The man is not standing in a well-lit area with clear details of his surroundings, as the description provided indicates he is in an indoor setting with tiled flooring and informational signage in the background, but no specific mention of lighting conditions.

Round 4 Rank: 143 → 122

Q: What color are the young man's jeans? A: The young man's jeans are medium blue.

Round 3

Rank: 124 → 143

Round 4 Rank: 3→2

1→1 Q: Is the man wearing a dark-colored shirt?

Q: Is the man's posture relaxed as he walks?

A: The man is wearing medium blue jeans that have a casual fit, with some visible wear, including a distressed detail.

ng with a relaxed posture.

A: The young man's jeans are medium

Rank:

A: No, the jacket the man is wearing does not have a Q: Is the young man wearing blue jeans that fit him well? logo on it. noticeable

with a relaxed posture?

Q: What color are the young man's je

Round 5 Rank:

122 → 85 Q: Is the young man wearing a jacket with a noticeable logo on it?

Rank: 52 → 90

Round 2 Rank: 52 → 4

A: The young man's jeans are medium blue.

Round A: The3man's hair is dark and medium-length, not long and light colored. Rank:

Rank: 143 → 122

Rank: 4→3

Rank: 90 → 124

Q: Is the man's hair long and light colored?

A: Yes, the man is walking with a relaxed posture.

A: No, the jacket the man is wearing does not have a noticeable logo on it.

Round 2 Rank: 52 → 4

Q: Is the man standing in a well-lit area with clear details of his surroundings?

Q: Is the man walking with a relaxed posture?

Q: Is the young man wearing a jacket with a noticeable logo on Round 3it?

RAVEL | final rank 1

Round 1

Rank: Round 52 → 90 1

Round 3

Rank: 52

Target

Target

LLaVA-ReID | final rank 85

Rank: 52

Round 4 Rank: 3→2

Rank: 2→1

ing a jacket with a noticeable logo on it?

Q: Is the man wearing a dark-colored shirt?

Q: What is the man carrying on his shoulder?

Q: What can you tell me about the man's hair?

an is wearing does not have a

A: Yes, the man is wearing a dark-colored jacket.

A: The man is carrying a backpack slung over one shoulder.

A: The man's hair is dark, medium-length, and slightly tousled.

Round 5 Rank: 122 → 85

Rank: 1→1

earing blue jeans that fit him well?

Q: Is the man's posture relaxed as he walks?

medium blue jeans that have a casual ar, including a distressed detail.

A: Yes, the man's posture is relaxed as he walks.

Figure 8: Qualitative comparison on a difficult case where the target starts at Rank 52. The initial description is followed by the LLaVA-ReID and RAVEL interaction trajectories.

19

Record · ID 1006890 · SHA-256 c1a40c04c46ea376
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.