ConceptioArchivearXiv CS
arXiv CSopen access

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Zihan Xu1* , Yanzhen Chen1* , Xiaocheng Zhang1* , Zhiting Fan1 , Weiqi Zhai2 , Hongxia Xu1,3 , and Zuozhu Liu1,3† 1

arXiv:2607.09322v1 [cs.AI] 10 Jul 2026

Zhejiang University, Hangzhou, China {zihan1.22,yanzhen.22,xiaocheng.22,zhiting.23}@intl.zju.edu.cn [email protected], [email protected] 2 Alibaba Group, Hangzhou, China [email protected] 3 Transvascular Implantation Devices Research Institute, Hangzhou, China

Abstract. In this work, we introduce LongMedBench, a real-world EHRbased benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized shortcontext knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model’s immediate context. Keywords: Computer-Aided Diagnosis · Medical Agents · EHR.

1 Introduction In clinical practice, the diagnostic process is inherently longitudinal and timedependent. Clinicians do not merely react to isolated symptoms; instead, they must synthesize evidence across multiple visits, diagnostic tests, and evolving * †

These authors contributed equally to the work. Corresponding author: Zuozhu Liu.

2

Z. Xu et al.

Table 1: Comparison of Medical Agent Benchmarks. Benchmark

Long.a Ctx.b EHR.c Dec.d Temp.e Focus

MedAgentBench[6] × × ✓ ✓ × EHR Tool Integration AgentClinic[17] × × × ✓ × Simulated Interaction DiagBench[15] × × ✓ ✓ × Diagnostic Trajectory MedBench v4[3] × ✓ ✓ × × Medical Knowledge QA ReflecTool[10] × × × ✓ × Reflective Tool-Use EHRSQL[9] × × ✓ × × Relational Querying LongMedBench ✓ ✓ ✓ ✓ ✓ Long-horizon Reasoning a Longitudinal: Reasoning across multiple discrete clinical visits. b Context: Whether the benchmark environment contains multi-turn context. c EHR: The dataset built upon realworld EHR. d Decision-making: Evaluating proactive clinical planning beyond static retrieval or SQL-based querying. e Temporal Sensitivity: Evaluating understanding of clinical timeline and urgency.

treatment responses over years [4]. This capacity of long-horizon clinical reasoning is fundamental to high-quality care. As Large Language Models (LLMs) transition to autonomous medical agents, their ability to navigate complex trajectories in real-world Electronic Health Records (EHR) [7] has become the critical benchmark for clinical readiness. Although general-purpose benchmarks have emerged to evaluate long-horizon or multi-turn interactions [19, 18, 8, 20, 5, 3], these benchmarks primarily emphasize explicit reasoning based on information retrieval—the ability to locate specific facts within a long context [11]. They test a model’s capacity to find a “needle in a haystack”, but overlook the temporal dynamics that are essential to clinical reasoning. In real-world medical scenarios, the challenge extends beyond retrieving isolated facts to understanding how a patient’s state evolves over dozens of visits and how this temporal progression informs treatment strategies. In parallel with these general limitations, current medical evaluation paradigms also fall short of this requirement, as shown in Table 1. Despite introducing simulated interactive environments, recent frameworks [15, 6, 17, 3, 10, 9] are constrained by limited context windows and session counts, emphasizing the agent’s tool-call abilities and immediate prediction rather than understanding of complete clinical trajectories. Consequently, they fail to assess how models utilize extensive medical history for future clinical decisions. To address these limitations, we introduce LongMedBench, a MIMIC-IV[7]based benchmark that overcomes constrained context lengths and simple factual recall by curating extensive longitudinal trajectories and designing reasoning tasks that challenge agent’s temporal sensitivity. By converting 335 patient records into longitudinal event streams (averaging 19.72 visits per patient), we construct a temporally dense environment and a framework explicitly targeting long-horizon clinical reasoning. Our contributions are: (1) A tri-level memory dataset architecture for different granularity history evaluation. (2) A progressive evaluation taxonomy that spans three hierarchical tasks: factual QA based on timestamp or relative positioning, targeting the fact retrieval limitation of general benchmarks; temporal reasoning for multi-visit and eventlevel ordering, addressing the lack of time-sensitive evaluation in existing medical

LongMedBench

3

frameworks; and long-horizon decision-making which directly challenges the agent’s ability to navigate extensive histories and autonomously plan next-step clinical actions. (3) Experimental findings: we showed that while state-ofthe-art LLMs can exploit explicit timestamps, they struggle with the implicit temporal reasoning required for visit-level understanding. Although RAG and memory systems improve fact retrieval performance, decision-making accuracy remains highly dependent on immediate context, highlighting a profound limitation in reasoning over long-term clinical trajectories.

2 Methodology 2.1 EHR Data Processing Pipeline LongMedBench is built on MIMIC-IV[7], a public database that contains medical records for > 100, 000 patients. To convert static EHRs into an interactive agent environment, we design a three-stage pipeline (Figure 1). STEP 1 : Event & Patient Filtering

P***.json

admission

MIMIC-IV Database patient

preception

 complete notes

CC, HR...

Visit 001

09/12/07 12:35 CBC, BMP...

labs ...

Y > 15 visits labevent

Visit 002

pharmacy

History visit:

Current visit:

P***-Vi-E1

09/12/07 19:35 Targeted liver ultrasound to target right hepatic lobe hyperechoic...

Timestamp:

09/12/07 11:44

€ abnormal flags

Hospitalization

STEP 3 : Memory Dataset Generation

STEP 2: Event Stream Construction

INPUT: EHR data

Note

IMAGING P***-Vi-E2

P***-Vi-E3

...

P***-Vi+1-E1 ...

P***-Vi+1-E2

...

...

...

P***-Vj-1-E1

P***-Vj-1-E2

...

Event Memory

P***-Vi-Em P***-V2-En ... P***-Vj-1-Ep

Contextual Memory

n > 15 ... ...

... radiology

microbio

procedure

discharge

diagnosis

admission

09/13/07 08:20 discharge

Visit n 355 Patients

6,999 Visits

314,294 Events

notes

Assistant User Assistant

09/12/07 15:27 order_imaging, "modality": "xray", "target": "chest"... 09/12/07 17:45 imaging_results: ... The lungs volumes are low... 09/12/07 15:27 order_imaging, "modality": "ultrasound", "target": ... ...

Note Memory

Note Summary OUTPUT: Memory datasets

Fig. 1: Data processing pipeline for LongMedBench

Event & Patient Filtering To reduce redundancy while preserving clinical signals, we retain only abnormal lab results and patients with complete admission/discharge records. Filtering for patients with ≥ 15 hospitalizations yields 355 patients and 6,999 visits. With an average of 19.72 visits per patient (median=18.00, SD=5.72), this dense multi-session structure rigorously evaluates the agent’s ability in cross-session memorization and temporal reasoning. Event Stream Construction A complete event stream will be constructed for each patient P. The event stream S contains a patient’s visit records S := {V1 , V2 , . . . , Vn }, where Vi is the ith visit of the patient and n ≥ 15. On average, each visit Vi contains 44.91 medical events (median=25.00, SD=70.32). A visit Vi begins with an admission event Eiadm , followed by a series of specific medical events, and ends with a discharge event Eidis . The kth medical event in Vi is denoted as Eik = {aki , tki , pki , oki }, where aki is an action within a typical clinic action space A, which includes imaging, lab tests, medication, etc. tki is the event timestamp, and pki is the corresponding action argument, such as the modality (CT, X-ray, MRI) of an imaging event. oki is the event result or clinical

4

Z. Xu et al.

observation, like the radiology report in an imaging event. Vi can therefore be expressed as a union of sequenced events: Vi = {Eiadm , Ei1 , Ei2 , . . . , Eidis }. In addition to the structured hospitalization data, each visit also includes two types of notes: one is radiology notes, which are parsed as imaging events; the other is discharge notes, which are summaries Ni from visits Vi , and can be parsed into admission info Niadm and discharge info Nidis according to logic. Memory Dataset Generation To evaluate how agents utilize long-context information, we design three memory modules with progressively finer granularity. From the beginning history visit Vi , current visit Vj and reasoning timestamp T , the agent’s memory access is strictly bounded to prevent event leakage. (i,j) 1. Note Memory: MN = {Ni , Ni+1 , . . . , Nj−1 }, which contains visit-level summaries from preceding Vi to Vj−1 . The summary is high-level compressed and emphasizing global trajectory rather than detailed event recall. ∪j−1 (i,j) 2. Event Memory: ME = k=i Vk , which contains all clinical events from Vi to Vj−1 . Events are flattened into a single chronological sequence, discarding visit-level boundaries, and forming a unified patient trajectory. (j,T ) 3. Contextual Memory: MC = {(r0 , m0 ), (r1 , m1 ), . . . , (rP , mP )}, where rk and mk denote the dialog role and message content. It is constructed by rewriting the event stream of current Vj into an LLM-style trajectory. Only events Ejk with tkj < T are retained. Each Ejk is transformed into a action–feedback pair ′ ′ using LLM: {(assistant, {akj , tkj , pkj }), (user, {fjk , tkj })}, where fjk and tkj denote simulated clinic feedback inferred from okj and corresponding timestamp. This memory simulates the instant interaction status for a medical agent.

2.2 Benchmark Question Generation (i,j)

(i,j)

(j,T )

Based on the memory modules {MN , ME , MC }, we construct three progressively challenging task families, from factual QA, temporal reasoning, to longhorizon decision making, reflecting increasingly global memory dependency.

Factual QA This task evaluates precise retrieval and temporal alignment within the agent’s event memory. From visit Vi , a ground truth event Eik ∈ Vi is randomly sampled, and is unified as questions using the following two formats: Explicit tasks include tki in the query text, evaluating direct retrieval capacity; while Relative questions provide only relative temporal relationships without timestamp (e.g. the medicine name in the 2nd medication event of Vi ). The questions require the agent to recover full factual details pki from a redundant (i−m,i+m+1) memory window ME provided, where m ≥ 0 is a hyperparameter to adjust the window size. We place Vi in the middle of the window to evaluate retrieval from non-recency-favored regions of long-context memory [11].

LongMedBench Task 3: Long-Horizon Decision

Task2: Temporal Reasoning

Task1: Factual QA

(1) Event Sorting:

(a) Agent input:

5

(a) Agent Input: Task context

(c) Score: Time-decay mechanism for T3-N/A

(a) Agent Input:

Event memory:

Working context

with explicit timestamp cue

Event stream: Type: Admission, Visit_id: V**, Eventid: P***_V**_adm, Timestamp: ......, Data: .... ... ...

Admission

[?]

[8:00] Admission

[9:00]

missing 1((g1)

Lab: Blood

[?]

[12:30]

missing i((gi)

Imaging

[?]

[20:30]

missing j((gj)

Discharge

Instructional prompt

(1st lab) Typeˆlab, Visit_id: V**, Eventi_id: P***_V**_E1, TimeStamp: ..... ,(Data: {˝INR(PT)˛: {˝value˛: ˝9˛, ˝unit˛:˛s˛}...}...

Visit n-2 Event

Shuffled event options:

Option 1

Option 3

Option 8

[9:15]

[13:00]

[11:50]

Visit n-1 ... ... (7th lab) Typeˆlab, Visit_id: V1, Eventi_id: P00001_V1_E114, TimeStamp: ..... ,(Data: {˝INR(PT)˛: {˝value˛: ˝15˛, ˝unit˛:˛s˛}...}...

Visit n Event

...

Medication:(Antibiotic

Procedure: TACE

...

Lab: biopsy

˛Now we will ... Your task is ...˛

Long-horizon memory ...

(b) Agent Output: a list of option IDs for missing items in event stream G=[g1, ..., gi, ..., gj, ...] Output: [1, ..., 3, ..., 8, ...]

(c) Score: Kendallˇs Ä

Multiple-choice question

Visit n+1 GT: [1, ... , 8, ..., 3, ...] Type: Discharge, Visit_id: V**, Eventid: P***_V**_dis, Timestamp: ......, Data: ˝Cheif Complaint: hypotension ...˛

GT

Discharge Visit n+2

(b) Agent Output

(2) Visit Sorting:

ti - t = 0 medication

order_labs

Score = 0

S = 0.0

... ... (a) Agent Input: 5 Shuffled visit summary of 5 contiguous visits

(b) Agent Output: sequence reconstruction

Questions: Explicit

What is the Chief Complaint and diagnosis in Visit n?

GT

hypotension and lethargy s/p RFA, Hepatitis C cirrhosis

with Implicit logic cue

T3-N: Next Action Prediction Output: [5, 4, 2, 3, 1]

GT

Yes, 15s.

T3-D: Discharge Decision

Summary 2

Summary 4

Summary 3 Admission: RUQ pain Discharge: procedure: Larparoscopic Cholecystectomy Summary 2 Admission: worsening symptoms Discharge: TACE procedure Summary 1

(c) Score: Kendallˇs Ä

Options:

Summary 1

Summary 4 Admission: nausea/vomiting, diarrhea Discharge: diarrhea improved

Which of the next step is most appropriate?

Output: [6, 2, ..., 9, 1]

Which 3 lab panels are most appropriate next?

Should the patient be discharged within 6h?

Summary 3

(b) Agent Output: Cheif Complaint: hypotension and lethargy s/p RFA(Discharge Diagnosis:Hepatitis C cirrhosis

T3-A: Argument Prediction

GT: [5, 2, 4, 3, 1]

Summary 5 Admission: lethargy, hypotension Discharge: lethargy improved In the 7th lab event of Visit n, was INR(PT) measured? If yes, what was the value? Relative

Answer

Question & GT Generation

Summary 5

ask_question, discharge, order_imaging, ...

Options:

Options:

blood cultures, LFT, urinalysis, stool culture ... Yes/No

GT: original temporal order

GT: [6, 2, ..., 9, 3] (c) Score: LLM-based matching score

GT

Keywords

GT:

(3) Joint Sorting:

[˝hypotension and lethargy s/p RFA˛, ˛hepatitis C cirrhosis˛]

GT:

GT: order_labs

Does the answer cover all keywords?

CBC, ABG, Coagulation Panel

No

Agent Input: 10 Shuffled visit admission/discharge clips split from note summary Split

Yes = correct

Summary 5

Clip 9

Clip 1

Score: Accuracy

Score: top-k Acc

Score: Accuracy

Fig. 2: Example of the three evaluation tasks in LongMedBench.

Temporal Reasoning This task evaluates the agent’s ability to reconstruct chronological order at different memory granularities. The evaluation score is measured by Kendall’s τ . Given a target event stream S or its visit Vi , we design three sub-tasks using the original event streams as ground truth, each targeting at a distinct level of temporal reasoning. Visit Cloze tests event-level reasoning: several intervention events Eik ∈ Vi with timestamps preserved are masked and provided to the agent as a shuffled list. The agent must insert each event back to its original position. Visit Sorting evaluates visit-level ordering: five shuffled (i,i+5) visit summaries N ∈ MN must be ordered solely by clinical progression cues with their timestamps removed. Joint Sorting assesses integration: five visits are split into ten admission/discharge clips. The agent must correctly pair the clips from one visit and order them in time sequence.

Long-Horizon Decision Making This task evaluates agent’s next-step planning under long-term historical context, a more comprehensive evaluation. From Vn , we sample a ground-truth event EiGT , and define the agent’s contextual memory as its working context S. The long-horizon note or event memory is also provided as a reference M . When generating question Q, we randomly select one of the following three sub-task formats:Next Action Prediction (T3-N) requires the agent to predict aGT i , Argument Prediction (T3-A) evaluates , and Discharge Decision (T3-D) queries if the patient can inference for pGT i be discharged within 6h. For two prediction tasks, as shown in Figure 2, a timedecay scoring mechanism is applied to recognize actions E24h , events that occur within 24h after tGT i . Unlike prior tasks, this task requires integrating immediate state S with long-term memory M to produce clinically coherent decisions. It evaluates true long-horizon reasoning rather than isolated memory access.

6

Z. Xu et al.

3 Experiments and Results 3.1 Experiment Environments To verify the robustness of medical agents in long-term clinical decision, we selected representative long-context LLMs as the agent’s backbone, including the closed-source frontier model gpt-5-mini[13], the lightweight one qwen-turbo[16], and the open-source high-performance one deepseek-v3.2[2]. All LLMs are deployed with default configurations. Referring to common retrieval methods in long-horizon tasks[19, 12], we set up three memory architectures: (1) Naive Long-Context. No external tools is used; memories form preceding visits are greedily injected into the LLM’s context window. (2) RAG. text-embedding-3-small[14] is used to vectorize memory entries; the agent can retrieve Top-K similar memory fragments when answering the question. (3) Agent Memory System. A dynamically updated external storage layer is included. When answering questions, the agent can spontaneously search based on the request and recall appropriate memories. We introduce Mem0[1], a product-ready agent memory architecture, in the following experiments. 3.2 Results Factual QA We evaluate Naive Long-Context LLMs using qwen-turbo, comparing window sizes {3, 5, 7} (i.e., m ∈ {1, 2, 3} in ME ) and full history (m = ∞) against the Agent Memory system (Mem0). Results in Table 2 report Lab-T (recall) and Lab-F (rejection) to assess hallucination impact. Table 2 reveals that naive-LLM’s performance decays severely as history grows, dropping from 0.882 to 0.423 in explicit retrieval, with low Lab-F scores (≤0.570) highlighting a pervasive hallucination bias. The performance of agent memory system is strongly correlated with the specific task type; while Mem0 achieves near-optimal results in explicit Lab-T (0.993) and Medication (0.983), its relative performance (Overall 0.331) remains inferior. This gap reveals that: first, agents struggle to generate precise queries for imaging reports even with few-shot prompting; interestingly, when falling back to vector-based search, relative semantic queries outperform explicit ID/timestamp matching, as the latter lacks sufficient embedding density in vector space; second, the agent’s memory architecture lacks a systematic indexing of relationships between events. Table 2: Performance on Factual Retrieval under Explicit and Relative Settings. Explicit

Relative

Model L-Tf L-Fg L-Oh Med. Img. Adm/Dis Avg. L-Tf L-Fg L-Oh Med. Img.

Avg.

m = 1 0.96 0.57 0.83 0.84 0.89 m = 2 0.92 0.51 0.78 0.71 0.81 m = 3 0.86 0.50 0.74 0.61 0.75 m = ∞ 0.63 0.29 0.51 0.28 0.41 Mem0 0.99 0.43 0.80 0.98 0.07

0.80 0.71 0.64 0.34 0.33

0.97 0.92 0.87 0.46 0.98

0.88 0.91 0.58 0.80 0.78 0.84 0.81 0.86 0.52 0.74 0.64 0.74 0.74 0.80 0.50 0.70 0.55 0.68 0.42 0.52 0.28 0.44 0.23 0.36 0.74 0.71 0.57 0.66 0.17 0.13

f L-T: Lab-T, recall of existing lab indicators. g L-F: Lab-F, rejection of not found lab indicators. h L-O: Overall accuracy of lab events.

LongMedBench

7

Table 3: Temporal reasoning performance (Kendall’s τ ) on LongMedBench. visit_cloze

visit_sorting

joint_sorting

Model

N

Mean

Std

N

Mean

Std

N

Mean

Std

deepseek-v3.2 gpt-5-mini deepseek-v3.2-thinking qwen-turbo

6776 6776 6776 6776

0.316 0.925 0.969 0.047

0.352 0.220 0.124 0.257

765 765 765 765

0.258 0.376 0.424 0.114

0.523 0.552 0.549 0.456

2878 2878 2878 2878

0.132 0.295 0.330 0.033

0.346 0.439 0.425 0.256

Temporal Reasoning Table 3 presents Kendall’s τ for temporal reasoning across three progressive tasks. In event-level visit cloze, top models achieve near-perfect performance (gpt-5-mini: 0.925), while weaker models struggle (qwenturbo: 0.046). Enabling thinking mode substantially boosts reasoning (deepseekv3.2-thinking: 0.969 vs. deepseek-v3.2: 0.316). This confirms that strong LLMs can effectively utilize explicit timestamp cues for event-level ordering, regradless of the option numbers (the length of the event stream), as illustrated in Figure 3a. Moving to Visit Sorting, where timestamps are absent so that models must sort five complete visit summaries based solely on clinical progression, performance drops sharply (best: deepseek-v3.2-thinking 0.423). This reveals that implicit temporal reasoning in visit-level remains challenging.

(a) Mean Kendall’s τ vs. number of options in visit_cloze.

(b) Comparison between joint_sorting and visit_sorting.

Fig. 3: Temporal reasoning analysis. The difficulty further escalates in Joint Sorting, where each visit summary is split into admission and discharge fragments, requiring simultaneous eventlevel pairing and visit-level sorting. Performance declines to 0.330, demonstrating the compounding complexity. Figure 3b shows this clear descending trend, confirming that when crucial information is fragmented at both event and visit level, implicit temporal reasoning in realistic multi-visit scenarios remains a major challenge.

8

Z. Xu et al.

Table 4: Long-context LLM performance on Long-Horizon Decision Making. Note Memory

Event Memory

Model

Ctxi T3-A T3-D T3-N

Avg

T3-A T3-D T3-N

Avg

qwen-turbo deepseek-v3.2 gpt-5-mini deepseek-v3.2-thinking

128K 128K 400K 128K

0.434 0.422 0.445 0.420

0.494 0.493 0.494 0.466

0.420 0.439 0.452 0.419

0.475 0.478 0.478 0.471

0.604 0.607 0.625 0.578

0.309 0.278 0.324 0.294

0.604 0.607 0.607 0.593

0.309 0.304 0.335 0.290

i Ctx: the maximum context length of LLM.

Long-Horizon Decision Making We first evaluate the performance of Naive Long-Context. All models are limited to a maximum context of 128K tokens. As shown in Table 4, structured event memory outperforms note memory in most models in decision making. gpt-5-mini performs best. Notably, for deepseek-v3.2, thinking patterns did not significantly improve decision performance. This is due to thinking mode makes the agent more conservative, attempting to ask question to get instant information rather than referring to past memories. To analyze the Lost-in-Middle[11] effect, using qwen-turbo as the baseline model, we (1) adjusted the number of injected historical visits and (2) implement RAG and Mem0 architecture as retrial source in separate experiments. As shown in Table 5, although the difference is subtle, close to that under full memory conditions, increasing the context length makes the agent perform worse. The performance differences between Mem0 and RAG are minimal, both close to the baseline performance, demonstrating that existing memory augmentation architectures offer limited gains for long-term decision tasks. Table 5: Ablation study. Left: visit injection (n). Right: memory architectures. Visit Injection n

Memory Architecture

Memory T3-A T3-D T3-N Avg Arch Memory T3-A T3-D T3-N

0 (Base)j –

0.49 0.61 0.32 0.45 RAG event 2 event 0.49 0.60 0.32 0.44 Mem0 event 5 event 0.48 0.60 0.31 0.43 RAG note 2 note 0.48 0.60 0.32 0.44 Mem0 note 5 note 0.48 0.60 0.32 0.44 j Baseline with no memory content injected.

0.45 0.51 0.29 0.47 0.62 0.31 0.48 0.60 0.31 0.46 0.56 0.29

Avg 0.40 0.44 0.44 0.41

Therefore, under the current LongMedBench setting, decision quality is not strongly correlated with the amount of retrieved historical information, while the agent performance depends more on the model’s clinical reasoning capacity within the immediate context. Accordingly, we need to design tasks that amplify deeper reasoning requirements and higher decision interdependencies in an attempt to better reveal the limitations of current models and more effectively assess their long-term clinical reasoning capabilities.

LongMedBench

9

4 Conclusion We introduce LongMedBench, a benchmark using real-world EHR data to evaluate medical agents in long-horizon clinical reasoning. By converting MIMIC-IV records into multi-session event streams, we assess agents across three memory types with varying tasks. Experiments reveal that while state-of-the-art models handle explicit timestamps well, implicit temporal reasoning remains a significant bottleneck. As history scales, models increasingly rely on immediate context rather than cross-session integration. Furthermore, while memory augmentation improves retrieval, decision-making performance is highly task-sensitive and remains limited by the model’s instant reasoning capacity. Our results highlight a critical gap in handling time-dependent clinical tasks for medical agents, necessitating future optimization of cross-session memory augmentation architectures.

References 1. Chhikara, P., Khant, D., Aryan, S., Singh, T., Yadav, D.: Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413 (2025) 2. DeepSeek-AI: Deepseek-v3.2: Pushing the frontier of open large language models (2025) 3. Ding, J., Lu, L., Ding, C., Bian, M., Chen, J., Pang, W., Chen, R., Peng, X., Lu, R., Ren, S., Zhu, G., Wu, X., Liu, Z., Zhang, R., Jiang, L., Han, B., Wang, Y., Xu, J.: Medbench v4: A robust and scalable benchmark for evaluating chinese medical language models, multimodal models, and intelligent agents (2025), https://arxiv.org/abs/2511.14439 4. Gong, E.J., Bang, C.S., Lee, J.J., Baik, G.H.: Knowledge-practice performance gap in clinical large language models: Systematic review of 39 benchmarks. Journal of medical Internet research 27, e84120 (2025). https://doi.org/10.2196/84120 5. Hsieh, C.P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., Ginsburg, B.: Ruler: What’s the real context size of your long-context language models? (2024), https://arxiv.org/abs/2404.06654 6. Jiang, Y., Black, K.C., Geng, G., Park, D., Zou, J., Ng, A.Y., Chen, J.H.: Medagentbench: A virtual ehr environment to benchmark medical llm agents. NEJM AI p. AIdbp2500144 (2025) 7. Johnson, A., Bulgarelli, L., Pollard, T., Gow, B., Moody, B., Horng, S., Celi, L.A., Mark, R.: MIMIC-IV. PhysioNet (Oct 2024). https://doi.org/10.13026/kpb9-mt58, https://doi.org/10.13026/kpb9-mt58, version 3.1 8. Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K.M., Melis, G., Grefenstette, E.: The narrativeqa reading comprehension challenge (2017), https://arxiv.org/abs/1712.07040 9. Lee, G., Hwang, H., Bae, S., Kwon, Y., Shin, W., Yang, S., Seo, M., Kim, J.Y., Choi, E.: Ehrsql: A practical text-to-sql benchmark for electronic health records. Advances in Neural Information Processing Systems 35, 15589–15601 (2022) 10. Liao, Y., Jiang, S., Wang, Y., Wang, Y.: Reflectool: Towards reflection-aware toolaugmented clinical agents (2024), https://arxiv.org/abs/2410.17657 11. Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P.: Lost in the middle: How language models use long contexts (2023), https://arxiv.org/abs/2307.03172

10

Z. Xu et al.

12. Maharana, A., Lee, D.H., Tulyakov, S., Bansal, M., Barbieri, F., Fang, Y.: Evaluating very long-term conversational memory of llm agents. arXiv preprint arXiv:2402.17753 (2024) 13. OpenAI: Gpt-5 mini model. https://developers.openai.com/api/docs/models/gpt5-mini (2026), accessed 22 Feb 2026 14. OpenAI: text-embedding-3-small model. https://developers.openai.com/api/docs/models/textembedding-3-small (2026), accessed 22 Feb 2026 15. Qiu, P., Wu, C., Liu, J., Zheng, Q., Liao, Y., Wang, H., Yue, Y., Fan, Q., Zhen, S., Wang, J., Gu, J., Wang, Y., Zhang, Y., Xie, W.: Evolving interactive diagnostic agents in a virtual clinical environment (2026), https://arxiv.org/abs/2510.24654 16. QwenTeam: Qwen3: Think deeper, act faster. https://qwen.ai/blog?id=qwen3 (2024), accessed: 22 Feb 2026 17. Schmidgall, S., Ziaei, R., Harris, C., Reis, E., Jopling, J., Moor, M.: Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments (2025), https://arxiv.org/abs/2405.07960 18. Shen, Y., Huang, Z., Wang, Z., Tian, M., Guo, Z., Zhang, C., Zhou, S., Hu, Z., Li, D., Xu, J., Wang, K., Liu, W., Li, T., Yue, F., Hong, F., Liu, C., Zeng, K.: Tripbench: A benchmark for long-horizon interactive agents in real-world scenarios (2026), https://arxiv.org/abs/2602.01675 19. Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.W., Yu, D.: Longmemeval: Benchmarking chat assistants on long-term interactive memory. https://arxiv.org/abs/2410.10813 (2024) 20. Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., Manning, C.D.: HotpotQA: A dataset for diverse, explainable multi-hop question answering. In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018)

Record · ID 361513 · SHA-256 2fd5ebeb574607c4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.