ConceptioArchivearXiv CS
arXiv CSopen access

ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection He Geng* , Yangmin Huang* † , Lixian Lai, Qianyun Du† , Hui Chu, Zhiyang He, Jiaxue Hu, Xiaodong Tao Xunfei Healthcare Technology Co., Ltd. {hegeng2, ymhuang9, lxlai2, qydu, huichu2, zyhe, jxhu2, xdtao}@iflytek.com

Accuracy

ProMedical-Bench

Latent Reward Landscape

arXiv:2604.08326v1 [cs.AI] 9 Apr 2026

Communication Quality

Expert Adjudication

Instruction Following Medical Experts

Rubric-by-Rubric Review Feedback Document

The Alignment Gap

The main criteria and score item 6 are unreasonable…

Training Signal

Binary Feedback

Only tells which better

Clinical Standards

Completeness

Fine-grained Rubrics

The first bonus item was incorrectly judged…

Clinical Instruction Response A

Response B

Rubrics-level Evaluation Iterative process Gold-Standard Labels

Safety, Proficiency and Excellence

Motivation: Alignment Gap

Contextual Awareness

Pairwise Data

ProMedical (Ours)

Learning from Binary Feedback

(b) Algorithm

(a) Data

Revision Proposals

100% double-blinded verified

(c) Evaluation

Figure 1: Motivated by the alignment gap between coarse binary signals and the high-dimensional latent reward landscape of clinical standards, we introduce the ProMedical suite: Data: ProMedical-Preference-50k, incorporating fine-grained clinical rubrics, hierarchical score vectors, and language feedback; Algorithm: a rubricdriven alignment paradigm that strictly enforces safety compliance and enhances reasoning depth; Evaluation: ProMedical-Bench, establishing a rigorous benchmark via double-blinded expert adjudication.

Abstract

proving overall accuracy by 22.3% and safety compliance by 21.7%, effectively rivaling proprietary frontier models. Furthermore, the aligned policy generalizes robustly to external benchmarks, demonstrating performance comparable to state-of-the-art models on UltraMedical. We publicly release our datasets, reward models, and benchmarks to facilitate reproducible research in safety-aware medical alignment.

Aligning Large Language Models (LLMs) with high-stakes medical standards remains a significant challenge, primarily due to the dissonance between coarse-grained preference signals and the complex, multi-dimensional nature of clinical protocols. To bridge this gap, we introduce ProMedical, a unified alignment framework grounded in finegrained clinical criteria. We first construct ProMedical-Preference-50k, a dataset generated via a human-in-the-loop pipeline that augments medical instructions with rigorous, physician-derived rubrics. Leveraging this corpus, we propose the Explicit Criteria Injection paradigm to train a multi-dimensional reward model. Unlike traditional scalar reward models, our approach explicitly disentangles safety constraints from general proficiency, enabling precise guidance during reinforcement learning. To rigorously validate this framework, we establish ProMedicalBench, a held-out evaluation suite anchored by double-blind expert adjudication. Empirical evaluations demonstrate that optimizing the Qwen3-8B base model via ProMedical-RMguided GRPO yields substantial gains, im* †

1

Introduction

Large Language Models (LLMs) have demonstrated unprecedented potential in transforming healthcare. Recent studies indicate that proprietary models, such as Med-PaLM 2, MedFound and Lingshu, have achieved proficiency approaching that of clinicians (Singhal et al., 2025; Liu et al., 2025; Xu et al., 2025). These models are capable of assisting physicians in case analysis and clinical diagnosis while providing second opinions for decision-making(Mehandru et al., 2025; O’Sullivan et al., 2024). On the patient side, they facilitate tasks such as drafting preliminary treatment plans and performing medical triage(Hsu et al., 2025; Health, 2024). However, a critical misalignment persists. Although contemporary

Equal contribution Corresponding author

1

specific rubrics, while the latter provides a heldout evaluation protocol anchored by doubleblind expert adjudication, ensuring strict alignment with professional clinical criteria. • We propose the explicit criteria injection paradigm, which trains a multi-dimensional reward model to steer GRPO. By internalizing complex medical protocols as dense, hierarchical reward signals, this method effectively disentangles safety constraints from general helpfulness, ensuring robust compliance in high-stakes scenarios. • We develop and release ProMedical-RM, a rubric-aware reward model employed to steer policy optimization via GRPO. Empirical evaluations demonstrate that this paradigm secures a 22.3% gain in overall accuracy and a 21.7% enhancement in safety compliance on our expertadjudicated benchmark, while maintaining robust generalization on public datasets. We opensource our code and datasets to facilitate reproducible research in safety-aware medical alignment.

evaluation benchmarks increasingly emphasize fine-grained reasoning grounded in clinical facts, which necessitates expert-level analytical capabilities and logical deduction processes(Arora et al., 2025; Manes et al., 2024), the underlying training paradigms predominantly rely on coarse-grained, binary supervisory signals(Rafailov et al., 2023; Shao et al., 2024). This discrepancy between training objectives and evaluation paradigms constitutes a significant barrier to the widespread deployment of artificial intelligence in the medical domain(Kim et al., 2025). Despite significant strides in biomedical domain adaptation and clinician-informed alignment (Luo et al., 2022; Zhang et al., 2023a; Ouyang et al., 2022; Rafailov et al., 2023), current pipelines face intrinsic limitations when addressing high-stakes medical errors. The prevailing reliance on holistic preference pairs is fundamentally inefficient for capturing the long-tail distribution of clinical pitfalls, as it forces models to implicitly infer complex rationales from binary signals(Qiu et al., 2025; Tien et al., 2022). This creates spurious correlations where models conflate safety with surface-level fluency rather than internalizing precise medical logic(Pahde et al., 2025; Liao et al., 2023). Such coarse supervision stands in stark contrast to evolving evaluation standards that prioritize clinically grounded assessments of reasoning and hallucination control (Arora et al., 2025; Hosseini et al., 2024; Seo et al., 2024a; Manes et al., 2024). Consequently, rigorous rubric-based assessments are largely relegated to post hoc validation (Arora et al., 2025; Kim et al., 2024; Liu et al., 2023), a disconnect further corroborated by reward-model benchmarks that reveal limited generalization under structured constraints (Lambert et al., 2025; Gunjal et al., 2025; Wang et al., 2025). To bridge this gap, we propose ProMedical, a unified framework that incorporates instructionlevel, clinician-defined rubrics into preference construction, reward modeling, and evaluation. Rather than treating rubrics as an external diagnostic tool, ProMedical embeds rubric-based criteria directly into the alignment process, explicitly aligning training objectives with clinically grounded evaluation standards. Our contributions are three-fold: • We construct ProMedical-Preference-50k and ProMedical-Bench, establishing a rigorous data foundation for medical alignment. The former enriches training samples with instruction-

2

Rubrics

In this section, we introduce a unified automated clinical metric construction algorithm, upon which we build ProMedical-Rubrics. Representing a high-dimensional, multi-faceted preference evaluation strategy, this framework is designed to provide Reinforcement Learning with more finegrained reward representations, capturing subtle clinical nuances that coarse scalar metrics often overlook. We start by briefly outlining the preliminaries of preference construction, focusing on how current approaches determine the ordinal ranking of response pairs. 2.1

Background and Preliminary

In the context of aligning medical language models, preference modeling serves as the cornerstone for distinguishing high-quality clinical responses. Formally, for an instruction q sampled from the dataset D, we derive a set of K candidate responses Rq = {r1 , . . . , rK }. The underlying mechanism for learning from these responses typically relies on the BradleyTerry model(Sun et al., 2025), which posits that the probability of a preferred response yw prevailing over a dispreferred one yl is determined by the 2

This dimension incentivizes models to exceed standard clinical expectations. • Safety Veto (S3 ): Detects critical infractions like severe hallucinations or toxic advice. Unlike soft penalties, it imposes a hard constraint to enforce a strict safety lower bound.

difference in their latent reward scores: P (yw ≻ yl |q) = σ(rϕ (q, yw ) − rϕ (q, yl )), (1) where σ(·) is the sigmoid function and rϕ represents the reward model parameterized by ϕ. Based on this formulation, existing annotation paradigms predominantly categorize into Pointwise Scoring, Pairwise Comparison, and Generative Feedback. While these methods have established foundations for general alignment, they exhibit distinct limitations when applied to the high-stakes clinical domain, particularly regarding inter-annotator reliability and the granularity of feedback. We provide a comprehensive analysis of these paradigms in Appendix F. 2.2

Hierarchical Preference Ranking. A key innovation of our framework is that these three components do not simply sum up. Instead, we adopt a Lexicographical Comparison Protocol to strictly enforce safety constraints before evaluating proficiency or style. For two responses rA and rB , the preference relation is determined hierarchically:  A B  S3 < S3 , rA ≻ rB ⇐⇒ S1A > S1B , if S3A = S3B (5)   A S2 > S2B , otherwise

Tripartite Evaluation Schema and Hierarchical Scoring

Mechanistically, this formulation imposes a hard constraint on the optimization landscape, effectively severing the gradient trajectory towards unsafe regions. By establishing a rigid decision boundary, it ensures that proficiency gains (S1 ) cannot incentivize the model to traverse beyond ethical limits, thereby rigorously enforcing the Do No Harm imperative.

As illustrated in Figure 3, to emulate the sophisticated decision-making processes of clinical practitioners, we project the alignment objective from low-dimensional binary classification onto a high-dimensional clinical manifold via a Tripartite Evaluation Schema. Specifically, we decompose the clinical utility of a response r into three orthogonal dimensions: Proficiency, which serves as the primary evaluation metric; Excellence, acting as a bonus reward mechanism; and Safety. Diverging from the scalar deduction paradigms in HealthBench and K-QA, which risk permitting optimization algorithms to trade safety for utility, we operationalize Safety as a strict veto constraint to enforce non-negotiable clinical boundaries.

3

Figure 2 illustrates the schematic overview of the proposed framework. The ProMedical-Rubrics framework not only constitutes a robust evaluation metric but also facilitates versatile training paradigms for aligning LLMs with clinical standards. Leveraging GRPO as the underlying optimization backbone, we formalize two distinct alignment strategies: Implicit Outcome Alignment and Explicit Criteria Injection.

Tripartite Components Definition. Formally, the rubric Rq induces a quantitative triplet S = (S1 , S2 , S3 ), quantified via the indicator function I(·): X ωi · vi , (2) S1 =

3.1

S2 =

I(r |= c),

(3)

I(r ̸|= c),

(4)

c∈Cbonus

S3 =

X

Paradigm I: Implicit Outcome Alignment

The first paradigm adheres to the groupwise preference learning formulation. Here, the generated rubrics function as a hierarchical oracle to assign scalar rewards to a group of sampled responses. In this setting, the model is optimized to maximize the likelihood of high-reward outputs relative to the group baseline, enabling it to internalize the latent reward landscape without explicit rubric supervision.

ci ∈Cmain

X

Rubric-Enabled Alignment Paradigms

c∈Cveto

• Main Proficiency (S1 ): Quantifies fundamental clinical accuracy and completeness. It functions as the weighted baseline metric derived from point-specific importance ωi . • Excellence Bonus (S2 ): Rewards superior attributes such as empathy and logical coherence.

Formulation. Formally, let D = {(x, Rx )} denote the augmented dataset, where each instruction x is paired with an instruction-specific clinical 3

rubric Rx . During training, we sample a group of G outputs {y1 , . . . , yG } from the reference policy πref for each input x. Evaluation against Rx yields a triplet S(i) = (S1 , S2 , S3 ). To synthesize these dimensions into a scalar optimization signal, we propose a cumulative penalty mechanism. We define the proficiency score S1 as the weighted sum of essential criteria, strictly normalized such that the total weight sums to 1 (i.e., P wprof = 1). To incentivize the model to pursue excellence features (S2 ) beyond mere correctness, we formulate the reward ri with an extended upper bound: (i)

(i)

Formulation. Formally, we redefine the reward modeling task as estimating the conditional preference P (yw ≻ yl | x, c), where c represents a specific rubric dimension. To train this evaluator, we implement dimensional data expansion. For an instruction x with K applicable rubrics, we decompose a single response pair into K distinct instances, assigning preference labels independently for each criterion. The optimization objective minimizes the negative log-likelihood: LRM (ϕ) = −EDexp [log σ (∆rϕ (yw , yl | x, c))] , (8) where ∆rϕ (·) = rϕ (yw |x, c) − rϕ (yl |x, c) denotes the conditional reward margin. Upon convergence, this RA-RM serves as the precision oracle for Paradigm I, computing the granular dimension-wise scores that are hierarchically aggregated—strictly enforcing safety vetoes prior to summing weighted proficiency scores and excellence bonuses—to determine the final preference ranking.

(i)

ri = Clip(S1 + αS2 , 0, 1 + β) − λ · S3 , | {z } | {z } Extended Utility

Safety Penalty

(6) where α < 1, Clip(·, 0, 1 + β) normalizes the (i) positive utility, and S3 represents the count of safety violations. Crucially, we introduce a margin parameter β > 0 to prevent reward saturation: this ensures that excellence bonuses are not truncated even when proficiency is perfect (S1 = 1), thereby maintaining valid gradient signals for superior clinical reasoning. Conversely, the penalty coefficient λ ≥ 1 + β is set to ensure that a single safety infraction strictly dominates any potential utility gain, enforcing a hard constraint on harm. We employ GRPO to maximize the expected reward. The objective minimizes the following loss: G i 1 Xh LGRPO = − ρi Âi − βKL DKL , G

4

A primary impediment to current research lies in the structural limitations of existing preference datasets. Predominant approaches rely heavily on coarse-grained pairwise comparisons or simplistic LLM-based adjudication, which lack rulelevel granularity. Conversely, fully manual expert rubrics remain scarce due to scalability bottlenecks and are often prone to inherent subjectivity. This dichotomy creates a significant dissonance between training signals and the standards of meticulously constructed evaluation benchmarks. To bridge this gap, we open-source ProMedicalPreference-50k, the first large-scale medical preference dataset aligned with fine-grained evaluation benchmarks, designed to reconcile model training paradigms with rigorous clinical standards. In this section, we detail the synthesis of instructions and responses. The formulation of the corresponding fine-grained rubrics, which serve as the alignment anchor, is discussed separately in Section 2.

(7)

i=1

i |x) where ρi = ππrefθ (y (yi |x) denotes the importance sam-

pling ratio, Âi represents the advantage computed from the rewards, and DKL = DKL (πθ ||πref ) serves as the trust region constraint. 3.2

Dataset

Paradigm II: Explicit Criteria Injection

While implicit alignment optimizes outcomes, reliance on scalar rewards often obscures the specific rationale behind preference labels, a phenomenon known as scalar conflation. To resolve this opacity, we introduce Explicit Criteria Injection via a Rubric-Aware Reward Model (RARM). This paradigm shifts from holistic scoring to criteria-conditioned evaluation, explicitly disentangling supervision signals to capture finegrained clinical nuances such as safety and empathy independently.

4.1

Instruction Curation Pipeline

The ProMedical-Preference-50k instruction corpus is constructed via a rigorous four-stage curation pipeline—encompassing data sourcing, semantic deduplication, difficulty curation, and expert-guided hierarchical classification—to ensure high quality and diversity, with detailed pro4

ProMedical-Train-Construction Public Medical Datasets Instruction (~ 823 k)

… Instruction Curation Pipeline (~ 50 k)

ProMedical-Rubrics framework Original GRPO Policy Model

q

Medical-Rubric GRPO (ours) Policy Model

q

chosen

5~9

Human-in-the-Loop Rubric Construction

r!

O"

Reward Model

r"

A!

Safety-aware

A"

Group Computation

A! Group Computation

A" A$

r$

O! O"

Expert Template

O$

Reference Model Category Classification

Reference Model

O$

scoring 0~10

Semantic Deduplication Difficulty Curation

O!

Specific Rubrics

Rubric Construction Priority: S% > S& > S'

r!

𝒊

𝒊

𝒊

𝒓𝒊 = Clip 𝑺𝟏 + 𝜶𝑺𝟐 , 𝟎, 1 + 𝜷 − 𝝀 ⋅ 𝑺𝟑

r"

Extended Utility

Safety Penalty

🔥Rubric-Aware Reward Model

𝐌𝐚𝐢𝐧 𝐏𝐫𝐨𝐟𝐢𝐜𝐢𝐞𝐧𝐜𝐲 (𝐒! ) 𝐄𝐱𝐜𝐞𝐥𝐥𝐞𝐧𝐜𝐞 𝐁𝐨𝐧𝐮𝐬 (𝐒" ) 𝐒𝐚𝐟𝐞𝐭𝐲 𝐕𝐞𝐭𝐨 (𝐒𝟑 )

r$

A$ Filtered Instructions Rubric Construction Expert-Anchored Template Injection Hierarchical Scoring

ProMedical Preference Datasets (~50 k) Datasets

ProMedical-Bench

Response Generation

Instruction

Rubric

Preference

Feedback

Medical experts Adjudication

Implicit Data(795)

Explicit Data(5505)

Instruction Response A

Instruction Response B

Response A

Response B

Score A

Score B

Score A

Score B

S" +Rationale

Rubric-wise Rationale

Rubric-wise Rationale

S$ +Rationale

S! + Rationale

Figure 2: Overview of the ProMedical framework. (Left) Construction of the ProMedical-Preference-50k dataset via a human-in-the-loop pipeline that transforms coarse medical instructions into fine-grained, rubric-enriched training samples. (Top Right) The proposed Medical-Rubric GRPO paradigm, which leverages a Rubric-Aware Reward Model to calculate hierarchical reward scalars based on Main Proficiency (S1 ), Excellence Bonus (S2 ), and Safety Veto (S3 ) to steer policy alignment. (Bottom) The ProMedical-Bench evaluation suite, establishing a robust clinical gold standard through double-blind expert adjudication with rubric-wise rationales.

4.3

tocols provided in Appendix A. The resulting taxonomy distribution is visualized in Figure 6. Furthermore, to facilitate the online generation phase of GRPO, we curated a distinct subset of 10k instructions from the source corpus. This subset adheres to the same quality control protocols while ensuring strict decontamination from both the preference training set and the evaluation benchmarks (details in Appendix A.6).

4.2

Human-in-the-Loop Rubric Construction Protocol

Guided by the protocols defined in Section 2, we construct the rubrics for ProMedicalPreference-50k using an iterative Humanin-the-Loop (HITL) framework designed to ensure clinical rigor at scale. We employ Gemini-3-Pro-thinking (DeepMind, 2025) to instantiate rubrics, conditioning the model on a dual-component prompt: a static expert-defined system instruction and a dynamic pool of few-shot demonstrations. In each alignment cycle, medical professionals adjudicate a stratified batch of 500 generated instances to rectify factual hallucinations and logical omissions. Crucially, these expert-refined gold standards are recursively injected back into the demonstration pool, dynamically updating the few-shot context for subsequent generation cycles. This continuous feedback mechanism ensures the generation quality rapidly converges to professional proficiency, evidenced by a 96.40% pass rate under strict expert evaluation. Following the same process, we employ GPT-4.1 (OpenAI, 2025) as the authoritative judge to annotate the labels of each criteria based on the instantiated rubrics

Response Generation

Drawing inspiration from UltraMedical(Zhang et al., 2024), we establish a diverse candidate pool by leveraging three distinct models spanning both proprietary and open-source landscapes to generate responses. Specifically, our model pool comprises Qwen3-235B-Thinking, Claude-Sonnet-4.5-Thinking, and Deepseek-R1(Yang et al., 2025; Anthropic, 2025; DeepSeek-AI, 2025). This heterogeneous selection strategy allows us to capture a wide spectrum of reasoning patterns and linguistic styles, effectively mitigating the self-reinforcement bias often observed in single-model generated datasets. 5

across K distinct criteria. By filtering out nondiscriminative instruction-rubric pairs, we curated a refined set of 5,505 expanded instances, including 3,625 Proficiency, 1,650 Excellence and 230 Safety pairs dedicated to fine-grained pairwise adjudication. Drawn from the curated corpus described in Section 4, this benchmark maintains a stratified distribution across the five major medical categories, ensuring a balanced representation of diverse clinical scenarios while strictly excluding any instances used during training.

Clinical Instruction

Is endurance sports a risk factor for atrial fibrillation after ablation for atrial flutter? Instruction-Specific Clinical Rubrics Dimension:Accuracy, Rubric: Correctly states that high-level endurance sport is an established risk factor for the de novo development of atrial fibrillation and flutter. weight:0.08… Proficiency The response balances the discussion of risk by also noting that ablation for flutter is highly successful and that a safe return to exercise is often possible with medical guidance, thus alleviating potential anxiety …

Excellence

The response contains discriminatory language against athletes or any group, or discloses non-public patient data. …

Rubric-Guided Expert Adjudication. Distinct from benchmarks dependent on automated metrics or crowd-sourced workers, ProMedical-Bench adopts a rigorous Double-Blind Expert Adjudication Protocol. We engaged a cohort of licensed physicians to conduct an exhaustive, instancelevel annotation of the entire 795-sample corpus. This labor-intensive undertaking necessitated the meticulous verification of every single response against its specific rubric Rx , explicitly scrutinizing adherence to granular checkpoints spanning the tripartite evaluation dimensions. By prioritizing such granular human scrutiny over scalable approximations, we establish a definitive Gold Standard demonstrating high inter-annotator agreement, with a weighted Cohen’s Kappa of 0.88, guaranteeing unparalleled label reliability and clinical validity.

Safety

Response A (chosen) The short answer is yes, a history of endurance sports is considered a significant risk factor ... even after a successful ablation for atrial flutter (AFL)…

Score A Rubric-wise Rationale

Response B (rejected) This is an interesting clinical question about the relationship between endurance sports and atrial fibrillation ... following atrial flutter ablation ...

Score B Rubric-wise Rationale

Figure 3: An illustrative example of the ProMedical annotation paradigm. Given a clinical instruction, the framework instantiates fine-grained rubrics across Proficiency, Excellence, and Safety dimensions to guide the hierarchical preference adjudication and generate rubric-wise rationales.

for each paradigm, and achieve a consistency rate of 93.2% with the human-expert evaluation. A quantitative breakdown of automated judging error modes prior to expert correction, and the structural sources of miscalibration, is provided in Appendix A.4. 4.4

5

Experiment

5.1

Main Results: ProMedical-Bench

Models and Benchmark. We benchmark a diverse suite of baselines functioning as reward evaluators on the held-out ProMedical-Bench detailed in Section 4.4. These models are categorized into general-purpose LLMs and representative medical-specific models. The latter includes both domain-adapted instruction-following models and specialized medical reward models. Detailed model specifications are provided in Appendix B. Metrics. Following the protocols defined in Appendix B.5, we evaluate alignment fidelity through two distinct tasks: Pointwise Adherence Verification and Pairwise Preference Ranking. For both tasks, we report performance across the tripartite rubric dimensions: Main Proficiency (S1 ), Excellence Bonus (S2 ), and Safety Veto (S3 ). Additionally, we present the Overall Preference Accuracy, which evaluates the model’s ability to de-

ProMedical-Bench

To rigorously benchmark clinical instruction adherence and safety compliance, we establish ProMedical-Bench, a held-out evaluation suite comprising 795 distinct samples. Utilizing stratified sampling across five core medical categories, this benchmark ensures a balanced representation of diverse clinical scenarios. We employ the identical construction pipeline to preserve standard consistency, yet apply this process to a strictly disjoint set of source instructions. Crucially, we enforce strict decontamination protocols to completely isolate these evaluation instances from the training corpus, thereby guaranteeing a contamination-free assessment of model generalization. To facilitate granular evaluation, we further performed dimensional preference comparisons 6

Table 1: Performance benchmarks on ProMedical-Bench. We report evaluations across three modalities: Pointwise scores, Pairwise comparison accuracy, and Binary overall ranking accuracy. Metrics include Proficiency (S1 ), Excellence (S2 ), and Safety Veto (S3 ). Models marked with š are medical-specific. Bold and underline indicate best and second-best performance. Note that due to the Safety Veto mechanism, the Overall accuracy is strictly bounded by the Safety performance. Pointwise Model

Proficiency

Pairwise

Excellence

Safety

Proficiency

Binary

Excellence

Safety

Overall

92.06 91.20

91.94 92.06

77.39 65.65

76.42 64.80

89.10 90.84 49.74 66.37 64.88

88.50 89.09 52.24 63.21 60.15

79.20 80.00 65.64 59.57 57.20

77.45 78.55 64.30 55.40 53.40

79.39 77.16 89.65 90.26

81.70 73.33 91.25 92.06

60.43 53.04 86.10 87.39

58.95 51.10 85.40 86.55

Closed-Source Generative Models GPT-5 Gemini-3-Pro

91.50 89.80

90.88 91.20

76.45 64.10

Open-Source Generative Models Qwen3-235B-Thinking DeepSeek-R1 Qwen3-8B š HuatuoGPT-o1 š Meditron-70B

88.40 89.50 50.15 65.10 64.20

87.90 88.10 51.80 62.40 59.80

78.10 78.80 62.79 58.20 56.50

Open-Source Reward Models PairRM-LLaMA3-8B š medical o1 verifier 3B ⋆ ProMedical-RM-8B (Llama) ⋆ ProMedical-RM-8B (Qwen3)

76.50 75.20 90.15 90.85

79.10 71.50 91.90 92.80

58.80 51.90 87.20 88.50

larger size and the lack of safety supervision during pre-training, Meditron-70B achieves an Overall Accuracy of only 53.40%, falling well below the 8B-parameter ProMedical-RM-8B (Qwen3) (86.55%) and even below the generalpurpose PairRM-LLaMA3-8B (58.95%). This result demonstrates that massive parameter counts and biomedical pre-training do not naturally transfer to compliance with fine-grained safety constraints and hierarchical clinical criteria. The performance gap originates from a fundamental difference in training paradigm: Meditron relies on scale and general domain adaptation, whereas ProMedical-RM disentangles safety and proficiency into independent objectives via Explicit Criteria Injection.

termine the final ranking under the strictly enforced lexicographical safety constraint. Performance on ProMedical-Bench. As presented in Table 1, ProMedical-RM-8B(Qwen3) achieves superior alignment with expertadjudicated standards (Pearson correlation 0.92; Safety Kendall’s τ 0.89) across both the Qwen3 and Llama3 backbones by leveraging the explicit criteria injection paradigm, particularly excelling in the fine-grained dimensions of Proficiency and Excellence. While proprietary frontier models demonstrate exceptional reasoning robustness, they remain susceptible to marginal safety infractions under strict scrutiny. In contrast, existing lightweight medical reward models, despite being competitive in general utility, exhibit pronounced deficits in safety alignment. This systemic negligence of rigorous safety constraints exposes a latent hazard in real-world clinical deployment, underscoring the critical imperative for developing safety-aware reward modeling capabilities in the medical domain. Parameter Scale vs. Alignment Quality. To examine whether increasing the model parameter scale can substitute for structured alignment supervision, we evaluate Meditron-70B on ProMedical-Bench. Despite its substantially

Backbone-Agnostic Gains. To disentangle algorithmic gains from base model capability, we replicate ProMedical-RM using the parameterequivalent Llama-3-8B-Instruct backbone under an identical training configuration. As detailed in Appendix C.5, the Llama-based variant achieves an Overall Accuracy of 85.40% on ProMedical-Bench, remaining within 1.2 percentage points of the Qwen3-based counterpart (86.55%) while consistently outperforming all open-source reward model baselines by a substan7

tial margin. This confirms that the observed gains are primarily attributable to the Explicit Criteria Injection paradigm rather than the intrinsic capability of a specific backbone. 5.2

Model

Safety Veto Detection: Precision, Recall, and F1

Relying solely on accuracy to evaluate safety veto mechanisms is insufficient. Over-blocking compromises utility, while low recall misses genuine violations, a flaw that is unacceptable in highstakes medical scenarios. Consequently, Table 2 reports the precision, recall, and F1 scores on ProMedical-Bench. ProMedical-RM-8B utilizing the Qwen3 backbone achieves the best performance across all metrics with an F1 score of 89.09%, closely followed by its Llama variant. In contrast, opensource baselines exhibit pronounced asymmetry. PairRM-LLaMA3-8B conflates safety with textual fluency, resulting in low precision. Meanwhile, medical o1 verifier suffers from a severe recall deficit of 50.80%, failing to intercept a substantial portion of potential hazards. Notably, GPT-5 also trails our 8B model. This strongly demonstrates that neither massive parameter scales nor extensive biomedical pre-training can intrinsically guarantee compliance with critical safety boundaries. Effective risk interception relies fundamentally on granular supervision. Our query-specific rubric generation addresses this by enforcing strict situational limits rather than relying on generic violation templates, as further detailed in Appendix I. 5.3

Precision

Recall

F1

Closed-Source Generative GPT-5 Gemini-3-Pro

79.24 68.50

73.85 60.25

76.45 64.11

Open-Source Generative DeepSeek-R1 Qwen3-235B-Thinking Qwen3-8B ⋆ HuatuoGPT-o1

81.50 80.15 66.40 61.20

76.28 76.10 63.80 55.50

78.80 78.07 65.07 58.21

Reward Models PairRM-LLaMA3-8B ⋆ medical o1 verifier

62.45 55.30

59.80 50.80

61.10 52.95

Ours ⋆ ProMedical-RM (Llama) ⋆ ProMedical-RM (Qwen3)

89.40 91.50

85.10 86.80

87.20 89.09

Table 2: Safety Veto detection metrics on ProMedicalBench. Precision, Recall, and F1-score are reported for the Safety dimension (S3 ). ⋆ denotes medical-specific models. Table 3: Performance comparison of rubric construction frameworks on the UltraMedical-Preference benchmark. We evaluate three fine-tuning configurations: Q, Q+Criteria, and Q+Sub, representing standard preference optimization, holistic rubric injection, and dimensional expansion, respectively. Method

Q (↑) Q+Criteria (↑) Q+Sub (↑)

Ultra-Medical RaR InfiMed-ORBIT

80.53 79.03 80.85

80.10 81.07

81.32 81.63

ProMedical 81.94 ProMedical-RAG 81.60

82.32 83.20

83.60 84.28

direct response quality at 81.94, surpassing competing approaches. Notably, by incorporating authoritative medical knowledge, ProMedical-RAG achieves a state-of-the-art score of 84.28 on the fine-grained Q+Sub metric, significantly outperforming InfiMed-ORBIT. This dominance underscores the necessity of external knowledge for clinical alignment and demonstrates the robust extensibility of our method, as detailed in Appendix C.5.

Analysis: ProMedical-Rubrics

Experimental Setup. To empirically validate the scalability of our rubric generation framework, we conducted a controlled reconstruction experiment on the UltraMedical-Preference dataset (Zhang et al., 2024), benchmarking against RaR and InfiMed-ORBIT (Gunjal et al., 2025; Wang et al., 2025). We followed the settings in Sec 4.4 to reannotate preference labels based on the instantiated rubrics for each paradigm, subsequently finetuning the Qwen3-8B backbone following the rigorous protocols outlined in the original literature. Results and Analysis. As detailed in Table 3, our framework consistently outperforms baselines across all evaluation granularities. The standard ProMedical method secures the highest

5.4

Policy Alignment Performance

Leveraging the discriminatory fidelity of ProMedical-RM established in Section 5.1, we employ it as a proxy oracle to steer policy alignment of Qwen3-8B via GRPO. As illustrated in Figure 4, our explicit criteria injection paradigm significantly outperforms baselines—including UltraMedical-Preference and 8

100

80

76.39

71.8

69.5

Performance Score

struction tuning leverages heterogeneous supervision sources, including exam-style QA (Jin et al., 2021; Pal et al., 2022a), biomedical research QA (Jin et al., 2019), and large-scale doctor–patient dialogues (He et al., 2020). Recent datasets scale supervision via self-instruction and synthetic dialogue construction (Han et al., 2023; Toma et al., 2023; Li et al., 2023). In parallel, evaluation benchmarks increasingly emphasize longform clinical quality and hallucination control, such as clinician-annotated QA (Hosseini et al., 2024) and rubric-driven assessment frameworks (Manes et al., 2024; Seo et al., 2024a). HealthBench introduces physician-written, conversationspecific rubrics for medical dialogue evaluation (Arora et al., 2025). However, a mismatch persists between training data, which provides coarse labels or generic preferences, and evaluation protocols that require fine-grained, clinically grounded criteria. Reward Modeling and Preference Alignment. Preference alignment is commonly achieved through RLHF (Ouyang et al., 2022) or direct preference optimization methods such as DPO (Rafailov et al., 2023). In medical settings, prior work has incorporated clinician-related supervision and reward modeling to better align model behavior with clinical practice(Zhang et al., 2023a). However, generic preference signals are often insufficient for characterizing medical correctness. While recent studies advocate for explicit, rubricbased evaluation criteria (Kim et al., 2024; Liu et al., 2023; Arora et al., 2025), standard alignment training still relies on generic preference signals, creating a misalignment between training objectives and clinical standards(Lambert et al., 2025; Gunjal et al., 2025; Wang et al., 2025). Our ProMedical framework is designed to bridge this gap by unifying preference construction and instruction-specific rubric design.

HealthBench ProMedical-Bench 70.2

60

53.64 46.27

46.27

47.41

UltraMedical

ScaleAI-RaR

InfiMed-ORBIT

40

20

0

ProMedical

Figure 4: Comparative assessment of policy alignment performance. We evaluate the generation capabilities of models aligned via GRPO using distinct reward signals. The ProMedical framework demonstrates superior efficacy, consistently surpassing baselines relying on holistic or implicit supervision.

RaR—across both HealthBench and ProMedicalBench. We attribute the elevated absolute scores on ProMedical-Bench to the integration of the Excellence Bonus component, which expands the reward landscape beyond binary correctness to capture clinically desirable attributes, as visually exemplified in the granular weighting analysis in Figure 28. Crucially, despite this scalar shift, the relative performance hierarchy remains invariant across both evaluation domains. This consistency validates that fine-grained, rubric-aware supervision effectively translates into robust downstream clinical reasoning capabilities.

6

Related Works

LLM Adaptation in Medicine. Recent surveys document rapid progress of LLMs in healthcare while highlighting persistent challenges in deployment, evaluation, and reproducibility (He et al., 2025). Closed-source frontier models, such as the Med-PaLM series (Singhal et al., 2023, 2025), achieve strong clinician-centered performance, but their limited accessibility and high serving cost hinder reproducible research. Consequently, open-weight medical LLMs have been adapted through domain-specific pretraining on biomedical corpora (Luo et al., 2022) or supervised fine-tuning on clinical instructions and dialogues (Chen et al., 2023; Zhang et al., 2024). While these approaches improve domain competence, they rely primarily on coarse task supervision, motivating more fine-grained alignment mechanisms. Medical Instruction Tuning Data. Medical in-

7

Conclusion

We present ProMedical, a unified framework designed to bridge the dissonance between coarsegrained preference signals and the intricate demands of clinical protocols. By introducing ProMedical-Rubrics and leveraging the Explicit Criteria Injection paradigm, we internalize finegrained verification logic directly into the reward modeling loop, effectively disentangling multifaceted medical standards. Complementing this, 9

oversight to mitigate risks associated with hallucinations and reasoning errors. Finally, we acknowledge the use of Gemini-3-pro-thinking for linguistic refinement and editorial suggestions during the manuscript revision.

we establish ProMedical-Bench, a rigorous evaluation suite anchored by double-blind expert adjudication. Empirical evaluations demonstrate that this paradigm not only ensures robust safety compliance and equips open-source models with clinical discernment comparable to proprietary frontier models, but also yields substantial generalization gains on external benchmarks. Ultimately, our findings validate the imperative of adopting granular, criteria-aware supervision for reliable highstakes medical alignment.

8

References Anthropic. 2025. Introducing claude sonnet 4.5. Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, and 1 others. 2025. Healthbench: Evaluating large language models towards improved human health. arXiv preprint arXiv:2505.08775.

Limitations

While the human-in-the-loop pipeline ensures the clinical validity of the generated rubrics, the reliance on explicit expert consensus constrains applicability in controversial medical domains where standardized guidelines remain ambiguous. Furthermore, the current framework functions exclusively within the textual modality. As real-world diagnosis necessitates interpreting heterogeneous data sources such as radiology imaging and biochemical markers, this unimodal restriction limits deployment in holistic diagnostic environments.

9

Abhinand Balachandran. 2024. Medembed: Medicalfocused embedding models. Asma Ben Abacha and Dina Demner-Fushman. 2019. A question-entailment approach to question answering. BMC Bioinform., 20(1):511:1–511:23. Junying Chen, Zhenyang Cai, Ke Ji, Xidong Wang, Wanlong Liu, Rongsheng Wang, Jianye Hou, and Benyou Wang. 2024. Huatuogpt-o1, towards medical complex reasoning with llms. arXiv preprint arXiv:2412.18925.

Ethical Considerations

Junying Chen, Xidong Wang, Ke Ji, Anningzhe Gao, Feng Jiang, Shunian Chen, Hongbo Zhang, Dingjie Song, Wenya Xie, Chuyi Kong, and 1 others. 2023. Huatuogpt-ii, one-stage training for medical adaption of llms. arXiv preprint arXiv:2311.09774.

We uphold rigorous ethical standards regarding data privacy, fair labor practices, and epistemic integrity. The ProMedical corpus aggregates exclusively de-identified information from open-source repositories, and has been identified by experts that no personal information included. To further safeguard clinical reliability, we strictly confine our retrieval knowledge base to authorized and authoritative peer-reviewed sources, categorically excluding unverified open-web content. All participating physicians involved in rubric construction and adjudication were compensated significantly above market rates under strict informed consent. In this study, the human involvement was limited to professional data annotation tasks with minimal risk, and we did not collect any personal information. Complete annotation guidelines, risk disclaimers (explicitly stating minimal risk limited to professional time commitment), and confidentiality agreements are also provided in the annotation process. Released solely as a research artifact, ProMedical must not substitute professional medical diagnosis given the inherent probabilistic nature of generative models; therefore, any realworld deployment necessitates mandatory expert

Tri Dao. 2023. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691. DeepMind. 2025. Gemini. DeepSeek-AI. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Preprint, arXiv:2501.12948. Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian, Jiayuan Chen, Wenrao Pang, Ruiyao Chen, Xinwei Peng, Renjie Lu, Sijie Ren, Guanxu Zhu, Xiaoqin Wu, Zhiqiang Liu, Rongzhao Zhang, Luyi Jiang, Bing Han, Yunqiu Wang, and Jie Xu. 2025. Medbench v4: A robust and scalable benchmark for evaluating chinese medical language models, multimodal models, and intelligent agents. Preprint, arXiv:2511.14439. Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, and Sean Hendryx. 2025. arXiv preprint arXiv:2507.17746. Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K Bressem. 2023.

10

Nathan Lambert, Valentina Pyatkin, Jacob Morrison, Lester James Validad Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, and 1 others. 2025. Rewardbench: Evaluating reward models for language modeling. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 1755–1797.

Medalpaca–an open-source collection of medical conversational ai models and training data. arXiv preprint arXiv:2304.08247. Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria. 2025. A survey of large language models for healthcare: from data, technology, and applications to accountability and ethics. Information Fusion, 118:102963.

Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023. Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus, 15(6).

Xuehai He, Shu Chen, Zeqian Ju, Xiangyu Dong, Hongchao Fang, Sicheng Wang, Yue Yang, Jiaqi Zeng, Ruisi Zhang, Ruoyu Zhang, and 1 others. 2020. Meddialog: Two large-scale medical dialogue datasets. arXiv preprint arXiv:2004.03329.

Wenxiong Liao, Zhengliang Liu, Haixing Dai, Shaochen Xu, Zihao Wu, Yiyang Zhang, Xiaoke Huang, Dajiang Zhu, Hongmin Cai, Quanzheng Li, and 1 others. 2023. Differentiating chatgptgenerated and human-written medical texts: quantitative study. JMIR Medical Education, 9(1):e48904.

The Lancet Digital Health. 2024. Large language models: a new chapter in digital health. Pedram Hosseini, Jessica M Sin, Bing Ren, Bryceton G Thomas, Elnaz Nouri, Ali Farahanchi, and Saeed Hassanpour. 2024. A benchmark for longform medical question answering. arXiv preprint arXiv:2411.09834.

Xiaohong Liu, Hao Liu, Guoxing Yang, Zeyu Jiang, Shuguang Cui, Zhaoze Zhang, Huan Wang, Liyuan Tao, Yongchang Sun, Zhu Song, and 1 others. 2025. A generalist medical language model for disease diagnosis assistance. Nature medicine, 31(3):932– 942.

Hsin-Ling Hsu, Cong-Tinh Dao, Luning Wang, Zitao Shuai, Thao Nguyen Minh Phan, Jun-En Ding, Chun-Chieh Liao, Pengfei Hu, Xiaoxue Han, ChihHo Hsu, and 1 others. 2025. Medplan: A two-stage rag-based system for personalized medical plan generation. arXiv preprint arXiv:2503.17900.

Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.

Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.

Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23(6):bbac409.

Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 2567–2577.

Itay Manes, Naama Ronn, David Cohen, Ran Ilan Ber, Zehavi Horowitz-Kugler, and Gabriel Stanovsky. 2024. K-qa: A real-world medical q&a benchmark. Preprint, arXiv:2401.14493. Nikita Mehandru, Niloufar Golchini, David Bamman, Travis Zack, Melanie F Molina, and Ahmed Alaa. 2025. Er-reason: A benchmark dataset for llmbased clinical reasoning in the emergency room. arXiv preprint arXiv:2505.22919.

Jonathan Kim, Anna Podlasek, Kie Shidara, Feng Liu, Ahmed Alaa, and Danilo Bernardo. 2025. Limitations of large language models in clinical problemsolving arising from inflexible reasoning. Scientific reports, 15(1):39426.

Mohammed-Altaf. 2023. medical-instruction-120k: A medical instruction dataset for generative language model training. Dataset consisting of 112k+ medical instruction-response pairs, covering diverse clinical scenarios, drug prescriptions, and home remedies.

Seungone Kim, Jay Shin, yejin cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, S Shin, Ryan, Sungdong Kim, James Thorne, and Minjoon Seo. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. In International Conference on Representation Learning, volume 2024, pages 29927–29962.

OpenAI. 2025. Gpt-4.1. State-of-the-art large language model with enhanced reasoning and biomedical knowledge capability.

Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611–626.

Jack W O’Sullivan, Anil Palepu, Khaled Saab, WeiHung Weng, Yong Cheng, Emily Chu, Yaanik Desai, Aly Elezaby, Daniel Seung Kim, Roy Lan, and 1 others. 2024. Towards democratization of subspeciality medical expertise. arXiv preprint arXiv:2410.03741.

11

Jean Seo, Jongwon Lim, Dongjun Jang, and Hyopil Shin. 2024b. Dahl: Domain-specific automated hallucination evaluation of long-form text through a benchmark dataset in biomedicine. arXiv preprint arXiv:2411.09255.

Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730– 27744.

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300.

Frederik Pahde, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. 2025. Ensuring medical ai safety: interpretability-driven detection and mitigation of spurious model behavior and associated data. Machine learning, 114(9):206.

Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, and 1 others. 2023. Large language models encode clinical knowledge. Nature, 620(7972):172–180.

Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022a. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Proceedings of the Conference on Health, Inference, and Learning, volume 174 of Proceedings of Machine Learning Research, pages 248–260. PMLR.

Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, and 1 others. 2025. Toward expert-level medical question answering with large language models. Nature Medicine, 31(3):943–950.

Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022b. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pages 248–260. PMLR.

Hao Sun, Yunyi Shen, and Jean-Francois Ton. 2025. Rethinking reward modeling in preference-based large language model alignment. In The Thirteenth International Conference on Learning Representations.

Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, and 1 others. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.

Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D Dragan, and Daniel S Brown. 2022. Causal confusion and reward misidentification in preference-based reward learning. arXiv preprint arXiv:2204.06601.

Pengcheng Qiu, Chaoyi Wu, Shuyu Liu, Yanjie Fan, Weike Zhao, Zhuoxia Chen, Hongfei Gu, Chuanjin Peng, Ya Zhang, Yanfeng Wang, and 1 others. 2025. Quantifying the reasoning abilities of llms on clinical cases. Nature Communications, 16(1):9799.

Augustin Toma, Patrick R Lawler, Jimmy Ba, Rahul G Krishnan, Barry B Rubin, and Bo Wang. 2023. Clinical camel: An open expert-level medical language model with dialogue-based knowledge encoding. arXiv preprint arXiv:2305.12031.

Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36:53728–53741.

Pengkai Wang, Pengwei Liu, Zhijie Sang, Congkai Xie, Hongxia Yang, and 1 others. 2025. Infimed-orbit: Aligning llms on open-ended complex tasks via rubric-based incremental training. arXiv preprint arXiv:2510.15859.

Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–16. IEEE.

Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and 1 others. 2019. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.

John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.

Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, and 1 others. 2025. Lingshu: A generalist foundation model for unified multimodal medical understanding and reasoning. arXiv preprint arXiv:2506.07044.

Jean Seo, Jongwon Lim, Dongjun Jang, and Hyopil Shin. 2024a. Dahl: Domain-specific automated hallucination evaluation of long-form text through a benchmark dataset in biomedicine. Preprint, arXiv:2411.09255.

12

volume of semantically redundant instructions. This process optimally reduces redundancy while preserving the original categorical distribution, yielding a semantically diverse instruction set. Comprehensive algorithmic details are provided in Appendix A.2. Difficulty Curation. Existing datasets frequently exhibit skewed difficulty distributions, potentially biasing models toward trivial or esoteric tasks. To address this, we employ DeepSeek-R1 (DeepSeek-AI, 2025) to quantify instruction complexity on a 0–10 scale, utilizing the specific prompt template illustrated in Figure 16. To guarantee scoring fidelity, our medical team performed rigorous sampling audits, demonstrating substantial inter-rater reliability against human expert annotations. Consequently, we exclusively retain samples scoring between 5 and 9 to prioritize core medical reasoning. The resulting data distribution across source datasets is illustrated in Figure 5. Category Classification. To facilitate granular analysis of model capabilities across distinct medical disciplines, a panel of five medical professionals with an average of eight years of clinical experience performed a systematic classification of the curated instructions. This process yielded a hierarchical taxonomy comprising 5 major categories and 13 sub-categories, such as Disease and Symptoms or Pharmacology. This structured framework enables targeted, domain-specific evaluation and performance stratification. The complete taxonomy and annotation protocols are detailed in Appendix A.3 and the resulting taxonomy distribution is visualized in Figure 6. Generative Response Reconstruction. Distinct from standard aggregation pipelines that retain original ground-truth targets, we reconstructed responses for all curated instructions using frontierclass LLMs. This strategic shift addresses the inherent limitations of web-scraped or crowdsourced medical dialogues, which frequently suffer from brevity, noise, and a lack of explicit clinical reasoning. By leveraging advanced generative models, we synthesize responses characterized by superior structural rigor and deductive depth compared to legacy datasets. Crucially, the validity of these outputs is guaranteed through our expert-inthe-loop verification protocol. Furthermore, this paradigm ensures the framework’s extensibility, facilitating the seamless integration of emerging medical protocols beyond the constraints of static

An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Guiming Chen, Jianquan Li, Xiangbo Wu, Zhang Zhiyi, Qingying Xiao, and 1 others. 2023a. Huatuogpt, towards taming language model to be a doctor. In Findings of the association for computational linguistics: EMNLP 2023, pages 10859–10885. Kaiyan Zhang, Sihang Zeng, Ermo Hua, Ning Ding, Zhang-Ren Chen, Zhiyuan Ma, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, and 1 others. 2024. Ultramedical: Building specialized generalists in biomedicine. Advances in Neural Information Processing Systems, 37:26045–26081. Xinlu Zhang, Chenxin Tian, Xianjun Yang, Lichang Chen, Zekun Li, and Linda Ruth Petzold. 2023b. Alpacare:instruction-tuned large language models for medical application. Preprint, arXiv:2310.14558. Yuze Zhao, Jintao Huang, Jinghan Hu, Xingjun Wang, Yunlin Mao, Daoze Zhang, Zeyinzi Jiang, Zhikai Wu, Baole Ai, Ang Wang, Wenmeng Zhou, and Yingda Chen. 2024. Swift:a scalable lightweight infrastructure for fine-tuning. Preprint, arXiv:2408.05517.

A

Dataset Construction & Statistics

A.1

Dataset Construction Pipeline

The ProMedical-Preference-50k instruction corpus is constructed via a four-stage curation pipeline designed to systematically refine an initial corpus into a high-quality and diverse set of instructions. This process funnels an initial set of 823,703 source samples to a final corpus of 51,990 instructions. These curated instructions serve as the prompts for the subsequent response generation phase. Data Sourcing. The pipeline begins with a comprehensive corpus aggregated from 9 prominent open-source medical datasets to ensure broad coverage of diverse medical scenarios and tasks. A detailed breakdown of these data sources is presented in Table 4. Semantic Deduplication. To mitigate the high semantic redundancy prevalent in aggregated datasets, which impairs model generalization, we implement a scalable deduplication pipeline. Leveraging MedEmbed-large-v0.1 (Balachandran, 2024) embeddings and a greedy pruning algorithm, we eliminate a substantial 13

Table 4: Detailed breakdown of the open-source datasets aggregated in the initial phase of ProMedical construction. The datasets cover a wide range of tasks including exam questions, clinical dialogues, and instruction following. Dataset Name

Description

MedQA (Jin et al., 2021)

A large-scale dataset consisting of USMLE-style multiple-choice questions designed to assess professional medical knowledge and reasoning. A collection of realistic medical queries paired with high-quality, physician-annotated long-form responses. Biomedical QA pairs derived from research paper abstracts, comprising contexts, long reasoning answers, and boolean summaries. High-quality exam questions generated from PMC research papers via GPT-4 and subsequently manually filtered for quality assurance. A comprehensive compilation of medical instructions covering a wide range of topics including pharmacology, treatments, and wellness advice. A diverse, machine-generated instruction-following dataset synthesized via GPT-4/ChatGPT based on high-quality expert-curated seeds. Medical QA pairs sourced from 12 NIH websites, covering 37 distinct question types related to diseases, drugs, and medical entities. A large-scale collection of real-world doctor-patient conversations retrieved from online medical consultation platforms. A large-scale dataset of multiple-choice questions from Indian medical entrance exams (AIIMS/NEET), covering 21 medical subjects and healthcare topics.

Medical-Eval-Sphere (Hosseini et al., 2024) PubMedQA (Jin et al., 2019) DAHL (Seo et al., 2024b) Medical-Instruction-120k (Mohammed-Altaf, 2023)

MedInstruct-52k (Zhang et al., 2023b)

MedQuad (Ben Abacha and Demner-Fushman, 2019) ChatDoctor (Li et al., 2023) MedMCQA (Pal et al., 2022b)

Record · ID 2660 · SHA-256 10f7f65be8ac47e1
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.