Benchmark Review

HealthBench Professional

Reviewed: Sep. 10, 2026

Benchmark creator: OpenAI

Verdict: Flawed

HealthBench Professional is a benchmark that builds on OpenAI’s HealthBench. It aims to evaluate whether LLMs give useful and accurate answers on real healthcare-related tasks that clinicians brought to ChatGPT in the course of their work. We randomly sampled 50 tasks, stratified by source slice, and found that 25 (50% of the sample) presented clear errors. Of these, at least 14 (28% of the sample) were structural: their issues are not related to medical accuracy. The other 11 had issues clearly identifiable with minimal domain knowledge. 4 (8% of the sample) were impossible to answer as written or ill-posed, 17 (34% of the sample) contained false negatives, 8 (16% of the sample) contained false positives, and 6 (12% of the sample) did not measure the claimed capability.1 A full review by medical professionals would have a broader scope and likely flag additional problems, but given the high error rate in our random sample, we are confident the true rate is at least 20%, resulting in a Flawed verdict.

Methodology

We randomly sampled 50 tasks out of 525 stratified by three categories,2 all represented proportionally: 24 out of 256 (9.4%) for Good Faith Typical,3 18 out of 191 (9.4%) for Red Teaming Difficult,4 and 8 out of 78 (10.3%) for Good Faith Difficult. We then prompted Claude Fable 5.1 and GPT-6 Astra to flag any of the following issues, as defined in our rubric: Ill-posed or Impossible to solve; False Negatives; False Positive; Wrong Measured Capability.

Finally, we manually reviewed all errors. A minority did require minimal medical knowledge to verify, such as task a5778c7ecdb4eeccf9d252631e18a274, which requires knowing broad test scoring norms.

A HealthBench Professional score is built in three phases.

  1. Each task starts from a prompt a clinician actually wrote, using the model as they would in their own clinical or research work.
  2. Other physicians write the rubric: they reconstruct the prompter’s intent and encode their own judgement, assigning a weight to each criterion that a correct answer should contain or avoid.5
  3. An LLM judge (GPT-5.4 at Low reasoning effort) then checks each criterion independently and assigns a binary pass/fail for that criterion.

Representative Errors

Impossible questions may be ill-posed, such as asking for interpretation of a test without giving details such as the test version or age norms, or asking a model to provide the formulation of a proprietary product’s composition.

False Negatives include no credit for asking for missing information, penalties for challenging one physician’s dubious views as fact, and penalties for content that is supported by information established in the task’s earlier turns.

False Positives can include tasks that check form rather than content: we found examples of cases where any citation passes (including fabricated ones), and where declining to write the requested documents earns a 100% score.

Wrong Measured Capability includes tasks where, at least partially, the rubric scores adjacent facts while leaving the actual question unscored. Examples include tasks involving translation or penalization for formatting tendencies.

A few structural choices may make this benchmark more prone to content-related errors, which we did not exhaustively assess:

  • Since tasks are taken from prompts clinicians actually wrote, they may be vague, ill-posed, or under-specified.
  • Since other physicians write the rubric, biases and errors can be inadvertently included due to differences in standards of care and differences of opinion where a ground truth is not established
  • The LLM judge is limited in giving a pass/fail verdict, so it cannot flexibly credit partial merit where available.

Errors

To prevent contamination, benchmark authors request that examples from the dataset not be published in plain text or images online. We respect their request by providing examples of classes of real errors in Representative Errors and LLM-summarized Notes here.

Error types: 1) Impossible, 2) False Negative, 3) False Positive, 4) Wrong Measured Capability.

Task IDError typeNotesDefect category
a5778c7ecdb4eeccf9d252631e18a274Impossible, False NegativePrompt’s numeric inputs are internally inconsistent and the reference norm is unstated.Clinical
c70db1fcf0a0e2e2b86e19141c7ab39eFalse Positive, Wrong Measured CapabilityThe clinician asks for drug-interaction data. The sole OR-criterion scores adjacent content and never checks the interaction claims, so a hallucinated claim passes.Structural
83a295db73b1bc551ed5941434655de1False NegativePenalty treats a comparison between two techniques as settled when the guideline it rests on takes no position; a positive criterion rewards citations without checking them.Clinical
35b9ab1e8dc0ef2244dde26f4872d185False NegativeThe positive criterion requires one specific anatomical answer, but the correct answer is conditional on a detail the prompt omits; asking for it earns nothing.Clinical
c3fc32bce9986b51c688e46c988116feFalse Positive, Wrong Measured CapabilityThe clinician asks for a dose but the dose is never scored, so a dangerous value with a generic risk caveat gets full marks.Structural
503b75b4cba8e5abd656cd2512a449bbFalse NegativePenalizes a choice that is defensible once the relevant lab is checked, and rewards only alternatives that are ill-suited to the scenario while omitting the standard choice.Clinical
282e788a7f8be989ae55aa236fcd9f88False PositiveBoth criteria are satisfied by any response of the requested genre; content accuracy is unchecked.Structural
99750e7472125ff9f35a896e532bb77aFalse PositiveThe only criterion rewards spotting an inconsistency; the requested document has zero weight, so declining to write it scores 100%.Structural
54f98e63dbdd317b1a6bf3d8f84f7bf0False NegativeAll-of-four enumeration includes a niche therapy; a sound first-line answer that omits it scores zero on that item.Clinical
6d7385494242595c03cdc6d2812f4a3dFalse Negative, False Positive, Wrong Measured CapabilityRequested dosing is never checked. Penalties include using a citation twice and citing a drug that is part of the standard regimen.Structural
d09add4f588755d3ccddfe9a7226993dFalse Negative, Wrong Measured CapabilityRubric scores two peripheral features of the condition rather than the specific list the clinician asked for; a precise answer can score zero.Structural
dbec60751876c6e24fbd07c5fab139e9False NegativeHonest uncertainty risks the penalty, and the ground truth is a moving target.Structural
7b30b74bfdd1d0e781a5d1ff9430a6cbImpossibleA positive criterion rewards responding to a patient characteristic that does not appear anywhere in the prompt (apparent rubric/prompt version mismatch).Structural
de721391f646b3844d7a52fa7e4e5511False NegativeThe task awards zero credit to an answer containing correct, directly relevant timing advice because it omits additional unasked-for context.Clinical
7113e3e9fa7c36ba46ddb6b6bc76d94aFalse Negative, Wrong Measured CapabilityPenalizes a position the earlier turns’ own cited evidence supports. Rewards recommending a test already performed in a prior turn; the final question is unscored.Structural
c6b2da2a85e39bd9b9018cf135476cf0Impossible, False NegativeDemands one exact score, but one of the clinician’s inputs maps to two adjacent point values, so the physician’s equally valid total scores zero.Clinical
ff227f1cf3089a77df73966ccb24ea4cFalse PositiveOperative technique scored only on “provides a stepwise description”; any numbered list passes, wrong steps cost nothing.Structural
696a797226704ff41bd9ef444778d73eImpossible, False NegativeDemands one exact risk score while leaving two scoring items unspecified in the prompt; several totals are defensible and only one is rewarded.Clinical
ca453ab13aa0ee164387f4cd8ef2e197False PositiveNaming any one of several generic approaches passes; answering in the wrong language is unpenalized.Structural
0535beb34f00feaeb7fe28c697caa476False NegativeOne criterion requires knowledge of a national policy change from the current year; a terminology penalty fires on wording that official national sources themselves use.Clinical
050d51fae49ae14008ad94db8b035e45False NegativeThe only positive criterion overstates guidance (avoid vs. screen-first) and is tangential to the question; a careful correct answer scores zero.Clinical
f3c8ffa378bf114ce33d694db43d8412False Negative, Wrong Measured CapabilityAll criteria hinge on a very recent policy disagreement between two national bodies; a model with an earlier cutoff loses everything for stating the prior consensus.Structural
458c5d4af48964eb26cf49f27dcdf968False NegativePenalty is broad enough to catch harm-reduction language in a good answer that a specialist would include.Clinical
3b6768c0c7ec45fe4fb1aa6d9a4a8c8eFalse Positive”References a case report from country X” is satisfied by a fabricated citation as readily as a real one.Structural
924b069395928cd26576cd3fb1aceab0False NegativeOne rubric item is two criteria merged by a formatting error (stray point marker), one of which is a subjective comparative judgement.Structural

Other Tasks

Below is the list of the unflagged 25 task IDs. These tasks may still present errors that a medical professional might flag, but they passed our analysis.

  • 4f08ae480b16ef825cf098eca6530e68
  • e156871820fef362c392aedfc8429c48
  • c820187649a64fca4be8aa237bd93fd6
  • 5b139efb2c8c5881d782135358d5ec25
  • dc4f101b1b9ead76d90067debe5c7d36
  • 41ba225b4e2a8210c1d03112a34c0ab1
  • 7791cf509bfaa1defdc4efaec9b29eb9
  • bdde297153ec143877b82dd9dd47b04f
  • c443f1a285bc015ef677e65ea361113d
  • 5060922f2c0e9525177a1d251594bb9a
  • 05569de0534eb183cc9935ff9498a7cf
  • 6c393aec7444ed28f3db46c33459fa15
  • f3f1182e86b37be585cfb0af4eb44ccd
  • ac7ce5c7f6b5fb860cd2feb7f9e4a6fe
  • eec81e8363a30d53edf8ed8b55bd535c
  • 2346e3e2d47f1326ba63811a6a4d0910
  • f3be6c11c68c492d3ef08af1be07f547
  • 78b02f28913d81712b01941bec4e9660
  • 7b0ad484ccf93f69612c2533d044c4b7
  • 40346f3978ae08ebb99213b99290fa90
  • 68ea82a57829bbb36ba708868b818a83
  • 4f8e076d432c53c2c56fb24215143490
  • 35b3491359ab7e2232c335e8c5b01ed2
  • fe90fa8aecea9ecd618059ca246e149b
  • f03edbfec507df466d00e5eb0364eefb

Rubric

1. Reviewability

LevelMeaningStatus
FullAll tasks and scoring logic inspectable*, and harness/API settings used for each model (reasoning effort, token/time limits, tool access, system prompts) are fully disclosed [proceed to 2]
PartialA representative sample of tasks and scoring logic inspectable, with full harness/API settings disclosed [proceed to 2]
InadequateLimited, biased, or no inspectable tasks or scoring logic, harness/API settings undisclosed [stop → NEI]

* Either publicly or privately to reviewers.

2. Scoring

This is the minimum standard required to be Verified; failing any of these items results in a Flawed verdict.

Scoring
Examples
  • Essentially impossible to correctly answer as written (e.g. underspecified task, hidden requirement, missing file/tool)
  • False Negative (e.g. overly strict scorer, stale/incorrect ground truth, dependent on live external state that can drift, sandbox failure independent of the agent)
  • False Positive (e.g. lax scorer, reward-hackable environment, stated skill can be bypassed via a shortcut such as exploiting an error in the scoring logic or retrieving the answer from the harness/web)
  • Task egregiously doesn’t measure the claimed capability
Default threshold for Flawed≥20% of inspected sample† contains errors or there is an issue that corrupts grading at scale
Status
Pass Flag [stop → Flawed] Not reviewed
Notes
Benchmark Consistency
Examples
  • Scorer, instructions, or ground truth changed without a version bump
Default threshold for FlawedLeaderboard has incomparable results from different versions
Status
Pass Flag [stop → Flawed] Not reviewed
Notes
Elicitation
Examples

Model elicitation is extremely constraining and is not the focus of the benchmark:

  • Under-resourced relative to task size (token/turn/time limit, sandbox resources)
  • Poor context management
  • Lack of agentic environment where it would be natural to provide one
  • Excessive non-voluntary termination for agentic benchmarks
Default threshold for FlawedSubstantially reduced performance compared to reasonable alternatives for the tasks
Status
Pass Flag [stop → Flawed] Not reviewed
Notes
Bias in Evaluation Setup
Examples
  • Uneven compute/token budgets
  • Unfair scaffold choice (e.g. only a subset of models optimized)
Default threshold for FlawedMaterial model-specific advantage found
Status
Pass Flag [stop → Flawed] Not reviewed
Notes

† For benchmarks with >50 available tasks, we will sample a random set of 50 (stratified by category, when present). If the observed error rate is 15–25%, we will expand the sample to 100. For benchmarks with ≤50 available tasks, all available tasks will be assessed.

3. Evaluation Quality

This is the standard we would like all benchmarks to meet, but it is not necessarily disqualifying to omit or fail these items.

QuestionStatusNotes
Elicitation and resource adequacy: Are the resources given to models (reasoning token/turn budget, tool access, etc.) sufficient for them to perform near their ceiling?
Sufficient Constraining Unreasonably constraining Unknown Not reviewed
Scaffold fairness: What scaffold does the leaderboard report?
Shared common scaffold Mix of model-specific and common scaffolds Model-specific scaffolds Not reviewed
Is there evidence/risk of contamination?
Not reviewed
Has human completability been assessed?
All tasks Representative set of tasks Poor implementation (Unrepresentative set of tasks, unreasonable set of participants) Not established Not reviewed
Score range (if possible to estimate)
— floor, — ceiling Not reviewed
Statistical adequacy: how many runs/model (≥ 5 recommended for error bars)
— runs/model Unknown Not reviewed
Construct Validity
Measures stated capabilities Partially measures stated capabilities Does not measure stated capabilities Not reviewed

Disclaimer

This is a partial review. We stopped inspecting once we discovered errors that exceeded our published threshold for a flawed benchmark.

We inspected 50 tasks, sampled at random, stratified by category. Tasks not inspected may contain additional defects.

We attempted to reach out to the benchmark developer prior to launch, and if they write a response to this review we will link it here. If you’re the benchmark developer, please email us at reviews@epoch.ai with a link to your response if you would like us to include it here.

Notes
  1. A task can carry more than one flag, therefore categories overlap and sum isn’t 50. Return

  2. The authors refer to this categorization as “by source slice(s)”, in Section 3.5. Return

  3. Good faith questions are real prompts clinicians asked models during the course of their practice. These are further split based on the expert’s perceived difficulty. Return

  4. Red teaming questions are adversarial prompts designed to probe weaknesses. Return

  5. 191 out of 525 (36.4%) tasks have negatively weighted criteria: a model will get penalized if it includes information a physician might consider potentially wrong or dangerous. Return

About Benchmark Reviews

Epoch AI’s Benchmark Reviews are independent reviews of external AI benchmarks. Our documentation describes the rubric behind this verdict, how we choose which benchmarks to review, and answers frequently asked questions.

Go to sourceRead the documentation and FAQs