HealthBench Professional is a benchmark that builds on OpenAI’s HealthBench. It aims to evaluate whether LLMs give useful and accurate answers on real healthcare-related tasks that clinicians brought to ChatGPT in the course of their work. We randomly sampled 50 tasks, stratified by source slice, and found that 25 (50% of the sample) presented clear errors. Of these, at least 14 (28% of the sample) were structural: their issues are not related to medical accuracy. The other 11 had issues clearly identifiable with minimal domain knowledge. 4 (8% of the sample) were impossible to answer as written or ill-posed, 17 (34% of the sample) contained false negatives, 8 (16% of the sample) contained false positives, and 6 (12% of the sample) did not measure the claimed capability.1 A full review by medical professionals would have a broader scope and likely flag additional problems, but given the high error rate in our random sample, we are confident the true rate is at least 20%, resulting in a Flawed verdict.
Methodology
We randomly sampled 50 tasks out of 525 stratified by three categories,2 all represented proportionally: 24 out of 256 (9.4%) for Good Faith Typical,3 18 out of 191 (9.4%) for Red Teaming Difficult,4 and 8 out of 78 (10.3%) for Good Faith Difficult. We then prompted Claude Fable 5.1 and GPT-6 Astra to flag any of the following issues, as defined in our rubric: Ill-posed or Impossible to solve; False Negatives; False Positive; Wrong Measured Capability.
Finally, we manually reviewed all errors. A minority did require minimal medical knowledge to verify, such as task a5778c7ecdb4eeccf9d252631e18a274, which requires knowing broad test scoring norms.
A HealthBench Professional score is built in three phases.
- Each task starts from a prompt a clinician actually wrote, using the model as they would in their own clinical or research work.
- Other physicians write the rubric: they reconstruct the prompter’s intent and encode their own judgement, assigning a weight to each criterion that a correct answer should contain or avoid.5
- An LLM judge (GPT-5.4 at Low reasoning effort) then checks each criterion independently and assigns a binary pass/fail for that criterion.
Representative Errors
Impossible questions may be ill-posed, such as asking for interpretation of a test without giving details such as the test version or age norms, or asking a model to provide the formulation of a proprietary product’s composition.
False Negatives include no credit for asking for missing information, penalties for challenging one physician’s dubious views as fact, and penalties for content that is supported by information established in the task’s earlier turns.
False Positives can include tasks that check form rather than content: we found examples of cases where any citation passes (including fabricated ones), and where declining to write the requested documents earns a 100% score.
Wrong Measured Capability includes tasks where, at least partially, the rubric scores adjacent facts while leaving the actual question unscored. Examples include tasks involving translation or penalization for formatting tendencies.
A few structural choices may make this benchmark more prone to content-related errors, which we did not exhaustively assess:
- Since tasks are taken from prompts clinicians actually wrote, they may be vague, ill-posed, or under-specified.
- Since other physicians write the rubric, biases and errors can be inadvertently included due to differences in standards of care and differences of opinion where a ground truth is not established
- The LLM judge is limited in giving a pass/fail verdict, so it cannot flexibly credit partial merit where available.
Errors
To prevent contamination, benchmark authors request that examples from the dataset not be published in plain text or images online. We respect their request by providing examples of classes of real errors in Representative Errors and LLM-summarized Notes here.
Error types: 1) Impossible, 2) False Negative, 3) False Positive, 4) Wrong Measured Capability.
| Task ID | Error type | Notes | Defect category |
|---|---|---|---|
| a5778c7ecdb4eeccf9d252631e18a274 | Impossible, False Negative | Prompt’s numeric inputs are internally inconsistent and the reference norm is unstated. | Clinical |
| c70db1fcf0a0e2e2b86e19141c7ab39e | False Positive, Wrong Measured Capability | The clinician asks for drug-interaction data. The sole OR-criterion scores adjacent content and never checks the interaction claims, so a hallucinated claim passes. | Structural |
| 83a295db73b1bc551ed5941434655de1 | False Negative | Penalty treats a comparison between two techniques as settled when the guideline it rests on takes no position; a positive criterion rewards citations without checking them. | Clinical |
| 35b9ab1e8dc0ef2244dde26f4872d185 | False Negative | The positive criterion requires one specific anatomical answer, but the correct answer is conditional on a detail the prompt omits; asking for it earns nothing. | Clinical |
| c3fc32bce9986b51c688e46c988116fe | False Positive, Wrong Measured Capability | The clinician asks for a dose but the dose is never scored, so a dangerous value with a generic risk caveat gets full marks. | Structural |
| 503b75b4cba8e5abd656cd2512a449bb | False Negative | Penalizes a choice that is defensible once the relevant lab is checked, and rewards only alternatives that are ill-suited to the scenario while omitting the standard choice. | Clinical |
| 282e788a7f8be989ae55aa236fcd9f88 | False Positive | Both criteria are satisfied by any response of the requested genre; content accuracy is unchecked. | Structural |
| 99750e7472125ff9f35a896e532bb77a | False Positive | The only criterion rewards spotting an inconsistency; the requested document has zero weight, so declining to write it scores 100%. | Structural |
| 54f98e63dbdd317b1a6bf3d8f84f7bf0 | False Negative | All-of-four enumeration includes a niche therapy; a sound first-line answer that omits it scores zero on that item. | Clinical |
| 6d7385494242595c03cdc6d2812f4a3d | False Negative, False Positive, Wrong Measured Capability | Requested dosing is never checked. Penalties include using a citation twice and citing a drug that is part of the standard regimen. | Structural |
| d09add4f588755d3ccddfe9a7226993d | False Negative, Wrong Measured Capability | Rubric scores two peripheral features of the condition rather than the specific list the clinician asked for; a precise answer can score zero. | Structural |
| dbec60751876c6e24fbd07c5fab139e9 | False Negative | Honest uncertainty risks the penalty, and the ground truth is a moving target. | Structural |
| 7b30b74bfdd1d0e781a5d1ff9430a6cb | Impossible | A positive criterion rewards responding to a patient characteristic that does not appear anywhere in the prompt (apparent rubric/prompt version mismatch). | Structural |
| de721391f646b3844d7a52fa7e4e5511 | False Negative | The task awards zero credit to an answer containing correct, directly relevant timing advice because it omits additional unasked-for context. | Clinical |
| 7113e3e9fa7c36ba46ddb6b6bc76d94a | False Negative, Wrong Measured Capability | Penalizes a position the earlier turns’ own cited evidence supports. Rewards recommending a test already performed in a prior turn; the final question is unscored. | Structural |
| c6b2da2a85e39bd9b9018cf135476cf0 | Impossible, False Negative | Demands one exact score, but one of the clinician’s inputs maps to two adjacent point values, so the physician’s equally valid total scores zero. | Clinical |
| ff227f1cf3089a77df73966ccb24ea4c | False Positive | Operative technique scored only on “provides a stepwise description”; any numbered list passes, wrong steps cost nothing. | Structural |
| 696a797226704ff41bd9ef444778d73e | Impossible, False Negative | Demands one exact risk score while leaving two scoring items unspecified in the prompt; several totals are defensible and only one is rewarded. | Clinical |
| ca453ab13aa0ee164387f4cd8ef2e197 | False Positive | Naming any one of several generic approaches passes; answering in the wrong language is unpenalized. | Structural |
| 0535beb34f00feaeb7fe28c697caa476 | False Negative | One criterion requires knowledge of a national policy change from the current year; a terminology penalty fires on wording that official national sources themselves use. | Clinical |
| 050d51fae49ae14008ad94db8b035e45 | False Negative | The only positive criterion overstates guidance (avoid vs. screen-first) and is tangential to the question; a careful correct answer scores zero. | Clinical |
| f3c8ffa378bf114ce33d694db43d8412 | False Negative, Wrong Measured Capability | All criteria hinge on a very recent policy disagreement between two national bodies; a model with an earlier cutoff loses everything for stating the prior consensus. | Structural |
| 458c5d4af48964eb26cf49f27dcdf968 | False Negative | Penalty is broad enough to catch harm-reduction language in a good answer that a specialist would include. | Clinical |
| 3b6768c0c7ec45fe4fb1aa6d9a4a8c8e | False Positive | ”References a case report from country X” is satisfied by a fabricated citation as readily as a real one. | Structural |
| 924b069395928cd26576cd3fb1aceab0 | False Negative | One rubric item is two criteria merged by a formatting error (stray point marker), one of which is a subjective comparative judgement. | Structural |
Other Tasks
Below is the list of the unflagged 25 task IDs. These tasks may still present errors that a medical professional might flag, but they passed our analysis.
- 4f08ae480b16ef825cf098eca6530e68
- e156871820fef362c392aedfc8429c48
- c820187649a64fca4be8aa237bd93fd6
- 5b139efb2c8c5881d782135358d5ec25
- dc4f101b1b9ead76d90067debe5c7d36
- 41ba225b4e2a8210c1d03112a34c0ab1
- 7791cf509bfaa1defdc4efaec9b29eb9
- bdde297153ec143877b82dd9dd47b04f
- c443f1a285bc015ef677e65ea361113d
- 5060922f2c0e9525177a1d251594bb9a
- 05569de0534eb183cc9935ff9498a7cf
- 6c393aec7444ed28f3db46c33459fa15
- f3f1182e86b37be585cfb0af4eb44ccd
- ac7ce5c7f6b5fb860cd2feb7f9e4a6fe
- eec81e8363a30d53edf8ed8b55bd535c
- 2346e3e2d47f1326ba63811a6a4d0910
- f3be6c11c68c492d3ef08af1be07f547
- 78b02f28913d81712b01941bec4e9660
- 7b0ad484ccf93f69612c2533d044c4b7
- 40346f3978ae08ebb99213b99290fa90
- 68ea82a57829bbb36ba708868b818a83
- 4f8e076d432c53c2c56fb24215143490
- 35b3491359ab7e2232c335e8c5b01ed2
- fe90fa8aecea9ecd618059ca246e149b
- f03edbfec507df466d00e5eb0364eefb
Rubric
1. Reviewability
| Level | Meaning | Status |
|---|---|---|
| Full | All tasks and scoring logic inspectable*, and harness/API settings used for each model (reasoning effort, token/time limits, tool access, system prompts) are fully disclosed | [proceed to 2] |
| Partial | A representative sample of tasks and scoring logic inspectable, with full harness/API settings disclosed | [proceed to 2] |
| Inadequate | Limited, biased, or no inspectable tasks or scoring logic, harness/API settings undisclosed | [stop → NEI] |
* Either publicly or privately to reviewers.
2. Scoring
This is the minimum standard required to be Verified; failing any of these items results in a Flawed verdict.
| Scoring | |
|---|---|
| Examples |
|
| Default threshold for Flawed | ≥20% of inspected sample† contains errors or there is an issue that corrupts grading at scale |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
| Benchmark Consistency | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Leaderboard has incomparable results from different versions |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
| Elicitation | |
|---|---|
| Examples | Model elicitation is extremely constraining and is not the focus of the benchmark:
|
| Default threshold for Flawed | Substantially reduced performance compared to reasonable alternatives for the tasks |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
| Bias in Evaluation Setup | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Material model-specific advantage found |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | — |
† For benchmarks with >50 available tasks, we will sample a random set of 50 (stratified by category, when present). If the observed error rate is 15–25%, we will expand the sample to 100. For benchmarks with ≤50 available tasks, all available tasks will be assessed.
3. Evaluation Quality
This is the standard we would like all benchmarks to meet, but it is not necessarily disqualifying to omit or fail these items.
| Question | Status | Notes |
|---|---|---|
| Elicitation and resource adequacy: Are the resources given to models (reasoning token/turn budget, tool access, etc.) sufficient for them to perform near their ceiling? |
Sufficient
Constraining
Unreasonably constraining
Unknown
Not reviewed | — |
| Scaffold fairness: What scaffold does the leaderboard report? |
Shared common scaffold
Mix of model-specific and common scaffolds
Model-specific scaffolds
Not reviewed | — |
| Is there evidence/risk of contamination? |
Not reviewed | — |
| Has human completability been assessed? |
All tasks
Representative set of tasks
Poor implementation (Unrepresentative set of tasks, unreasonable set of participants)
Not established
Not reviewed | — |
| Score range (if possible to estimate) | — floor, — ceiling
Not reviewed | — |
| Statistical adequacy: how many runs/model (≥ 5 recommended for error bars) | — runs/model
Unknown
Not reviewed | — |
| Construct Validity |
Measures stated capabilities
Partially measures stated capabilities
Does not measure stated capabilities
Not reviewed | — |
Disclaimer
This is a partial review. We stopped inspecting once we discovered errors that exceeded our published threshold for a flawed benchmark.
We inspected 50 tasks, sampled at random, stratified by category. Tasks not inspected may contain additional defects.
We attempted to reach out to the benchmark developer prior to launch, and if they write a response to this review we will link it here. If you’re the benchmark developer, please email us at reviews@epoch.ai with a link to your response if you would like us to include it here.
-
A task can carry more than one flag, therefore categories overlap and sum isn’t 50.
-
The authors refer to this categorization as “by source slice(s)”, in Section 3.5.
-
Good faith questions are real prompts clinicians asked models during the course of their practice. These are further split based on the expert’s perceived difficulty.
-
Red teaming questions are adversarial prompts designed to probe weaknesses.
-
191 out of 525 (36.4%) tasks have negatively weighted criteria: a model will get penalized if it includes information a physician might consider potentially wrong or dangerous.
About Benchmark Reviews
Epoch AI’s Benchmark Reviews are independent reviews of external AI benchmarks. Our documentation describes the rubric behind this verdict, how we choose which benchmarks to review, and answers frequently asked questions.