SimpleQA Verified successfully measures whether a model can recall obscure facts with no tool usage. The 1,000 questions were filtered from OpenAI’s earlier work, SimpleQA, to be unambiguous and difficult for 2025-era models. Each question has an attached human-sourced answer and at least two supporting URLs.
The benchmark claims to measure parametric factual knowledge,1 but the headline number also reflects a model’s willingness to guess, as abstentions are marked wrong. All questions and answers have been public since 2024 as part of SimpleQA, so the benchmark is not contamination-resistant, and Epoch has already flagged one likely contaminated model. Frontier models answer about three quarters of the questions correctly as of the review date.
In a random, topic-stratified sample of 50 questions, we found 5 (10%) with score-affecting defects. Of those five, there are four questions where the answer key does not cover all of the technically correct answers, and one where the answer key is outright incorrect. This error rate is well below our 20% threshold for declaring a benchmark flawed, but still worth being aware of.
Interpretation
SimpleQA Verified tests recall for one obscure fact per question, answered in a single turn with no tools, and graded by an LLM scorer against a human-sourced answer key as CORRECT, INCORRECT, or NOT_ATTEMPTED. Epoch’s hub reports the percentage of questions answered correctly, with NOT_ATTEMPTED counted as incorrect. Epoch uses GPT-5.6 Terra as the grader with the official grading prompt, and appends an instruction that guessing carries no penalty, developed and tested in collaboration with Generality Labs. A high scorer would be able to retrieve long-tail facts about people, places, dates and numbers across ten topic areas without retrieval or additional tools. Relatedly, the dataset card claims the benchmark is trivial with tools.
Potential confounds
- Abstention handling: under the percent correct metric, “I don’t know” scores the same as a wrong answer, so willingness to guess moves scores by several points. In Epoch’s August 2026 runs on the bare prompt, Claude Fable 5 (xhigh) declined 11% of questions and scored 69.2% correct (77.8% of attempts), while Gemini 3.1 Pro (high) declined 2% and scored 75.2% correct (76.7% of attempts).
- Convention-dependent answers: several questions’ resolutions lean on an unclear convention (e.g. what stage of a project’s beginning counts as “set up”). A model that knows the facts but follows the “wrong” convention is graded INCORRECT. We estimate this caps the score of a well-informed model near 90% (see Task Analysis).
- LLM grading: the grader is itself a model and we did not measure its error rate. OpenAI’s original SimpleQA autograder produced 15 false negatives in 1,000 human answers.2 We use the revised grading prompt from SimpleQA Verified along with a stronger grader model (GPT-5.6 Terra compared to o3-mini), so we expect a similar or lower false-negative rate.
Task Analysis
All 1,000 questions and answers, along with topic, answer-type labels, and source URLs are public on Hugging Face. The questions demand recall of obscure facts, including the third ionization energy of molybdenum, the name of the first EP by Kinoko Teikoku, and who replaced Bradie Tennell at the 2021 Skate America. We inspected the full dataset and audited a random sample of 50 questions against independently searched sources.
The questions are a filtered subset of the 4,326 in OpenAI’s SimpleQA, which were written and answered by human trainers and independently re-answered by a second trainer. Google’s pipeline removed questions that share a source URL (3,095 remaining), near-duplicates by embedding and TF-IDF similarity (2,664), and questions whose sources disallow crawling (1,855), rebalanced topics and answer types (1,218), dropped questions whose sources conflicted (with a 5% tolerance for numbers) (1,073), and selected for difficulty by preferring questions that GPT-4o, Gemini 2.0 Flash and Claude 3.7 Sonnet all answered incorrectly (1,000). Numeric answers carry explicit acceptable ranges: exact for small counts, about 1% for mid-sized quantities, and 5% for large quantities.
- Topics: Politics 176, Science and technology 160, Art 145, Sports 117, Geography 111, Music 102, Other 102, History 52, TV shows 20, Video games 15
- Answer types: Other 249, Date 222, Person 198, Number 185, Place 146
Question Audit
Following our standard protocol, we drew a random sample of 50 questions,3 stratified in proportion to topic counts: Politics 9, Science and technology 8, Art 7, Sports 6, Geography 5, Music 5, Other 5, History 3, TV shows 1, Video games 1. For each question, Fable 5.1 was asked to do a web search to establish the ground truth independently from the benchmark-provided source. 10 Claude Fable 5.1 research agents ran at least 3 web searches and read at least 2 authoritative pages per question. We manually checked the agents’ rationale, with special scrutiny applied to claims that the answer key was wrong. Each question was classified as OK, Impossible (no determinate answer as written), False negative (a reasonable, correctly formatted correct answer could be graded INCORRECT) or False positive (an incorrect answer could be graded CORRECT).
We found 5 of the 50 questions (10%) to be defective; 4 of them are false negatives (questions in which a correct answer could be marked wrong) and 1 of them is likely impossible.
| Question (original_index) | Topic | Error type | Notes |
|---|---|---|---|
| 2928. In what year did Danish mathematician and astronomer Niels Erik Norlund set up seismographic stations in Denmark and Greenland? Answer key: 1925 | Science and technology | False negative | 1925, the reference answer, is when funding was secured and instruments ordered. The Copenhagen station was installed in late 1926 and began monitoring in February 1927.4 1926 and 1927 are graded INCORRECT, which is a handicap for any model that might use the installation date or start-of-monitoring date as the year that those stations were “set up.” |
| 4159. In which month and year, during a visit to Washington, D.C., did Golda Meir agree with Henry Kissinger’s peace proposal based on “security versus sovereignty”? Answer key: February 1973 | Politics | False negative | The visit ran from February 26 to March 8, 1973. The reference answer is based on Rabin’s memoir of a private meeting on 28 February, but the official State Department record has Meir saying, “We want security; we are not concerned with sovereignty,” and accepting the proposal on March 1, 1973, which is also the date scholarly accounts use.5 Unfortunately, “March 1973” is graded INCORRECT. |
| 659. Other than being a choreographer for the TV series In Living Color, what other job did Rosie Perez do on the show? Answer key: segment producer | TV shows | Impossible, false negative | The reference answer can be traced back to one line in a 1999 interview with Roger Ebert. However, Perez had more roles than just “segment producer” - formal credits list her both as director of the dance segments and as talent coordinator.6 In this benchmark, “segment director” and “talent coordinator” are spuriously graded INCORRECT. |
| 765. What month, day, and year was the Köln Frankfurter Straße station opened? Answer key: 13 June 2004 | Geography | False negative | The station sits on an airport loop that was ceremonially opened on June 12, 2004, and the German Wikipedia article on the station gives that date as its opening. June 13, 2004 is simply the date on which scheduled service began.7 However, “12 June 2004” is graded INCORRECT, which is an improper grading if a ceremonial opening counts as an opening. |
| 1149. Which architect designed the old post office at the corner of Jacques-Cartier and Saint-Jacques streets in Saint-Jean-sur-Richelieu, which was completed in 1909? Answer key: J. E. H. Benoît | Art | False negative | Quebec’s heritage registry and the Biographical Dictionary of Architects in Canada both name the correct architect as Joseph-E.-Alexandre (J. E. A.) Benoît.8 However, the answer key mistakenly contains the answer “J. E. H. Benoît,” and one of the sources has no citations (as opposed to J. E. A. Benoît’s ownership of this project, which is well-backed by other sources). |
The answer keys of the other 45 sampled questions were found to be correct, with original_index values as follows: 8, 79, 141, 356, 386, 394, 477, 502, 815, 1132, 1367, 1513, 1542, 1742, 1744, 1847, 1851, 1952, 1976, 2010, 2079 2146, 2277, 2374, 2447, 2622, 2943, 2945, 3021, 3030, 3162, 3171, 3271, 3304, 3306, 3391, 3443, 3495, 3503, 3531, 3547, 3626, 3737, 4188 and 4283.
Elicitation & Scaffolding
- Elicitation setup: one turn generation, with no system prompt beyond the question and no tools. Epoch runs each model once at a high or maximum reasoning setting, grades with GPT-5.6 Terra (medium reasoning) using the same prompt and structured tool-call output, and reports percentage correct. Since 2026-08-27, Epoch appends an anti-abstention instruction, chosen from 9 candidate wordings tested on 12 models.9
- Scaffold variants: none. All models see the same prompt and grader. The only variant is Epoch’s anti-abstention instruction.
Limitations
All 1,000 questions and answers have been public since SimpleQA’s release in October 2024, and there is no held-out split, so it’s difficult to distinguish whether the model has memorized the answer key or not. Epoch already suspects contamination in one case: Qwen3-Max-Instruct scored 67%, the top score at the time, while ranking much lower on AA-Omniscience, a closed factuality benchmark of similar design.
The final selection step during benchmark construction only kept questions that GPT-4o, Gemini 2.0 Flash, and Claude 3.7 Sonnet all answered incorrectly. This skews the benchmark towards facts that weren’t present in those three models’ training corpora, likely artificially depressing their (and similar models’) scores, though we have not quantified the effects of this problem. In any case, we don’t consider this a major issue today, as those models are not near the frontier.
Epoch’s scores use accuracy as a metric, whereas Google’s paper reports F1 on the bare prompt10 (the former punishes abstention, while the latter rewards it). Others who run the benchmark, such as model developers, tend to report accuracy and don’t use the new anti-abstention instruction.
One question was reworded on 2026-03-09 without a version bump. The effect of that change is up to 0.1 points, but there are some public mirrors of the dataset which now disagree on the wording of that question.
Rubric
1. Reviewability
| Level | Meaning | Status |
|---|---|---|
| Full | All tasks and scoring logic inspectable*, and harness/API settings used for each model (reasoning effort, token/time limits, tool access, system prompts) are fully disclosed | [proceed to 2] |
| Partial | A representative sample of tasks and scoring logic inspectable, with full harness/API settings disclosed | [proceed to 2] |
| Inadequate | Limited, biased, or no inspectable tasks or scoring logic, harness/API settings undisclosed | [stop → NEI] |
* Either publicly or privately to reviewers.
2. Scoring
This is the minimum standard required to be Verified; failing any of these items results in a Flawed verdict.
| Scoring | |
|---|---|
| Examples |
|
| Default threshold for Flawed | ≥20% of inspected sample† contains errors or there is an issue that corrupts grading at scale |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | 5 of 50 sampled questions (10%) defective, all false negatives (four ambiguous conventions, one wrong name); no false positives found. |
| Benchmark Consistency | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Leaderboard has incomparable results from different versions |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | One question reworded on 2026-03-09 without a version bump, only a 0.1-point effect. |
| Elicitation | |
|---|---|
| Examples | Model elicitation is extremely constraining and is not the focus of the benchmark:
|
| Default threshold for Flawed | Substantially reduced performance compared to reasonable alternatives for the tasks |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | Single-turn, tool-free by design. Abstention handling (metric or instruction) is the only material choice. |
| Bias in Evaluation Setup | |
|---|---|
| Examples |
|
| Default threshold for Flawed | Material model-specific advantage found |
| Status |
Pass
Flag [stop → Flawed]
Not reviewed |
| Notes | Same prompt and grader for every model. |
† For benchmarks with >50 available tasks, we will sample a random set of 50 (stratified by category, when present). If the observed error rate is 15–25%, we will expand the sample to 100. For benchmarks with ≤50 available tasks, all available tasks will be assessed.
3. Evaluation Quality
This is the standard we would like all benchmarks to meet, but it is not necessarily disqualifying to omit or fail these items.
| Question | Status | Notes |
|---|---|---|
| Elicitation and resource adequacy: Are the resources given to models (reasoning token/turn budget, tool access, etc.) sufficient for them to perform near their ceiling? |
Sufficient
Constraining
Unreasonably constraining
Unknown
Not reviewed | It’s designed as a single-turn no-tools benchmark, so it’s constrained but not unreasonably so. |
| Scaffold fairness: What scaffold does the leaderboard report? |
Shared common scaffold
Mix of model-specific and common scaffolds
Model-specific scaffolds
Not reviewed | See “Elicitation & Scaffolding”. |
| Is there evidence/risk of contamination? | As of Sep. 10, 2026: 100% of tasks public, 100% of solutions public
Not reviewed | One suspected case (Qwen3-Max-Instruct); see “Limitations”. |
| Has human completability been assessed? |
All tasks
Representative set of tasks
Poor implementation (Unrepresentative set of tasks, unreasonable set of participants)
Not established
Not reviewed | The answer keys were originally human-verified, but there has not been a human test-taker baseline. |
| Score range (if possible to estimate) | 0% floor, 90-100% ceiling
Not reviewed | — |
| Statistical adequacy: how many runs/model (≥ 5 recommended for error bars) | 1 runs/model
Unknown
Not reviewed | For reference, 1,000 questions give a standard error of about 1.4 points at 75%. This seems sufficient. |
| Construct Validity |
Measures stated capabilities
Partially measures stated capabilities
Does not measure stated capabilities
Not reviewed | See “Interpretation”. |
Disclaimer
We attempted to reach out to the benchmark developer prior to launch, and if they write a response to this review we will link it here. If you’re the benchmark developer, please email us at reviews@epoch.ai with a link to your response if you would like us to include it here.
-
The SimpleQA Verified paper describes its target as “a model’s ability to recall facts directly from its internal parameters” and calls this domain “parametric factuality,” which “isolates the model’s stored knowledge from external aids” (Haas et al. 2025, Section 1). Essentially, it seeks to measure the amount of knowledge a model can produce from its trained weights alone, without tools.
-
From the original SimpleQA paper: “Of the 5.6% of answers (56/1,000) classified as incorrect by the prompted ChatGPT grader, we did a manual examination of all examples and found that fifteen were false negatives from the autograder.”
-
The sample was drawn from the Hugging Face dataset google/simpleqa-verified. Per-topic quotas are 50 times each topic’s share of the 1,000 questions. Topics are processed in descending order of question count, and within each topic the quota is drawn without replacement from the topic’s original_index values using a numpy.random.default_rng(20260910) generator.
-
Hjelme, Voss and Larsen (2014), GEUS Bulletin: “Establishing seismological stations from 1925 to 1927”; “By 17 February 1927 all instruments were in place and monitoring began”; https://geusbulletin.org/index.php/geusb/article/download/4424/10145/29149. The reference answer follows MacTutor: “in 1925, set up seismographic stations in Denmark and Greenland.”
-
Foreign Relations of the United States, 1969–1976, vol. XXV, doc. 35, Memorandum of Conversation, 1 March 1973: https://history.state.gov/historicaldocuments/frus1969-76v25/d35. The 28 February date is Rabin’s (The Rabin Memoirs, 1996, p. 215) via Wikipedia.
-
TV Guide credits list her under Director, Choreographer and Talent Coordinator: https://www.tvguide.com/celebrities/rosie-perez/credits/3000557410/. IMDb lists segment director (Dance Bumpers, 59 episodes) and talent coordinator (26 episodes). The reference answer follows Ebert, “Rosie Perez On A Roll,” 17 February 1999.
-
https://de.wikipedia.org/wiki/Bahnhof_K%C3%B6ln_Frankfurter_Stra%C3%9Fe (“wurde am 12. Juni 2004 … in Betrieb genommen”); https://en.wikipedia.org/wiki/K%C3%B6ln_Frankfurter_Stra%C3%9Fe_station (13 June 2004, citing the NRW rail archive).
-
Quebec heritage registry person record: https://www.patrimoine-culturel.gouv.qc.ca/rpcq/detail.do?id=22796&methode=consulter&type=pge; Hill, Biographical Dictionary of Architects in Canada: http://dictionaryofarchitectsincanada.org/node/1099 (“ST. JEAN, QUE. Post Office, Cartier Street, 1906”). The reference answer’s initials appear in the registry’s building record and two local history pages.
-
Without any such instructions, the abstention rate was up to 82% depending on the model. The best wordings reduced this rate to under 1% for 10 of 12 models. Of all the instructions, we picked the one with the lowest average abstention rates.
-
The F1 score is in the range [0,1], the harmonic mean of the share of all questions answered correctly and the share correct among questions attempted. Models that answer when they know and abstain when they don’t are thus rewarded. The F1 score is distinct from the percentage of problems answered correctly.
About Benchmark Reviews
Epoch AI’s Benchmark Reviews are independent reviews of external AI benchmarks. Our documentation describes the rubric behind this verdict, how we choose which benchmarks to review, and answers frequently asked questions.