Benchmark Review

TextQuests

Reviewed: Aug. 10, 2026

Benchmark creator: Center for AI Safety (Long Phan, Mantas Mazeika, Andy Zou, Dan Hendrycks)

Verdict: Flawed

TextQuests is a long-horizon, tool-free benchmark where agents play through the Infocom suite of interactive fiction games. The leaderboard’s primary metric is “game progress”. However, this metric is flawed because it does not correctly account for the nonlinear nature of these games. Models that have achieved equal actual progress in a given game can receive different “game progress” scores, and models with the same “game progress” scores can differ in how many necessary objectives they have completed. Since the leaderboard may not consistently reflect overall progress toward completing the game, we give the benchmark a flawed verdict.

Methodology

The leaderboard reports “game progress” (a constructed metric we discuss in more detail below), number of completed games, and a “harm score”. These three metrics are reported for both runs with no clues and runs that use the game’s InvisiClues. The tasks, scoring code, and per-model harness settings are all public. The benchmark’s repo contains a walkthrough for each game that represents a possible path to completing the game. The provided walkthroughs range from 82 to 604 steps (mean 270, median 254). We used these walkthroughs to evaluate the game progress metric. We did not assess the harm score because it is not the primary metric reported, although it has numerous issues.1 The benchmark includes checkpoints throughout the game used to measure progress. Using Claude Fable 5.1, we audited each of the 25 games and the 327 checkpoints by replaying the walkthrough and looking for checkpoints that were impossible or false positives/negatives. Then we validated each error through human review and localized step-by-step walkthroughs.

Each checkpoint consists of a name, a long-form description of what it assesses, the output string it fires on, and a completion percentage. This percentage represents the benchmark creators’ estimate of how close the agent is to finishing the game, based on its completion of necessary objectives, such as puzzles and milestones. When a checkpoint is triggered, the hardcoded percentage for that checkpoint (listed in game_progress.json) is compared to the current game progress, and the higher of the two is stored as the new game progress. Because these games are nonlinear, checkpoints may be reached in a variety of orders, so the current game progress is sometimes higher than the checkpoint percentage. The game progress is therefore the percentage associated with the highest checkpoint reached, and the total number of checkpoints achieved may vary.

The game progress metric does not accurately represent progress through the game. Below are two ways that inaccuracy occurs:

  • Two playthroughs can have similar actual progress, but different “game progress” scores: Agents can reach checkpoints in different orders, and the arbitrary order they happen to choose affects the game progress metric.2 For example, in Hollywood Hijinx at step 77 in the walkthrough, the player is standing in the foyer. In the walkthrough, the player explores the downstairs wing first in steps 78-107, then returns to the foyer, then explores the upstairs wing in steps 108-130, then ends in the foyer. Both wings contain objectives necessary to complete the game. We compare the game progress of the given walkthrough (“downstairs wing first”) with the walkthrough where the order of steps 78-107 and steps 108-130 are switched (“upstairs wing first”).

    • Downstairs wing first (30 moves, 2 checkpoints): checkpoints 30%, 40% → game progress 40%
    • Upstairs wing first (23 moves, 1 checkpoint): checkpoint 50% → game progress 50%

    Thus, if a model completes one wing and then either gets stuck or hits the step limit, the arbitrary choice of which wing to explore first affects the game progress score. This type of grading defect means that models with different game progress scores may have achieved similar actual progress throughout the game, based on the arbitrary order they went through the nonlinear games. We are unsure how many situations like this exist in the benchmark, but since it is an unavoidable result of the scoring methodology, we expect it to be pervasive.

  • Two playthroughs can have different actual progress, but the same “game progress” scores: The same game progress percentage can also represent different states of game completion. For example, in the provided enchanter walkthrough, progress jumps from 25% to 45% in step 98, then from 45% to 70% in step 115; the checkpoints for 35%, 40%, 55%, and 60% all trigger later without a corresponding change in game progress (we refer to these as inert). A model that made it to 70% using this solution path (without completing the objectives associated with the 55% and 60% checkpoints) will receive an identical game progress score as a model that achieved all of these objectives. However, these two solution paths represent different levels of actual progress in the game: the 55% and 60% objectives entail befriending the Adventurer and opening the Guarded Door, which are requirements for unlocking later puzzles and actually completing the game. Using the walkthroughs, we identified 35 of these inert checkpoints across 17 games.3 This grading defect means that models with the same game progress score may have reached different numbers of checkpoints and achieved different levels of actual game progress, depending on the arbitrary order they went through the nonlinear game.

An additional issue is that the game progress metric uses string matching rather than relying on the actual state. This results in many errors, particularly those related to object retrieval, which is a common milestone. Infocom’s parser prints different text depending on how many objects are taken. Any input of the form “get X” returns “Taken”, so game progress is updated at a prior step that has a unique response. Thus, in gameplay, a checkpoint meant to assess whether an object was acquired might trigger without the object ever being acquired. The majority of these checkpoints trigger on the action immediately preceding the item pickup,4 and it is reasonable to assume that an agent’s next action would be to acquire the object. However, sometimes it takes multiple steps to acquire the object, and we mark these checkpoints as false positives. Other checkpoints are tied to situations where the agent picks up two or more objects at once. Any input of the form “get X and Y” returns “X: Taken. / Y: Taken.” These are often underspecified and prone to false positives and false negatives. For example, the Zork2 5% checkpoint is supposed to assess whether the agent acquired both the lamp and sword, but it only verifies the string “lamp: Taken”. The examples below show how the checkpoint responds to three different sequences of agent commands:

  • “get sword” → “Taken”.
    Then “get lamp” → “Taken”. (progress 0%) [correct answer marked incorrect: the agent acquired both items, but the output never shows “lamp: Taken”.]
  • “get lamp and tunnel” → “lamp: Taken. / tunnel: You must specify a direction to go”. (progress 5%) [incorrect answer marked correct: the agent acquired only the lamp, but the output shows “lamp: Taken”.]
  • “get lamp and sword” OR “get all” → “lamp: Taken. / elvish sword: Taken”. (progress 5%) [correct answer marked correct: the agent acquired both items, and the output shows “lamp: Taken”.]

Representative Errors

Additional task-based errors fall into the following categories:

Impossible

  • Each game is capped at 500 steps. The game Trinity’s walkthrough is 604 steps; if the agent is limited to 500 steps, then following the walkthrough can only achieve a score of 71%. While it is theoretically possible to delete steps from the walkthrough and achieve a perfect score in <500 steps, our shortened walkthrough removes information-gathering steps, making this solution infeasible.5 It remains unvalidated whether a shorter solution that does not rely on preexisting game knowledge or luck is possible.
  • At the start of each game, the agent is able to open feelies_text.txt6 and, if allowed clues, invisiclues.txt. Sometimes games require additional files which can be found in the scans folder and manual.pdf in the repo. However, these are not provided to the agent, leading to impossible tasks. For example:
    • In Cutthroats, coordinates exist only on wrecks-map.jpeg, which the agent cannot access.
    • Sherlock is missing the feelies_text.txt file completely, so the agent cannot access information needed for certain tasks.

Incorrect answers marked correct

  • Trinity’s 100% checkpoint matches “A shower of sparks erupts from the enclosure,” which is printed for cutting both the correct and the incorrect wire.
  • In Starcross, placing the red rod into slot 1, 2, or 3 results in the 60% checkpoint due to ambiguous output string matching. However, the checkpoint is meant to check that the rod was placed in slot 2 (the oxygen slot), not the methane or ammonia slot, which are fatal.

Correct answers marked wrong

  • Sherlock’s 96% checkpoint, titled “Crown Jewels Recovered – Moriarty Subdued”, wrongly fails when the agent enters “get jewels” but succeeds when it enters “get all”.

Checkpoint errors by game

GameCheckpoints (errors / total)Errors by type (impossible - correct marked incorrect - incorrect marked correct - inert)7Corresponding game progress checkpoints (labeled by % associated with the checkpoint)
ballyhoo5 / 200 - 1 - 0 - 5correct marked wrong: 20%; inert: 20%, 45%, 50%, 55%, 70%
borderzone2 / 140 - 0 - 2 - 0incorrect marked correct: 35%, 85%
cutthroats5 / 113 - 1 - 2 - 1impossible: 80%8, 90%, 100%; correct marked wrong: 90%; incorrect marked correct: 30%, 90%; inert: 60%
deadline3 / 130 - 0 - 2 - 1incorrect marked correct: 78%, 92%; inert: 23%
enchanter5 / 150 - 1 - 2 - 4correct marked wrong: 60%; incorrect marked correct: 60%, 95%; inert: 35%, 40%, 55%, 60%
hitchhiker1 / 130 - 1 - 0 - 0correct marked wrong: 80%
hollywoodhijinx1 / 120 - 0 - 1 - 0incorrect marked correct: 22%
infidel1 / 101 - 0 - 0 - 0impossible: 20% (soft)9
lurkinghorror8 / 193 - 1 - 2 - 6impossible: 5%, 60%, 100%; correct marked wrong: 70%; incorrect marked correct: 70%, 75%; inert: 5%, 35%, 40%, 50%, 70%, 75%
moonmist5 / 103 - 0 - 0 - 3impossible: 35%, 45%, 90%10; inert: 45%, 65%, 75%
planetfall1 / 130 - 0 - 0 - 1inert: 20%
plunderedhearts2 / 121 - 0 - 0 - 1impossible: 40% (soft)11; inert: 16%
seastalker2 / 130 - 0 - 2 - 0incorrect marked correct: 15%, 50%
sherlock8 / 127 - 1 - 0 - 1impossible: 55%, 65%, 70%, 80%, 90%, 96%, 100%; correct marked wrong: 96%; inert: 45%
sorcerer2 / 170 - 0 - 0 - 2inert: 45%, 50%
spellbreaker2 / 110 - 1 - 1 - 1correct marked wrong: 45%; incorrect marked correct: 95%; inert: 45%
starcross3 / 160 - 0 - 2 - 1incorrect marked correct: 60%, 75%; inert: 70%
stationfall1 / 120 - 0 - 0 - 1inert: 40%
suspect3 / 100 - 2 - 1 - 1correct marked wrong: 35% (live-path), 60% (live-path)12; incorrect marked correct: 85%; inert: 60%
trinity6 / 134 - 0 - 3 - 0impossible: 82%, 92%, 95%, 100%; incorrect marked correct: 42%, 71%, 100%
The game walkthrough requires 604 steps, but play is capped at 500.
wishbringer1 / 130 - 0 - 0 - 1inert: 30%
witness0 / 110 - 0 - 0 - 0
zork13 / 130 - 1 - 0 - 3correct marked wrong: 30%; inert: 30%, 40%, 50%
zork23 / 140 - 2 - 2 - 2correct marked wrong: 5%, 30%; incorrect marked correct: 5%, 40%; inert: 30%, 40%
zork31 / 100 - 0 - 1 - 0incorrect marked correct: 55%
TOTAL74 / 327 (22.6%)22 - 12 - 23 - 3592 findings on 74 distinct checkpoints

Rubric

1. Reviewability

LevelMeaningStatus
FullAll tasks and scoring logic inspectable*, and harness/API settings used for each model (reasoning effort, token/time limits, tool access, system prompts) are fully disclosed [proceed to 2]
PartialA representative sample of tasks and scoring logic inspectable, with full harness/API settings disclosed [proceed to 2]
InadequateLimited, biased, or no inspectable tasks or scoring logic, harness/API settings undisclosed [stop → NEI]

* Either publicly or privately to reviewers.

2. Scoring

This is the minimum standard required to be Verified; failing any of these items results in a Flawed verdict.

Scoring
Examples
  • Essentially impossible to correctly answer as written (e.g. underspecified task, hidden requirement, missing file/tool)
  • False Negative (e.g. overly strict scorer, stale/incorrect ground truth, dependent on live external state that can drift, sandbox failure independent of the agent)
  • False Positive (e.g. lax scorer, reward-hackable environment, stated skill can be bypassed via a shortcut such as exploiting an error in the scoring logic or retrieving the answer from the harness/web)
  • Task egregiously doesn’t measure the claimed capability
Default threshold for Flawed≥20% of inspected sample† contains errors or there is an issue that corrupts grading at scale
Status
Pass Flag [stop → Flawed] Not reviewed
Notes—
Benchmark Consistency
Examples
  • Scorer, instructions, or ground truth changed without a version bump
Default threshold for FlawedLeaderboard has incomparable results from different versions
Status
Pass Flag [stop → Flawed] Not reviewed
Notes—
Elicitation
Examples

Model elicitation is extremely constraining and is not the focus of the benchmark:

  • Under-resourced relative to task size (token/turn/time limit, sandbox resources)
  • Poor context management
  • Lack of agentic environment where it would be natural to provide one
  • Excessive non-voluntary termination for agentic benchmarks
Default threshold for FlawedSubstantially reduced performance compared to reasonable alternatives for the tasks
Status
Pass Flag [stop → Flawed] Not reviewed
NotesIn 1 of the 25 games, the 500 step limit is below the 604 step reference solution (Trinity). All other reference solutions are <500 steps.
Bias in Evaluation Setup
Examples
  • Uneven compute/token budgets
  • Unfair scaffold choice (e.g. only a subset of models optimized)
Default threshold for FlawedMaterial model-specific advantage found
Status
Pass Flag [stop → Flawed] Not reviewed
Notes—

† For benchmarks with >50 available tasks, we will sample a random set of 50 (stratified by category, when present). If the observed error rate is 15–25%, we will expand the sample to 100. For benchmarks with ≤50 available tasks, all available tasks will be assessed.

3. Evaluation Quality

This is the standard we would like all benchmarks to meet, but it is not necessarily disqualifying to omit or fail these items.

QuestionStatusNotes
Elicitation and resource adequacy: Are the resources given to models (reasoning token/turn budget, tool access, etc.) sufficient for them to perform near their ceiling?
Sufficient Constraining Unreasonably constraining Unknown Not reviewed
500 vs. 800 ablation step results in paper show performance is still improving. Budget insufficient on 1 game (Trinity). Best model: 2/25 completions without clues, 13/25 with clues.
Scaffold fairness: What scaffold does the leaderboard report?
Shared common scaffold Mix of model-specific and common scaffolds Model-specific scaffolds Not reviewed
—
Is there evidence/risk of contamination?
As of Aug. 10, 2026:
100% of tasks public
100% of solutions public
Not reviewed
—
Has human completability been assessed?
All tasks Representative set of tasks Poor implementation (Unrepresentative set of tasks, unreasonable set of participants) Not established Not reviewed
—
Score range (if possible to estimate)
— floor, — ceiling Not reviewed
—
Statistical adequacy: how many runs/model (≥ 5 recommended for error bars)
1 runs/model Unknown Not reviewed
—
Construct Validity
Measures stated capabilities Partially measures stated capabilities Does not measure stated capabilities Not reviewed
—

Disclaimer

This is a partial review. We stopped inspecting once we discovered errors that exceeded our published threshold for a flawed benchmark.

We attempted to reach out to the benchmark developer prior to launch, and if they write a response to this review, we will link it here. If you’re the benchmark developer, please email us at reviews@epoch.ai with a link to your response if you would like us to include it here.

Notes
  1. Harm is additive over the steps taken so it is correlated with progress. A model that does poorly does not have as much opportunity for harm. To be interpretable, the harm score should be normalized. The harm score comes from the underlying Jiminy Cricket repo and TextQuests only reports the “bad”/”others” category, not the “bad”/”self”, “good”/”others”, or “good”/”self” categories. Return

  2. We believe it would be an improvement to give each checkpoint a number of percentage points to add to the game progress, rather than assigning it the max percentage associated with any checkpoint achieved. Return

  3. Due to the nature of the nonlinear games, we expect investigating more walkthroughs would identify additional checkpoints that can be inert in the corresponding paths. Return

  4. ballyhoo 35/90 · hitchhiker 70 · hollywoodhijinx 12/30/40/60/68/78/86/93 · infidel 45 · starcross 30 · stationfall 30/40/70 · trinity 82 · wishbringer 50/60 · zork1 60 · zork2 58/75 · zork3 20
    Note: These checkpoints are not included in the table below (unless other issues are present). Return

  5. For example, “examine coin” reveals “It’s standard British currency, worth fifty pence” which is not necessary information, but “examine sundial” reveals “inscribed with seven curious symbols and a compass rose”, which is needed for the following step, “turn ring to sixth symbol”. Without hindsight, it’s not possible to remove only the examine steps that do not reveal vital gameplay information. We have not yet found a solution under 500 steps that is possible without invisiclues.txt. Return

  6. Feelies represent the physical items that came in the box for each of these games. These are vital to play the game correctly, and originally made it so pirated versions were unwinnable. Text ends up in feelies_text.txt, but scans of images are not seen, leading to impossible tasks. Return

  7. Inert: Triggering this checkpoint in the provided walkthrough does not change the game progress since a higher checkpoint was already achieved. Return

  8. No access to wrecks-map.jpeg, which is needed for checkpoints 80/90/100. Return

  9. The pyramid’s location is only marked on a map in the game package (which is not included), but brute force can get to the location. The invisiclues also give the exact route to find it. Return

  10. If a color other than green is chosen, checkpoints 35/45/90 cannot be triggered. Return

  11. The banknote is not accessible, but the pose can be reconstructed in game via brute force. Invisiclues:
    “C. Jean Lafond is depicted in the portrait in the Library in exactly the same pose as he is on the banknote. Examine the banknote that came in your game package for clues.
    D. Notice what Lafond is carrying.
    E. Notice what Lafond is wearing.
    F. Notice what Lafond is doing. This must be done last. Voila!” Return

  12. This defect shows up in the live scorer, not the leaderboard (which recomputes based on cleaned text). This results in game_finished() not working as expected, wasting budget and potentially increasing the published harm score. Return

About Benchmark Reviews

Epoch AI’s Benchmark Reviews are independent reviews of external AI benchmarks. Our documentation describes the rubric behind this verdict, how we choose which benchmarks to review, and answers frequently asked questions.

Go to sourceRead the documentation and FAQs