Back to benchmarks

Earthborne Rangers (EBR-bench)

A test of AI systems' learning capabilities, measuring whether their scores improve across repeated playthroughs of the relatively obscure game Earthborne Rangers.

Earthborne Rangers (EBR-bench)

Overview

EBR-bench evaluates whether AI systems can improve at an unfamiliar, complex task through repeated attempts and note-taking — a proxy for the broader ability to “learn on the job.” Models repeatedly play Earthborne Rangers, a campaign-style board game in which a player explores a wilderness landscape, overcoming obstacles and pursuing objectives.

We chose this game because it is relatively obscure, so models are unlikely to have memorized it during training; it is almost entirely card-based with very little spatial reasoning, so weak multimodal reasoning is not the bottleneck; and it requires a layered mix of strategy and tactics — deck-building plus turn-by-turn play — over a long time horizon. A single playthrough of the segment we evaluate takes an experienced human 2–4 hours, and mastering the game can require dozens of playthroughs.

In our early experiments, we saw little evidence of learning from experience: scores stay roughly flat over repeated play, well below an expert human baseline. For the full launch analysis, see our publication, AI doesn’t get better at this board game with practice. As of the release of GPT-6 Astra, we now observe above-median-human average performance, and occasionally top-human performance, from frontier models.

Epoch has used Earthborne Rangers content solely for the purposes of research and commentary. This work is conducted independently of Earthborne Games. Epoch does not distribute or sell access to the underlying copyrighted material from the game.

Experimental settings and chart descriptions

We report the “topline score” for each model on default settings. Default settings consist of ten playthroughs in a sample, with each sample consisting of eight learning playthroughs and two scored playthroughs; the topline score for each model is the better of their two scored playthroughs. All models are run at a compaction setting of 250k, with the exception of Fable-series models that are run at 90% of their advertised context window.

As of September 2026, our EBR-specific learning experiments (learning trajectories across playthroughs, learning gain relative to ablated runs, the effect of a strategy guide, and deck diversity) are now analyzed internally rather than published as interactive charts.

Note that chart data represent single-agent runs, and that older models besides GPT-5.6 Sol, Fable 5.1, and Opus 5 were not rerun on the “Card ban” setting. All recent models are evaluated under the “Card ban” setting, and while their scores may not be directly comparable to older models’ results, we are reasonably confident older models’ scores would not be affected by the ban (see “Methodology updates” for more details). The four charts we publish are:

  • Score vs time plots each model’s score (as a percentage of 21 possible objectives) against its release date, with the top and average human baseline participants’ scores marked for reference. “Frontier trend only” restricts the view to models that were state of the art on this benchmark when released.
  • Score vs training compute plots score against the compute used to train each model. Coverage is limited to models with a published training compute estimate.
  • Score vs ECI plots score against the Epoch Capabilities Index, together with the score the index would predict. Models above the trend line perform better at Earthborne Rangers than their general capability level implies; models below it perform worse.
  • Leaderboard ranks every evaluated model by score, with 90% confidence intervals where a model was run multiple times.

Methodology

We primarily report each model’s “topline score”, their top score over the final 2 playthroughs of a run. Each model’s topline score is sampled 5 times (10 for older runs). Agents are informed in their system prompt that only their final 2 playthroughs are scored, and that the other playthroughs are free to be used for exploration and learning. By default, a sample or “run” consists of 10 playthroughs. Agents are encouraged to take notes as their only means of retaining specific information across compactions. As of September 2026, one card from the base game is banned; see “Methodology updates” for more details.

Each model is run at a compaction threshold of 250k tokens, with the exception of Claude Fable series models that we have found to be the only models which significantly benefit from a higher compaction setting. Previously, we ran models at two compaction settings and took the better result; for the benchmark leaderboard, we retain the better result from these older runs (roughly those run before September 2026).

Each playthrough consists of the first five in-game days of the game’s campaign. Across these days, and within the subset of the full game provided in our digital implementation, players can score up to 21 distinct objectives, consisting of a mix of missions completed, rewards unlocked, and nontrivial in-game “notable events” recorded. The top score from any human so far is 21 out of 21.

The agent’s system prompt contains basic information about the nature of the experiment (including that the agent is responsible for discovering the scored objectives), the scoring methodology, a tool-calling guide, and a compressed version of the game’s rulebook. The agent’s tools consist of a card lookup utility, an image of the game’s map, and a text-based pathfinding utility for models that struggle with vision. The agent can also read and edit a notes directory, and can read a set of whitelisted reference files, consisting of the game’s full rulebook, its rules glossary, and a digital interface guide that notes differences between the actual game and the provided digital implementation. Note that under the “Card ban” setting, the card in question is completely removed from all reference materials, as well as being removed from gameplay.

Models act through either a simple ReAct agent harness or a more complex Deep Agent harness that allows for subagents. See “Methodology updates” for more details.

As mentioned, we currently evaluate only the game’s initial five-day introductory segment; full campaigns span 20–25 days.

Human Baselines

Note: Human Baseline results were previously published in detail on this page, but have since been removed except as reference lines. See our tweet thread of human baseline results for more detailed data.

From July 15th to August 15th, 2026, we ran a human baseline experiment to measure the performance of humans with a variety of pre-existing experience levels at playing and learning this version of Earthborne Rangers. A total of 15 participants were recruited from among personal and professional networks of the benchmark creators, and participants were compensated at a rate of $30/hr, with a maximum participation time of 60 hours. A +25% bonus was provided for a topline score of 10 or above (“better than AI”); an additional +25% bonus (for a total of +50%) was provided for a topline score of 15 or above (“expert performance”).

Participants played in conditions as similar to the AI agents as possible. They were informed they had 10 playthroughs to use to learn the game, and their final two playthroughs would be scored; due to real-life considerations and the time-consuming nature of the experiment, participants were also given the option to declare they would end their participation early and pre-register their next two playthroughs (instead of their 9th and 10th) as their scored playthroughs.

Participants played the game through a private website created by Epoch, through a text-based interface built on the same game engine provided to the agents (with some affordances to make the UI more tolerable for humans, such as organizing card information in boxes and color-coding some keywords). To the greatest extent possible, humans and AI were provided with the exact same information about the game, and humans were strongly encouraged not to use the internet to try to help themselves (even though the internet broadly lacks useful information on performing well at this version of the game anyway).

Of the 15 initial participants, 2 dropped out early and have been excluded from the published dataset. We show the average and top participant scores as reference lines.

For our internal analysis of learning scores, we calculated them differently for human and AI agent players. AI agents’ learning scores are the difference between their topline (best-of-2) performance in a 10 or 18-playthrough run and their topline performance in an “ablated” 2-playthrough run where they are given no runway to learn and informed they must score as high as possible immediately. Since this type of ablation is not possible with humans, we instead simply take the difference between humans’ topline score from their terminal scored playthroughs and their initial two playthroughs. We do the same for the fatigue, injury, and travel metrics from playthrough associated with their topline score in each window. We call these “within-run learning scores” for humans, as opposed to “ablated learning scores” for AI.

This may inflate human learning scores, as humans were informed their initial playthroughs would not be scored, making it possible they would treat the playthroughs as exploratory or throwaway and not pursue objectives. However, Earthborne Rangers is generally considered a difficult game to learn, so we expect the initial playthrough scores and metrics would be relatively low anyways even if humans were instructed to try their best immediately. Additionally, based on a review of human gameplay transcripts, we did not observe anyone treating their initial playthroughs as “throwaway”. Participants made an earnest effort on their early playthroughs, and generally did not want to waste playthroughs not learning about any objectives.

Methodology updates

2026-09-22. We updated the benchmark to v4 methodology, making many changes to prepare for sustainably running it on rapidly advancing models until saturation. We also removed EBR-bench’s custom charts, due to operational difficulties in keeping them updated and readable.

First, we introduced an alternate multi-agent scaffold. Having observed models struggle with trying the same strategies repeatedly, we were curious whether subagents would help models better explore the strategic space of the game. The scaffold is based on Inspect’s Deep Agent harness for parity across providers, slightly augmented with head-agent-to-subagent messaging affordances to match a Claude Code/Codex capability. Up to four games of EBR can run concurrently, so we allow up to four subagents. The head agent can choose whether to play games itself, or delegate to subagents. Scored playthroughs are “locked” until all learning playthroughs complete.

Second, we introduced a substantive change to the content of the game by banning a card. The release of GPT-6 Astra showed a massive spike in performance that involved using a card that bypasses many of the strategic and tactical aspects of the game we hope to measure, such as fatigue management and efficient travel. While this is a legitimate approach that was also used by our top human baseliner, we believe the benchmark will present a better measurement of models’ holistic learning abilities if the card is banned, instead of measuring whether a model can figure out “one weird trick”.

While prior models have used this card, none figured out how to exploit it to its potential until Astra. We ran experiments that verified that several frontier or near-frontier models besides Astra did not experience a significant topline score impact from banning the card. Thus, we feel running the benchmark under this new setting, while unfortunately a retroactive change, will likely not affect existing score comparability.

The topline scores for Sol, Fable 5.1, and Opus 5 have been replaced with the results under this new “Card ban” setting, and GPT-6 Astra’s published score is also from a “Card ban” run; all other models’ scores predate the ban. Going forward, we plan to run new models on these three combinations of settings:

  • Card banned + single-agent (the default, and the leaderboard score)
  • Card unbanned + single-agent
  • Card banned + multi-agent

Since we decided to stop maintaining EBR-bench’s custom charts, the results of the above non-default experiments will not be reflected in the charts on this page. We plan to publish articles that feature any notable results from those side experiments instead.

Additionally, we made several other smaller changes:

  • We retired the practice of running every new model on two compaction settings to compare performance, due to budget constraints and the observation that only Claude Fable benefited from higher compaction. We plan to run all models at 250k compaction for the remainder of the benchmark’s lifespan, with the exception of Claude Fable series models which will be run at 90% of their advertised context window.

  • We updated the agent-facing prompt to more clearly specify they are responsible for discovering the 21 scoreable objectives, to bring it closer in line with the material given to our human baseline participants. This affects results from after August 21, 2026.

  • We reduced the sample size of larger experiments from 10 to 5, to accommodate the expanded number of experiments.

  • Due to budgetary constraints and consistent null results, we dropped 18-playthrough runs as a standard experiment.

2026-08-19. We added data from our human baseline experiment participants.

We recruited 15 participants who brought a variety of pre-existing experience with gaming in general and Earthborne Rangers specifically, playing under conditions as similar to the agents as possible. See the “Human Baselines” subsection of Methodology above for more details.

2026-08-11. We updated the benchmark to v3 methodology.

We found in further experiments that compaction threshold settings can have a significant impact (on the order of 1 to 3 objectives) on certain models with certain task parameters. As such, we now run every model at two compaction settings: 250,000 tokens, and a higher value around 90% of the model’s reported maximum context window (usually 880,000 for GPT models, since anything higher can cause issues with the compaction endpoint in rare cases, and around 950,000 for other models). We then report the better topline score between the two compaction settings as the model’s canonical topline score in the leaderboard.

Of the five models we’ve actively run experiments on (Claude Opus 4.8, Claude Opus 5, Claude Fable 5, GPT-5.5, and GPT-5.6 Sol), this change in methodology affected the topline scores for each as follows:

  • Claude Opus 4.8: Scored 3.4 percentage points better at the lower compaction setting.
  • Claude Opus 5: Scored 4.3 percentage points better at the lower compaction setting.
  • Claude Fable 5: Scored 11.9 percentage points worse at the lower compaction setting, so its reported topline score is left unchanged.
  • GPT-5.5: Scored 13.3 percentage points better at the lower compaction setting.
  • GPT-5.6 Sol: Scored 0.4 percentage points better at the lower compaction setting.

2026-07-27. We updated the benchmark to v2 methodology.

We changed the topline scoring from the mean of a sample’s final 20% of playthroughs to the best of its final two playthroughs, matching the setup of our ongoing human baseline studies. Each model is now evaluated with at least ten independent samples rather than one to three, considerably improving confidence intervals. We also introduced the ablated learning score.

We increased the context compaction threshold, from 250k tokens to approximately 90% of each model’s full context window. We are still conducting experiments to find a compaction policy that does not under-elicit model performance.

Scores for models not re-evaluated under v2 continue to reflect the v1 methodology; the two are not directly comparable. v1 datapoints can be identified visually by their lack of error bars.

We also changed the “strategy guide” used to maximize elicitation to a more detailed one with a specific deck recommendation and an explicit series of steps to follow in-game for a high score. This more accurately reflects our claim that the strategy guide is an “answer sheet” that reflects the expert human approach to maximizing score.

Human baseline experiments began during this period, so we also updated the agent-facing UI to be closer to the interface presented to the human baseliners. Mainly, agents now have a separate tool call that lets them view the contents of the game’s discard piles, rather than those discard piles’ contents always being in-context.

The agent-facing system prompt also now informs them of the human expert performance of 20/21 objectives.

Citations

Citation

Epoch AI, 'Earthborne Rangers (EBR-bench)'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks/ebr-bench' [online resource]. Accessed 29 Sep 2026.

BibTeX Citation

@misc{EpochEbrBench2026, title = {{Earthborne Rangers (EBR-bench)}}, author = {{Epoch AI}}, year = {2026}, month = {7}, url = {https://epoch.ai/benchmarks/ebr-bench}, note = {Accessed: 29 Sep 2026} }