Back to benchmarks

Epoch Automation Reports

By Epoch
Oct. 8, 2026

A benchmark of realistic, open-ended tasks drawn from Epoch AI’s own work. Models produce work outputs, which human graders evaluate against internal quality standards.

Agent
Epoch administered

Rubric scores by system and task category

Mean rubric scores across each task category. A score of 100% means the output meets the internal standard we’d expect from an Epoch employee. Hover over a cell to see individual task scores.

System
Graphic design
Data insight generation
Data explorer generation
AI data center research
Research design
Average
Claude Fable 5.1Claude Code
65%
GPT-6 AstraCodex
65%
Grok 4.6Grok Build
59%
Qwen 3.8 MaxQwen Code
53%
Kimi K3Kimi Code
52%
Gemini 3.8 FlashAntigravity
42%

Explore the rubric

Pick a system to see how it did on each task

Claude Fable 5.1

Claude Code

Wealth redistribution

Create an Epoch-style diagram for universal basic proposals given a rough sketch.

Score
70.7%
Time
1h 20m
Cost
$75.25
Tokens
26.2M
Final output
Claude Fable 5.1's final output for Wealth redistribution
This is an AI generated output, and should not be cited or shared as Epoch's work.
Full prompt
Attachment to the Wealth redistribution prompt

I’m working on a gradient updates article.

Create a diagram to illustrate this:

What can we do if AI takes all our jobs, and most people can’t work to pay their bills? People have thought about this question for a while now, and the answer seems to be Universal Basic ________.

What goes in the blank? There are a few options:

  • Universal Basic Income (UBI): The government pays everyone cash to meet their basic needs. This is the most well-known one by far, and has been proposed by the likes of Elon Musk, Vinod Khosla, and Geoffrey Hinton, to name but a few.
  • Universal Basic Services (UBS): Give everyone access to free public services. Think welfare states like Norway or Sweden, except that the government covers all basic needs, including things like food and housing which the Nordics currently don’t provide.
  • Universal Basic Capital (UBC): The government gives everyone their own capital assets (such as equity), which they can earn a return on to pay for their basic needs.

At first glance, they don’t sound too different — the government gives people something to make ends meet, whether cash, public services, or capital assets.

Previous gradient updates articles can be found here: https://epoch.ai/gradient-updates

Here is the figma file we use for graphics generation (treat this as read-only):
[Figma link]

Prioritize graphic design principles when creating your diagram.

I’ve also attached a conceptual diagram of what I want. Your task is to epochify this diagram, following this as a general guide (you do not have to follow it exactly and may make changes as you see fit).

Your final output should be an epochified diagram.

Rubric
Communication
40% weighted
2/3
  • Accurate
  • Helpful
  • Does not assume additional information
General design knowledge
40% weighted
3/5
  • Follows basic design rules
  • Has good color contrast
  • Has visual balance
  • Has intentional color selection
  • Avoids repetition
Epoch-specific rules
20% weighted
3/3
  • Follows Epoch’s general style
  • Found color ordering in Figma
  • Applied color ordering in final output

Epoch Automation Reports

Epoch Automation Reports is Epoch AI’s benchmark of realistic, open-ended tasks drawn from our own work. For each task, we give a model the relevant context and tools and ask it to produce an output. A human grader then evaluates the output using rubrics that reflect our standards. Tasks span a diverse range, from graphic design to AI data center research, and we plan to add more over time.

Most existing benchmarks focus on well-defined tasks with easily verifiable outputs. In contrast, real work often calls for creativity and judgment in ambiguous situations. Existing benchmarks do not fully capture these open-ended aspects, creating a gap between benchmark scores and real-world performance.

We aim to combine the systematic approach of benchmarks with the holistic impressions that come from testing models on real work. We hope this sheds light on what easily verifiable benchmarks might miss and informs broader questions about AI adoption. At the minimum, it should inform our own adoption: if models do very well, we can automate these tasks within Epoch.

For more details, see our report, Can AI automate Epoch?.

Methodology

Overview

We build our task suite based on the work people actually do within Epoch AI. This currently includes 11 tasks across five diverse categories. We run each model on its corresponding agent harness at its highest reasoning setting and provide task-specific resources. As a general principle, we give models as much context as we would give a human contractor. We also disable memory and skill writing so what a model learns on one task does not carry over to future tasks. A human grader then scores each output against rubrics that reflect our internal quality standards.

Graphic design

Task. Create an Epoch-style graphic that illustrates a given description.

Purpose. This task tests how well the model learns patterns from provided resources and communicates ideas clearly.

Provided tools/resources.

  • Figma Remote MCP server where supported (e.g., Claude Code), and Figma Desktop MCP server otherwise (e.g., Qwen Code).
  • A link to a Figma file with Epoch graphics from June-August 2026, along with Epoch’s color palette, design guidelines, templates, and core components.
  • A link to Epoch’s Gradient Updates page, which includes graphics that can be used as further reference.

Evaluation. Grading prioritizes communication and general design principles. Communication covers how well the graphic illustrates the description as simply as possible. A high-quality graphic should not simply restate the description; it should help viewers understand the message better than the text alone. General design principles, such as visual balance, spacing, and alignment, serve as the baseline for an effective graphic.

Data insight generation

Task. Create an Epoch-style Data Insight based on a provided direction.

Data Insights are short articles regularly published by Epoch AI, which break down complex AI trends into focused snapshots.

Purpose. This task tests how well the model identifies patterns in the provided resources, judges interesting topics, and communicates ideas clearly.

Provided tools/resources.

  • Google Drive MCP server, so models can read the provided Google Doc. Some agent harnesses configure this automatically (e.g., Claude Code), while others require manual setup through the Google Cloud Console (e.g., Qwen Code).
  • A link to an internal Google Doc containing the Data Insight template and a corresponding checklist. The template outlines key sections, each with a description of what should be contained. The checklist covers what is typically reviewed before publication, such as “Is the title 80–90 characters (or less)?”
  • A link to Epoch’s Data Insights page for reference.
  • An adapted version of Epoch’s internal website repository, scoped to relevant parts for data insight generation.

Evaluation. Grading weighs the title and plot most heavily — since most readers look only at the title and plot, with the text providing secondary context. These two elements alone should communicate the insight clearly.

Data explorer generation

Task. Create an Epoch-style Data Explorer based on a provided direction.

Epoch’s Data Explorers present key datasets tracking historical and contemporary progress in AI. They are designed so that viewers can easily browse and understand the data.

Purpose. This task tests how well the model learns patterns from provided resources, collects and analyzes external data, and presents it intuitively.

Provided tools/resources.

  • Epoch’s internal website repository, including the source code for all Data Explorers, reusable components, and markdown files containing specific guidelines on Epoch’s website design.
  • A link to Epoch’s Data Explorers page for reference.

Evaluation. Grading prioritizes overall design and consistency with Epoch’s style. Data Explorers should be intuitive to navigate and aid in understanding data. They should also match Epoch’s existing explorers, reusing components where possible and following established layouts.

AI data center research

Task. Research and analyze a given AI data center to add to Epoch’s AI Data Centers database.

Purpose. This task tests how well the model follows instructions, handles computer use, analyzes detailed images, and sustains progress over long horizons.

Provided tools/resources.

  • Airtable MCP server, connected to model-specific Airtable base.
  • SkyWatch MCP server, for searching satellite imagery.
  • Google Drive MCP server, for reading and writing Google Docs/Sheets, and uploading files to folders.
  • QGIS MCP server, for image analysis.
  • Model-specific Google account.
  • Model-specific SkyWatch Explore account, with credit card set up for automatic checkout.
  • Browser access through provider-specific extensions where available (e.g., Claude in Chrome), and through the Chrome DevTools MCP otherwise.
  • Computer use through native tools where available (e.g., Claude Computer Use), and through CUA Driver MCP otherwise.
  • A Google Doc link containing a detailed guide for researching and analyzing a data center. This is the same guide given to contractors, adapted for agents (e.g., allowing them to purchase images without approval).

Evaluation. Grading is based on accuracy, using a rubric adapted from an internal peer review checklist for AI data center entries. The gold reference is a complete, reviewed AI data center entry that is not yet published when the task runs, so the model cannot find it through Epoch’s AI Data Centers database.

If the gold reference does not contain enough information to assess a specific criterion (e.g., if there is no information on cooling), we exclude that criterion from the rubric and grade the model on the remaining items.

If an expected output is missing from the model’s Airtable entry, we check its other output files, such as its notes and calculations.

We grade evaluation criteria independently to prevent errors from cascading. For example, calculating cooling capacity depends on first getting the correct number of fans per unit. We grade the fan count and the calculation method separately, so a model that gets the fan count wrong but applies the correct method still earns the point for the method.

The guide instructs models to focus on buildings completed in 2024 or later, so we do not penalize models for how they handle earlier buildings. If a model includes pre-2024 buildings, we grade them normally; if they are left out, we only grade the included buildings.

Tasks occasionally require human intervention. The most common case involves SkyWatch imagery, which can take up to 48 hours to arrive after ordering. Models can set up an automated check for this to make the task fully autonomous. If the check does not work as intended or is not set up at all, a human has to notify the model when images arrive. Interventions are also needed when a model’s mistake causes API errors (e.g., sending an image too large for the API request). In these cases, we start a new session and point the agent to its previous trace so it can pick up where it left off. Each intervention deducts points. Permission approvals such as order approvals do not count as point deductions.

Research design

Task. Propose a novel research project and run a pilot experiment to support further experimentation.

Purpose. This task tests the model’s ability to generate novel ideas, design experiments, and interpret results.

Provided tools/resources.

Evaluation. Grading prioritizes research design. We assess whether the model’s experiments are set up to answer the research question and whether it draws the right conclusions from the results. We do not expect pilot experiments to succeed, and therefore do not grade based on experiment success. We focus on how the model interprets the results. A failed experiment can still reveal how future experiments should be redirected, so we grade the model’s analysis and approach to further experimentation rather than the results themselves.

Limitations

Small sample size. We have a limited task suite, and we run each model once per task. The scores are therefore noisy and may shift with more runs. These results should not be read as a general measure of what AI can or cannot do across all work tasks.

Subjective grading. Scoring rubrics involve subjective judgment, so numerical scores should not be treated as ground truth. Specific numbers could change with different rubric item definitions or different graders. We include these numbers to give a rough sense of how models compare and how close each task is to being fully automated. We believe that qualitative findings are more informative than these exact numbers, which we publish through our Epoch Automation Reports.

Epoch-specific tasks. Our tasks are scoped to Epoch work, and thus are not representative of all work tasks. We also evaluate these model outputs by Epoch’s standards, which may not align with how other work tasks should be evaluated.

Different agent harnesses. We test each model on its own harness, such as GPT-6 Astra in Codex. Some differences we observe may be influenced by the harness rather than the model itself. For instance, GPT-6 Astra in Codex with its built-in browser extension may help boost performance compared to Qwen 3.8 Max in Qwen Code with the Chrome DevTools MCP. We chose this setup to reflect how people actually use these models in their work, but we acknowledge that it may introduce advantages for some models over others.

Future work

We plan to expand this task suite over time, adding new tasks and retiring those that become saturated. As new models are released, we will also publish updated results along with commentary on their performance.

Citations

Citation

Epoch AI, 'Epoch Automation Reports'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/benchmarks/epoch-automation-reports' [online resource]. Accessed 8 Oct 2026.

BibTeX Citation

@misc{EpochEpochAutomationReports2026, title = {{Epoch Automation Reports}}, author = {{Epoch AI}}, year = {2026}, month = {10}, url = {https://epoch.ai/benchmarks/epoch-automation-reports}, note = {Accessed: 8 Oct 2026} }