Report
Oct. 8, 2026

Can AI automate Epoch?

Epoch Automation Reports show models can’t yet autonomously produce Epoch-quality work.

We gave six models real work tasks from Epoch, such as graphic generation and research design, to test how close they are to automating our work. While frontier models such as Claude Fable 5.1 and GPT-6 Astra perform reliably on well-defined tasks, they fail in the more open-ended aspects that prevent full automation.

Epoch Automation Reports captures these shortcomings. We curate a diverse set of tasks directly from our work at Epoch, provide the models with the necessary context, use the highest available reasoning settings, and have them attempt the tasks without further intervention. We then manually grade the outputs against our own employee standards.

These evaluations aim to capture gaps in typical benchmarks, which concentrate on easily verifiable tasks that capture only a subset of real work. These existing benchmarks may also overstate progress if a model has been optimized to score well on them.

By combining the systematic approach of benchmarks with the holistic impressions that come from testing models on real work, we expose capability gaps that existing benchmarks miss and inform broader questions about AI adoption.

Key findings:

  • Although Fable 5.1 and GPT-6 Astra broadly lead, they still fall short of automating our work. Open-weight models lag further behind, struggling even on the well-defined tasks that frontier models handle reliably.
  • Models struggle to pick up Epoch’s standards, even with ample reference material. They miss implicit conventions such as our visual style and the kinds of topics our audience cares about, or produce more information-dense outputs than our typical work.
  • Models can identify promising research directions, but lack judgment to successfully follow through. They struggle to design experiments that measure what they claim and treat flawed results as key findings.
  • Although several tasks involve open-ended exploration with a large idea space, models often converge on similar ideas.

Prior work has explored evaluating models on work automation, including GDPval, AutomationBench, and the Remote Labor Index (RLI). These benchmarks use realistic work tasks, but simplify them, limiting what can be tested. GDPval and RLI provide all necessary context and files up front: RLI keeps tasks self-contained, and GDPval excludes proprietary tools. In contrast, many of our tasks require models to gather context by using tools (such as Figma, Google Drive, and satellite imagery ordering sites) to reflect real work as closely as possible. AutomationBench tasks run in simulated environments scoped to what can be checked programmatically (e.g., was this email sent?). These automated checks are reliable and make evaluation easier to verify, but they leave the more open-ended qualities of these models underexplored.

Closest to our work is CRUX, which runs realistic, messy long-horizon tasks that are difficult to verify. Each CRUX evaluation focuses on a single task designed to test specific frontier capabilities. We share this emphasis on messy, open-ended evaluation, but we use a suite of everyday tasks from our own work to measure Epoch work automation rather than focus on task-specific capabilities.

Methodology

Tasks

Our task suite is based on what people do within Epoch. As of launch, this includes 11 tasks across five diverse categories.

Graphic Design. Create an Epoch-style graphic that illustrates a given description.

Tasks:

  • Create an Epoch-style diagram given a description and rough sketch of universal basic proposals.
  • Create an Epoch-style graphic to illustrate how model progress may be understated with a fixed token budget.
  • Create an Epoch-style conceptual diagram of the parity point: the point at which humans become more cost-effective than AIs.

Data Insight Generation. Create an Epoch-style Data Insight based on a provided direction.

Data Insights are short articles we regularly publish, which break down complex AI trends into focused snapshots.

Tasks:

  • Create an Epoch-style Data Insight using Epoch’s July 2026 polling data.
  • Create an Epoch-style Data Insight on AI usage in physics research with data from arXiv.
  • Create an Epoch-style Data Insight using any data source.

Data Explorer Generation. Create an Epoch-style Data Explorer based on a provided direction.

Data Explorers present key datasets tracking historical and contemporary progress in AI. They are designed so that viewers can easily browse and understand the data.

Tasks:

  • Create an Epoch Data Explorer on AI mentions in the Federal Reserve Board Beige Book over time, given the website repository and specific features to include.
  • Create an Epoch Data Explorer on AI mentions in the Federal Reserve Board Beige Book over time, given the website repository.

AI Data Center Research. Research and analyze a given AI data center to add to Epoch’s AI Data Centers database, given a detailed guide.

Tasks:

  • EcoDataCenter 1
  • Tesla Cortex 1

Research Design. Propose a novel research project and run a pilot experiment to support further experimentation.

Task:

Setup

We test models on readily available systems on their highest reasoning settings:

ModelHarnessReasoning Effort
GPT-6 AstraCodexUltra
Claude Fable 5.1Claude CodeUltracode
Grok 4.6Grok Buildxhigh
Gemini 3.8 FlashAntigravityHigh
Kimi K3Kimi CodeMax
Qwen 3.8 MaxQwen CodeMax

For each task, we provide models with the relevant resources and full permissions within their own workspace and Google account. For example, for graphic design tasks, we provide models with a Figma integration and a Figma file with as much relevant information as possible (previous designs, reusable components, design guidelines). In general, we follow the principle of providing as much context to the model as we would to a new hire.

Evaluation

Each task category has a rubric that was developed with Epoch employees who have expertise in that task. We further developed the rubrics through an initial qualitative evaluation of these tasks, which identified common failure patterns and turned them into evaluation criteria.

Each rubric includes a mix of objective and subjective criteria. For example, the Data Insight Generation task includes:

  • Objective: “Pre-chart and post-chart paragraphs each under 150 words”
  • Subjective: “Title is clearly communicated”

Each output is manually reviewed and graded by a single human grader. We currently use this approach for grading expedience, and may add multiple graders in the future.

Our key findings come from qualitative observations, as we believe they best capture gaps in open-ended tasks. These observations come not only from the model outputs, but also from the trajectories, which help us understand what may have led to them. We use the rubric alongside these observations to compare models and track progress; however, these numbers alone do not capture the full picture and should be considered alongside the qualitative observations.

Fable 5.1 and GPT-6 Astra are broadly tied in the lead

Of the six models we evaluated, Fable 5.1 and GPT-6 Astra achieve the highest aggregate scores and lead on most individual task categories as well.

Horizontal bar chart of average task performance across five Epoch work task categories for six leading models showing leading models cannot yet autonomously complete Epoch work tasks.

Their lead comes mostly from reliability on the well-defined parts of our tasks, such as coding and computational analysis, where their outputs were consistently accurate. Computer use was also far more capable than we expected. Earlier this year, frontier closed-weight models failed at tasks such as porting an article to Substack. When we tried this again with newer models, they succeeded.

We even removed a graphic design task from our suite since Fable 5.1 produced an Epoch-quality output. This design task was more straightforward than others: given a Matplotlib chart and our website repository of instructions and reusable chart components, evaluated models were asked to restyle the chart to match Epoch’s house style. We now largely automate this step in our own work.

Even so, these frontier closed-weight models fall short of fully automating our work due to gaps in the more open-ended parts. One common pattern was failing to pick up Epoch’s standards, despite being given ample reference material. These models consistently missed implicit conventions that humans would pick up, such as our visual style or what our audience cares about. Their research judgment was also weak: they struggled to design informative experiments and often presented results from flawed setups as key findings.

Open-weight models struggle where closed-weight models are reliable

Open-weight models lag further behind. They struggle not only on the open-ended parts of tasks, but also on the well-defined parts that closed-weight models handle reliably.

We can see this through one of our Data Insight Generation tasks, where models are asked to produce a data insight using arXiv data on physics papers. Kimi K3 based its entire insight on a data filtering error:

Inaccurate line chart produced by Kimi K3 as part of “arXiv” task, showing the share of physics submissions mentioning machine learning versus large language models from 2000 to 2026, with a ChatGPT launch marker at Nov 2022.

Kimi K3 output on “arXiv” task

Although AI mentions in physics papers have increased since ChatGPT, they have not increased as sharply as presented here. Kimi K3 computed these numbers across all arXiv papers, not just physics papers, and never checked whether its filter worked. None of the frontier closed-weight models made factual errors in their data insights.

These factual inaccuracies show up in other task outputs as well, such as Kimi K3’s wealth redistribution diagram where returns flow from people into capital assets:

Inaccurate flow diagram by Kimi K3 for “Wealth Redistribution” task, comparing three redistribution proposals — Universal Basic Income, Universal Basic Services, and Universal Basic Capital — showing how government revenue reaches people as cash, free services, or income-generating assets.

Kimi K3 output on “Wealth Redistribution” task

The prompt describes people earning returns from capital assets, and the generated caption refers to “assets that pay a return”, but the visual does not follow this described flow. Despite Kimi K3 repeatedly working on this part of the diagram, its fixation on minor design details caused it to miss this larger communication error. In contrast, all frontier closed-weight model graphics were factually accurate.

This capability gap is larger than existing benchmark scores suggest. For example, Kimi K3 scores 158 on the Epoch Capabilities Index (ECI), which combines results from many AI benchmarks into a single measure of general capability. This is roughly tied with closed-weight models like Grok 4.6, despite Kimi K3 struggling on basic tasks that Grok 4.6 handles more reliably. Realistic open-ended tasks like these may reveal weaknesses that existing benchmarks miss.

Models fail to learn generalizable patterns from examples

Since we are testing for the ability to complete Epoch-quality work end-to-end, it is important for models to be able to pick up and understand Epoch’s standards. We provide references for task inputs to facilitate this process.

In the following task, we ask models to generate a graphic and provide sufficient reference material to convey Epoch’s simple house style and general design principles. We provide a link to our Gradient Updates page, which includes reference graphics, and a link to our Figma file, containing several months of design work, our color palette, design guidelines, templates, and reusable components.

I’m working on a gradient updates article.

Create a diagram to illustrate this: “Models have been improving over time in both performance and a fixed token budget but also their ability to productively use large token budgets. This means that constant-elicitation understates recent progress relative to max elicitation.“

Previous gradient updates articles can be found here: https://epoch.ai/gradient-updates

Here is the figma file we use for graphics generation (treat this as read-only): [Figma link]

Prioritize graphic design principles when creating your diagram. Your final output should be an epochified diagram.

Prompt for “Model Performance Chart” task

Screenshot of a Figma design file named "Epoch AI Designs," with a page list of resources and canvas frames grouped into "Web graphs", "Social images", and "Data" sections.

Screenshot from Figma file

The underlying idea is simple, so the diagram should be too. An Epoch designer given the task illustrated it with this simple chart.

Schematic line chart of performance versus token budget, made by an Epoch designer for the “Model Performance Chart” task, showing a newer model's curve above an older model's, with annotations noting larger gains at higher token budgets.

Epoch designer output for “Model Performance Chart”

Fable 5.1 produced the following instead. It is far more information-heavy than our typical graphics and, to our eye, the additional information serves no explanatory purpose.

Overly complex two-panel illustrative diagram, made by Fable 5.1 for the “Model Performance Chart” task: left line chart, labeled A, of benchmark score versus token budget for 2024, 2025 and 2026 models, right chart, labeled B, comparing constant versus max elicitation progress by release year.

Fable 5.1 on “Model Performance Chart” task

Fable 5.1 borrows from a particularly complex reference example it found, shown below. This example uses the “A” and “B” badges to label the two concepts being compared, and they earn their place because each concept shows up in more than one part of the diagram. However, Fable 5.1 did not learn this standard and simply used these badges for one-off labeling.

Schematic Epoch-produced diagram, from Fable 5.1’s reference search, comparing microbump (labeled A) and hybrid bonding chip interconnects (labeled B) showing cross-sections, pitch and density figures, and a third dot-grid illustration of ~44× more connections per area with A and B labels used.

Epoch diagram from Fable 5.1’s reference search

This failure to learn Epoch’s standards is also demonstrated by GPT-6 Astra’s attempt to choose an interesting topic that would appeal to our audience.

We asked the model to write an Epoch-style Data Insight, a short article on an AI trend. We gave the model an article focus (AI usage in physics research) and a link to all our published Data Insights, which show a clear pattern of appealing to a general audience interested in AI.

For comparison, we published a Data Insight on AI usage in math research shortly after running the models on this task. This Data Insight also uses arXiv paper data (from math) and focuses on a broad trend that is easy for a general audience to understand.

Line chart from actual Epoch “arXiv” Data Insight showing the share of math preprints acknowledging AI use, for any purpose and for substantial research.

Actual Epoch “arXiv” Data Insight

However, GPT-6 Astra chose a topic that is too niche for such an audience:

Line chart, produced by GPT-6 Astra as part of the “arXiv” task, of papers per year in the HEP-ML Living Review, showing normalizing flows and diffusion models overtaking GANs in particle-physics AI bibliography.

GPT-6 Astra on “arXiv” task

This topic is specific to the physics domain rather than a general AI trend that Epoch’s audience would find interesting. Although the task involves physics arXiv data, GPT-6 Astra should have enough evidence in its references to determine a more fitting topic.

Models generate promising research directions, but struggle to design experiments that are actually informative

We task models with proposing a research project based on Epoch’s “9 big questions benchmarks can help answer”. This process involves choosing a novel research question, designing and running a pilot experiment, writing up the pilot results, and then proposing follow-up experiments.

The most important part of this process is experiment design. It’s not difficult to identify a valuable general direction for an experiment. For instance, an experiment that reveals exactly which jobs AI will automate next year would clearly be worth running. The challenge is in figuring out how to execute this experiment in a way that actually measures what we are looking for, which is what models struggle with.

For example, GPT-6 Astra chose the following research question, which could genuinely be a promising direction:

When an AI agent fails to improve with practice, has it failed to collect useful evidence, or failed to understand evidence it already has?

Its proposed experiment is also directionally good, as it involves realistic tasks that would help make the results more generalizable:

Excerpt of model output by GPT-6 Astra as part of the “Benchmarking (9 Questions)” task, outlining an experiment to test a diagnostic on a realistic task, describing the unlikely execution of an out-of-sample transfer test.

GPT-6 Astra on “Benchmarking (9 Questions)” task

However, it’s unclear how this experiment would be executed in practice. For example, the description does not specify what “matched externally selected experience” refers to or how “simulator-derived or restricted-policy reference bounds” would be used, even though both are critical parts of the experiment. A research proposal is only valuable if its experiments can be carried out and actually measure what they claim to.

Models present results driven by flaws in their own experiment setup as key findings

Alongside designing vague experiments, models also make mistakes setting up the experiments that they do run. Such mistakes are common in research, so we do not penalize models for them. What matters is how they handle the flawed results in informing next steps. We find that models sometimes recognize these mistakes, but treat the flawed results as key findings that drive the rest of the proposal.

For example, consider GPT-6 Astra’s research proposal. It asks whether AI agents fail to improve because they don’t collect useful evidence or use existing evidence. In the pilot experiment, it tested whether models could figure out the hidden rule behind a simple puzzle by trying a limited number of inputs and seeing their outputs. GPT-6 Astra compared how well they did when choosing their own inputs versus when given inputs designed to reveal the rule.

The setup flaw was that it gave models too low a token budget. The models Astra tested in its experiment could output up to 4096 tokens while choosing the input, which cut off 61 of 280 responses before they could name an input at all. These cut-off responses still counted as valid attempts, so the tested models had to guess the hidden rule with less information than they were supposed to get. GPT-6 Astra recognized this error and increased the token budget, which led to a significant improvement across the tested models. These results show that the 4096-token budget is insufficient to elicit a complete response, which is not relevant to the core research question.

In its write-up, GPT-6 Astra mentions that the low token budget lowered scores and that raising it brought them back up. However, it claims that these results reveal “sensitivity to the acquisition budget”, where acquisition budget refers to the token budget models get for choosing inputs:

Text excerpt from GPT-6 Astra as part of the “Benchmarking (9 Questions)” task, reporting pilot findings on rule recovery rates for tested models, token allowance exhaustion and truncations, and total pilot spending across API calls.

GPT-6 Astra on “Benchmarking (9 Questions)” task

This framing is misleading because it implies models had a fair chance to produce a complete response and failed, when in reality they were unable to produce any answer. The performance differences therefore reflect incomplete responses rather than meaningful differences in reasoning effort during acquisition.

Fable 5.1 produces outputs with extremely high information density

In general, we want output to be presented in a simple and intuitive manner. For instance, a research report’s executive summary should be as concise as possible while preserving the main points. But models tend to introduce unnecessary complexity, especially Fable 5.1.

For example, consider the summary section of Fable 5.1’s research proposal and pilot experiment write-up.

Text excerpt from Fable 5.1 as part of the “Benchmarking (9 Questions)” task, showing a verbose executive summary.

Fable 5.1 on “Benchmarking (9 Questions)” task

This general pattern shows up not only in writing but also in other tasks, like graphic design. For example, Fable 5.1 leans towards more verbose captions, which contrasts with our core design principle of communicating clearly with as little detail as possible.

Flow diagram produced by Fable 5.1 in the “Wealth Redistribution” task, comparing three post-AGI wealth distribution proposals — Universal Basic Income, Services, and Capital — each with lengthy descriptions.

Fable 5.1 on “Wealth Redistribution” task

The reasoning for this is revealed in Fable 5.1’s rounds of self-critique when it pushes for more descriptive captions:

Descriptions are thin paraphrases: UBS ‘Everyone gets free public services’ drops ‘covering all basic needs, including food and housing’; UBC ‘Everyone gets capital that earns a return’ drops ‘their own’

Feedback from Fable 5.1 subagent

Models converge on similar ideas, despite a large idea space

On open-ended tasks with many possible directions, we may expect models to produce a diverse range of ideas. However, we find that models often converge on similar ideas.

For example, when we tasked models with creating a Data Insight using Epoch’s polling data, Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 all converged on the same topic:

Side-by-side comparison of stacked bar charts produced by Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 in the “Polling” task, each showing work versus personal use shares for models.

Fable 5.1, Gemini 3.8 Flash, and Grok 4.6 on “Polling” task

The idea space is somewhat more limited here, since models have to work within the polling dataset and avoid topics that have already been published. Yet there are still enough candidates that we did not expect three out of six models to pick the same one.

We also observe this with Kimi K3 and Qwen 3.8 Max on our research design task where they could generate any idea out of nine broad directions:

Side-by-side screenshots of two AI-generated research proposals, by Kimi K3 and Qwen 3.8 Max in the “Benchmarking (9 Questions)” task, each showing a title relating to self-evaluation, plus a summary that is similar in direction.

Kimi K3 and Qwen 3.8 Max on “Benchmarking (9 Questions)” task

Limitations

We acknowledge the following limitations of this benchmark:

Small sample size. We have a limited task suite, and we run each model only once per task. Given that manual grading is time-intensive, we prioritize in-depth qualitative analysis over more runs. Scores are therefore noisy and may shift with more runs. These results should not be read as a general measure of what AI can or cannot do across all work tasks.

Subjective grading. Many of our evaluation criteria involve subjective judgment, so the numerical scores should not be treated as ground truth. Specific numbers could change with different rubric item definitions or different graders. We include these numbers to give a rough sense of how models compare and how close each task is to being fully automated. We believe that the qualitative findings are more informative than these exact numbers.

Epoch-specific tasks. Our tasks are scoped to Epoch work, and thus are not representative of all work tasks. We also evaluate these model outputs by Epoch’s standards, which may not align with how other work tasks should be evaluated.

Different agent harnesses. We test each model on its own harness, such as GPT-6 Astra in Codex. Some observed differences may reflect the harness rather than the model itself. For instance, GPT-6 Astra in Codex with its built-in browser extension may help boost performance compared to Qwen 3.8 Max in Qwen Code with the Chrome DevTools MCP. We chose this setup to reflect how people actually use these models in their work, but we acknowledge it may introduce advantages for some models over others.

Conclusion

The frontier is evolving rapidly, but we still see gaps in the more open-ended parts of work. The best-performing closed-weight models are reliable on well-defined tasks, such as coding, data analysis, and computer use. However, they consistently fall short on the judgment that drives real work. This includes learning generalizable patterns from examples, designing good research experiments, and drawing meaningful conclusions from results.

In studying AI capabilities at open-ended tasks drawn from Epoch, we have two goals. First, we aim to provide a fuller view of AI capabilities than current benchmarks, which typically focus on well-defined, easy-to-grade tasks. Second, we want to shed light on how close AI is to automating work end-to-end. We find that it cannot yet replace workers, at least not at Epoch. We plan to add results as new models are released, informing questions about AI adoption and automation over time.