Overview
We build our task suite based on the work people actually do within Epoch AI.
This currently includes 11 tasks across five diverse categories.
We run each model on its corresponding agent harness at its highest reasoning setting and provide task-specific resources.
As a general principle, we give models as much context as we would give a human contractor.
We also disable memory and skill writing so what a model learns on one task does not carry over to future tasks.
A human grader then scores each output against rubrics that reflect our internal quality standards.
Graphic design
Task. Create an Epoch-style graphic that illustrates a given description.
Purpose. This task tests how well the model learns patterns from provided resources and communicates ideas clearly.
Provided tools/resources.
- Figma Remote MCP server where supported (e.g., Claude Code), and Figma Desktop MCP server otherwise (e.g., Qwen Code).
- A link to a Figma file with Epoch graphics from June-August 2026, along with Epoch’s color palette, design guidelines, templates, and core components.
- A link to Epoch’s Gradient Updates page, which includes graphics that can be used as further reference.
Evaluation. Grading prioritizes communication and general design principles.
Communication covers how well the graphic illustrates the description as simply as possible.
A high-quality graphic should not simply restate the description; it should help viewers understand the message better than the text alone.
General design principles, such as visual balance, spacing, and alignment, serve as the baseline for an effective graphic.
Data insight generation
Task. Create an Epoch-style Data Insight based on a provided direction.
Data Insights are short articles regularly published by Epoch AI, which break down complex AI trends into focused snapshots.
Purpose. This task tests how well the model identifies patterns in the provided resources, judges interesting topics, and communicates ideas clearly.
Provided tools/resources.
- Google Drive MCP server, so models can read the provided Google Doc. Some agent harnesses configure this automatically (e.g., Claude Code), while others require manual setup through the Google Cloud Console (e.g., Qwen Code).
- A link to an internal Google Doc containing the Data Insight template and a corresponding checklist. The template outlines key sections, each with a description of what should be contained. The checklist covers what is typically reviewed before publication, such as “Is the title 80–90 characters (or less)?”
- A link to Epoch’s Data Insights page for reference.
- An adapted version of Epoch’s internal website repository, scoped to relevant parts for data insight generation.
Evaluation. Grading weighs the title and plot most heavily — since most readers look only at the title and plot, with the text providing secondary context.
These two elements alone should communicate the insight clearly.
Data explorer generation
Task. Create an Epoch-style Data Explorer based on a provided direction.
Epoch’s Data Explorers present key datasets tracking historical and contemporary progress in AI.
They are designed so that viewers can easily browse and understand the data.
Purpose. This task tests how well the model learns patterns from provided resources, collects and analyzes external data, and presents it intuitively.
Provided tools/resources.
- Epoch’s internal website repository, including the source code for all Data Explorers, reusable components, and markdown files containing specific guidelines on Epoch’s website design.
- A link to Epoch’s Data Explorers page for reference.
Evaluation. Grading prioritizes overall design and consistency with Epoch’s style.
Data Explorers should be intuitive to navigate and aid in understanding data.
They should also match Epoch’s existing explorers, reusing components where possible and following established layouts.
AI data center research
Task. Research and analyze a given AI data center to add to Epoch’s AI Data Centers database.
Purpose. This task tests how well the model follows instructions, handles computer use, analyzes detailed images, and sustains progress over long horizons.
Provided tools/resources.
- Airtable MCP server, connected to model-specific Airtable base.
- SkyWatch MCP server, for searching satellite imagery.
- Google Drive MCP server, for reading and writing Google Docs/Sheets, and uploading files to folders.
- QGIS MCP server, for image analysis.
- Model-specific Google account.
- Model-specific SkyWatch Explore account, with credit card set up for automatic checkout.
- Browser access through provider-specific extensions where available (e.g., Claude in Chrome), and through the Chrome DevTools MCP otherwise.
- Computer use through native tools where available (e.g., Claude Computer Use), and through CUA Driver MCP otherwise.
- A Google Doc link containing a detailed guide for researching and analyzing a data center. This is the same guide given to contractors, adapted for agents (e.g., allowing them to purchase images without approval).
Evaluation. Grading is based on accuracy, using a rubric adapted from an internal peer review checklist for AI data center entries.
The gold reference is a complete, reviewed AI data center entry that is not yet published when the task runs, so the model cannot find it through Epoch’s AI Data Centers database.
If the gold reference does not contain enough information to assess a specific criterion (e.g., if there is no information on cooling), we exclude that criterion from the rubric and grade the model on the remaining items.
If an expected output is missing from the model’s Airtable entry, we check its other output files, such as its notes and calculations.
We grade evaluation criteria independently to prevent errors from cascading.
For example, calculating cooling capacity depends on first getting the correct number of fans per unit.
We grade the fan count and the calculation method separately, so a model that gets the fan count wrong but applies the correct method still earns the point for the method.
The guide instructs models to focus on buildings completed in 2024 or later, so we do not penalize models for how they handle earlier buildings.
If a model includes pre-2024 buildings, we grade them normally; if they are left out, we only grade the included buildings.
Tasks occasionally require human intervention.
The most common case involves SkyWatch imagery, which can take up to 48 hours to arrive after ordering.
Models can set up an automated check for this to make the task fully autonomous.
If the check does not work as intended or is not set up at all, a human has to notify the model when images arrive.
Interventions are also needed when a model’s mistake causes API errors (e.g., sending an image too large for the API request).
In these cases, we start a new session and point the agent to its previous trace so it can pick up where it left off.
Each intervention deducts points.
Permission approvals such as order approvals do not count as point deductions.
Research design
Task. Propose a novel research project and run a pilot experiment to support further experimentation.
Purpose. This task tests the model’s ability to generate novel ideas, design experiments, and interpret results.
Provided tools/resources.
Evaluation. Grading prioritizes research design.
We assess whether the model’s experiments are set up to answer the research question and whether it draws the right conclusions from the results.
We do not expect pilot experiments to succeed, and therefore do not grade based on experiment success.
We focus on how the model interprets the results.
A failed experiment can still reveal how future experiments should be redirected, so we grade the model’s analysis and approach to further experimentation rather than the results themselves.
Limitations
Small sample size. We have a limited task suite, and we run each model once per task.
The scores are therefore noisy and may shift with more runs.
These results should not be read as a general measure of what AI can or cannot do across all work tasks.
Subjective grading. Scoring rubrics involve subjective judgment, so numerical scores should not be treated as ground truth.
Specific numbers could change with different rubric item definitions or different graders.
We include these numbers to give a rough sense of how models compare and how close each task is to being fully automated.
We believe that qualitative findings are more informative than these exact numbers, which we publish through our Epoch Automation Reports.
Epoch-specific tasks. Our tasks are scoped to Epoch work, and thus are not representative of all work tasks.
We also evaluate these model outputs by Epoch’s standards, which may not align with how other work tasks should be evaluated.
Different agent harnesses. We test each model on its own harness, such as GPT-6 Astra in Codex.
Some differences we observe may be influenced by the harness rather than the model itself.
For instance, GPT-6 Astra in Codex with its built-in browser extension may help boost performance compared to Qwen 3.8 Max in Qwen Code with the Chrome DevTools MCP.
We chose this setup to reflect how people actually use these models in their work, but we acknowledge that it may introduce advantages for some models over others.
Future work
We plan to expand this task suite over time, adding new tasks and retiring those that become saturated.
As new models are released, we will also publish updated results along with commentary on their performance.