Back to benchmarks

WeirdML (v3)

An agentic benchmark of 11 hand-made machine learning and data analysis tasks with unfamiliar data, unspecified goals, and limited feedback.

WeirdML (v3)

WeirdML, created by Håvard Tveit Ihle at the Norwegian Defence Research Establishment (NDRE), is a benchmark of complex, hand-made machine learning tasks. Its third version, WeirdML v3 (September 2026), is agentic: it contains 11 tasks, each of which requires the agent to explore and understand unfamiliar data, develop machine learning and data analysis pipelines, and produce appropriate results despite limited data, unspecified goals, or very limited feedback. Models run inside coding agents such as Codex CLI, Claude Code, and Gemini CLI.

WeirdML v3 uses a different task set from WeirdML v2, and scores on the two versions are not comparable. We funded most of the API costs for the v3 evaluations, and METR and NDRE funded the rest.

Methodology

We source results from the public WeirdML leaderboard. Our chart reports the leaderboard’s official score, which rewards token efficiency as well as final performance. The leaderboard’s Final Best Score (shown here as “Max performance”) and mean API cost per run are available in the metric dropdown, and the tooltip shows the score’s 95% confidence interval, harness version, reasoning effort, and number of runs.

For each model, the leaderboard plots its average best-so-far effective score across all 11 tasks against the number of tokens used, on a logarithmic token axis. The official score is 80% of the normalized area under this curve plus 20% of the curve’s final value; only the interval from 500k to 50M tokens contributes to the area. The curve averages runs within each configuration and then weights each of the 11 tasks equally, with the hinted and hintless versions of a task each receiving half of that task’s weight. Effective scores are normalized and include hint penalties, so they are not raw accuracies. The Final Best Score is the equally weighted mean of each task’s final effective score. The leaderboard only includes models with at least one valid run in every configuration.