FrontierSWE (v2)

FrontierSWE (v2)

FrontierSWE evaluates coding agents on hard, ultra-long-horizon software engineering tasks collected from real-world technical domains. The current version, FrontierSWE v2 (September 2026), contains 34 tasks across implementation, performance engineering, scientific computing, visual reasoning, and AI research; examples include reimplementing git in Zig, decoding speech from MEG brain recordings, and training a medium-range weather forecasting model. Each task is run for up to 20 hours and scored from 0 to 100%, and models are run five times per task.

The leaderboard’s headline number is the mean score over the five trials, averaged across all tasks. The site also reports the best and worst of the five trials, per-domain scores, and the average cost and wall-clock time per task. All models are run at their maximum reasoning effort inside the benchmark’s own minimal agent harness, “proximus”.

Methodology

We source results from the public FrontierSWE leaderboard and show its default mean@5 score. The data export also carries the best and worst of the five trials, the implementation, performance and research sub-scores, and the leaderboard’s average cost and duration per task.