APEX-Agents was created by Mercor to assess whether AI agents can execute long-horizon, cross-application tasks in professional services. The benchmark contains 480 tasks – 160 each in investment banking, management consulting, and corporate law – set within 33 simulated work environments, or “worlds”, each averaging over 160 files. The worlds and tasks were built by industry professionals with five to ten years of experience at top-tier firms, with each world constructed around a realistic project scenario, such as a week-long consulting engagement. Agents complete the tasks using the software a human professional would use – documents, spreadsheets, presentations, PDFs, email, chat, calendar, a file system, and code execution – with web search disabled for reproducibility.
Evaluation is rubric-based: each task has an expert-written rubric of one to ten binary pass/fail criteria (around four on average), and a judge LM grades each criterion against the agent’s output and produced artifacts. A task is passed only if every criterion is met, and the headline metric is the pass@1 rate. Tasks are estimated to take a human professional 1.8 hours on average, and require agents to plan over long horizons, navigate large file structures, and coordinate work across multiple applications. The benchmark is open source: the tasks and worlds are released under a CC-BY license, and the underlying environment infrastructure, Archipelago, is available as an open-source repository.
Have a question? Noticed something wrong? Let us know.
A benchmark evaluating whether AI agents can execute long-horizon professional tasks in investment banking, management consulting, and corporate law within realistic simulated work environments.