
Benjamin Ou is a benchmarking researcher at Epoch AI, exploring how challenging games can be used to evaluate AI capabilities. He previously worked at a private startup designing CPU pipelines for optimized floating-point units. He has a BA in physics and computer science from UC Berkeley.

Updates on EBR-bench, Epoch AI’s benchmark that tests models’ ability to learn from experience. GPT-6 Astra hit a perfect score on over half its EBR-bench attempts by exploiting one overpowered card, which we now ban. Plus multi-agent scaffold results.

Frontier models show no improvement across 30 playthroughs of the board game Earthborne Rangers, scoring far below expert humans. Epoch AI's EBR-bench probes whether AI can learn on the fly.