DTBench (Decision Theory Benchmark), created by the conceptual reasoning capabilities team at Redwood Research, tests models’ reasoning about the decision theory of Newcomb-like problems — situations in which an agent’s decisions are informative about the behavior of similar agents, such as copies of itself. The dataset contains 537 handcrafted multiple-choice questions, the vast majority original and written by decision theorist Caspar Oesterheld: 407 capability questions with objectively correct answers, and a further 130 questions that probe models’ decision-theoretic attitudes, such as whether they recommend one-boxing in Newcomb’s problem.
We source results from the public Conceptual Reasoning Index leaderboard. Our chart reports accuracy on the 407 capability questions; the attitude questions have no objectively correct answers and are excluded from scoring. Random guessing scores 40%, and because every answer was independently checked by a second domain expert, the benchmark’s authors believe ceiling performance is at or close to 100%.
Models are run with each lab’s default sampling parameters and the maximum available token limit, reasoning effort, and thinking settings. The share of a model’s answers on the attitudes section that matches the recommendation of evidential decision theory is kept in the data export.
For full details, see the DTBench paper and code.
A benchmark of handcrafted multiple-choice questions testing models' understanding of the decision theory of Newcomb-like problems.