LMCA

LMCA

LMCA (Language Model Conceptual Argumentation), created by the conceptual reasoning capabilities team at Redwood Research, measures how well models evaluate informal arguments in conceptual domains such as decision theory, philosophy, and AI safety — domains where the bottom-line answer is often contested, but the quality of individual arguments can still be judged. The dataset consists of 560 position texts and 1,461 arguments against them, each rated by at least one domain expert, for a total of 2,140 expert ratings. The dataset itself is not public; access can be requested from Redwood Research.

Methodology

We source results from the public Conceptual Reasoning Index leaderboard. Our chart reports the LMCA score: the Pearson correlation between a model’s “overall” ratings of the arguments and the average of the human experts’ “overall” ratings, multiplied by 100. The score’s 95% confidence interval is kept in the data export.

Arguments are rated by hand against a detailed rubric on seven dimensions, including centrality, strength, correctness, and an “overall” rating. Large disagreements between raters are discussed and revised, and a subset of around 50 arguments was independently rated by four to six people to estimate inter-rater agreement. On this basis, the benchmark’s authors estimate ceiling performance at a correlation of around 0.85, i.e. a score of 85. Models are run with each lab’s default sampling parameters and the maximum available token limit, reasoning effort, and thinking settings.

For full details, see the LMCA paper.