MuSR
MuSR (Multi-Step Reasoning) tests language models on complex problems that require chaining multiple reasoning steps together. The benchmark includes murder mystery puzzles, team allocation problems, and object placement tasks that demand tracking multiple entities, applying logical rules, and maintaining consistency across 5-10 reasoning steps. It exposes weaknesses in models that appear strong on simpler benchmarks.
Metrics
Accuracy (%) on multi-step reasoning tasks
Created By
Zayne Sprague et al.
Paper
View paper →Website
Visit website →Top Model Scores
| Rank | Model | Score | Date |
|---|---|---|---|
| 1 | GPT-5.2 | 71.2% | 2026-03 |
| 2 | Claude Opus 4.6 | 69.8% | 2026-02 |
| 3 | Gemini 3 Ultra | 67.4% | 2026-01 |
| 4 | Grok 4 | 63.1% | 2026-02 |
| 5 | DeepSeek-V4 | 60.5% | 2026-01 |
Related Reasoning Benchmarks
ARC (AI2 Reasoning Challenge)
The AI2 Reasoning Challenge contains 7,787 genuine grade-school science questions, split into Easy and Challenge sets. The Challenge set contains only questions that are answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm.
Top: GPT-5.2 — 98.2%
GPQA (Diamond)
Graduate-Level Google-Proof Q&A (GPQA) Diamond is a challenging benchmark of expert-level questions in biology, physics, and chemistry. Questions are designed to be answerable by domain experts but extremely difficult for non-experts, even with web search.
Top: GPT-5.2 — 94.7%
BBH (BIG-Bench Hard)
BIG-Bench Hard is a suite of 23 challenging tasks from the BIG-Bench benchmark where language models previously performed below average human raters. Tasks include boolean expressions, causal judgement, date understanding, disambiguation, and more.
Top: GPT-5.2 — 95.3%
DROP
Discrete Reasoning Over Paragraphs tests reading comprehension that requires discrete reasoning steps including addition, subtraction, counting, sorting, and other operations over text passages.
Top: GPT-5.2 — 93.1