WinoGrande
WinoGrande is a large-scale commonsense reasoning benchmark inspired by the original Winograd Schema Challenge. It presents fill-in-the-blank problems that require understanding context, physical commonsense, and social reasoning to resolve ambiguous pronoun references. The dataset contains 44,000 problems adversarially constructed to minimize annotation artifacts, making it a robust test of genuine commonsense understanding.
Metrics
Accuracy (%) on pronoun resolution tasks
Created By
Keisuke Sakaguchi et al. (AI2)
Paper
View paper →Website
Visit website →Top Model Scores
| Rank | Model | Score | Date |
|---|---|---|---|
| 1 | GPT-5.2 | 95.3% | 2026-03 |
| 2 | Claude Opus 4.6 | 94.9% | 2026-02 |
| 3 | Gemini 3 Ultra | 94.5% | 2026-01 |
| 4 | Grok 4 | 93.8% | 2026-02 |
| 5 | Llama 4 405B | 92.1% | 2026-01 |
Related Language Benchmarks
MMLU
Massive Multitask Language Understanding measures knowledge across 57 academic subjects including STEM, humanities, social sciences, and more. It tests both world knowledge and problem-solving ability at varying difficulty levels from elementary to professional.
Top: GPT-5.2 — 92.4%
AlpacaEval 2.0
AlpacaEval 2.0 is an automatic evaluation benchmark that measures instruction-following ability. It uses a length-controlled win rate against a reference model, reducing length bias that affected the original version.
Top: Claude Opus 4.6 — 72.1%
WildBench
WildBench evaluates AI models on challenging real-world user queries collected from the wild. It focuses on complex, multi-constraint instructions that test practical model capabilities beyond academic benchmarks.
Top: Claude Opus 4.6 — 68.7%
SimpleQA
SimpleQA evaluates factual accuracy on straightforward, unambiguous factual questions with short, verifiable answers. It specifically tests whether models provide correct factual information vs. hallucinating plausible-sounding but incorrect answers.
Top: GPT-5.2 — 52.8%