HellaSwag
HellaSwag tests commonsense natural language inference by asking models to predict the most plausible continuation of a given scenario. The dataset uses adversarial filtering to generate wrong answers that are superficially plausible but logically incorrect. Tasks span everyday activities like cooking, sports, and social interactions, testing whether models truly understand sequential reasoning about real-world events.
Metrics
Accuracy (%) on sentence completion tasks
Created By
Rowan Zellers et al. (AI2/UW)
Paper
View paper →Website
Visit website →Top Model Scores
| Rank | Model | Score | Date |
|---|---|---|---|
| 1 | GPT-5.2 | 97.6% | 2026-03 |
| 2 | Claude Opus 4.6 | 97.2% | 2026-02 |
| 3 | Gemini 3 Ultra | 96.9% | 2026-01 |
| 4 | Grok 4 | 96.3% | 2026-02 |
| 5 | Llama 4 405B | 95.0% | 2026-01 |
Related Language Benchmarks
MMLU
Massive Multitask Language Understanding measures knowledge across 57 academic subjects including STEM, humanities, social sciences, and more. It tests both world knowledge and problem-solving ability at varying difficulty levels from elementary to professional.
Top: GPT-5.2 — 92.4%
AlpacaEval 2.0
AlpacaEval 2.0 is an automatic evaluation benchmark that measures instruction-following ability. It uses a length-controlled win rate against a reference model, reducing length bias that affected the original version.
Top: Claude Opus 4.6 — 72.1%
WildBench
WildBench evaluates AI models on challenging real-world user queries collected from the wild. It focuses on complex, multi-constraint instructions that test practical model capabilities beyond academic benchmarks.
Top: Claude Opus 4.6 — 68.7%
SimpleQA
SimpleQA evaluates factual accuracy on straightforward, unambiguous factual questions with short, verifiable answers. It specifically tests whether models provide correct factual information vs. hallucinating plausible-sounding but incorrect answers.
Top: GPT-5.2 — 52.8%