LanguageEst. 2019

HellaSwag

HellaSwag tests commonsense natural language inference by asking models to predict the most plausible continuation of a given scenario. The dataset uses adversarial filtering to generate wrong answers that are superficially plausible but logically incorrect. Tasks span everyday activities like cooking, sports, and social interactions, testing whether models truly understand sequential reasoning about real-world events.

Metrics

Accuracy (%) on sentence completion tasks

Created By

Rowan Zellers et al. (AI2/UW)

Top Model Scores

RankModelScoreDate
1GPT-5.297.6%2026-03
2Claude Opus 4.697.2%2026-02
3Gemini 3 Ultra96.9%2026-01
4Grok 496.3%2026-02
5Llama 4 405B95.0%2026-01