LanguageEst. 2019

WinoGrande

WinoGrande is a large-scale commonsense reasoning benchmark inspired by the original Winograd Schema Challenge. It presents fill-in-the-blank problems that require understanding context, physical commonsense, and social reasoning to resolve ambiguous pronoun references. The dataset contains 44,000 problems adversarially constructed to minimize annotation artifacts, making it a robust test of genuine commonsense understanding.

Metrics

Accuracy (%) on pronoun resolution tasks

Created By

Keisuke Sakaguchi et al. (AI2)

Top Model Scores

RankModelScoreDate
1GPT-5.295.3%2026-03
2Claude Opus 4.694.9%2026-02
3Gemini 3 Ultra94.5%2026-01
4Grok 493.8%2026-02
5Llama 4 405B92.1%2026-01