IFEval
IFEval (Instruction Following Evaluation) tests whether language models can precisely follow specific formatting and content instructions. Tasks include writing responses with exact word counts, including or excluding specific phrases, formatting output as JSON or bullet points, and following complex multi-constraint instructions. It measures the practical reliability of models when users need outputs to conform to exact specifications.
Metrics
Strict accuracy (%) on verifiable instruction-following tasks
Created By
Jeffrey Zhou et al. (Google)
Paper
View paper →Website
Visit website →Top Model Scores
| Rank | Model | Score | Date |
|---|---|---|---|
| 1 | GPT-5.2 | 89.7% | 2026-03 |
| 2 | Claude Opus 4.6 | 88.3% | 2026-02 |
| 3 | Gemini 3 Ultra | 86.9% | 2026-01 |
| 4 | Grok 4 | 84.5% | 2026-02 |
| 5 | Llama 4 405B | 81.2% | 2026-01 |
Related Language Benchmarks
MMLU
Massive Multitask Language Understanding measures knowledge across 57 academic subjects including STEM, humanities, social sciences, and more. It tests both world knowledge and problem-solving ability at varying difficulty levels from elementary to professional.
Top: GPT-5.2 — 92.4%
AlpacaEval 2.0
AlpacaEval 2.0 is an automatic evaluation benchmark that measures instruction-following ability. It uses a length-controlled win rate against a reference model, reducing length bias that affected the original version.
Top: Claude Opus 4.6 — 72.1%
WildBench
WildBench evaluates AI models on challenging real-world user queries collected from the wild. It focuses on complex, multi-constraint instructions that test practical model capabilities beyond academic benchmarks.
Top: Claude Opus 4.6 — 68.7%
SimpleQA
SimpleQA evaluates factual accuracy on straightforward, unambiguous factual questions with short, verifiable answers. It specifically tests whether models provide correct factual information vs. hallucinating plausible-sounding but incorrect answers.
Top: GPT-5.2 — 52.8%