LanguageEst. 2023

IFEval

IFEval (Instruction Following Evaluation) tests whether language models can precisely follow specific formatting and content instructions. Tasks include writing responses with exact word counts, including or excluding specific phrases, formatting output as JSON or bullet points, and following complex multi-constraint instructions. It measures the practical reliability of models when users need outputs to conform to exact specifications.

Metrics

Strict accuracy (%) on verifiable instruction-following tasks

Created By

Jeffrey Zhou et al. (Google)

Top Model Scores

RankModelScoreDate
1GPT-5.289.7%2026-03
2Claude Opus 4.688.3%2026-02
3Gemini 3 Ultra86.9%2026-01
4Grok 484.5%2026-02
5Llama 4 405B81.2%2026-01