SafetyEst. 2021

TruthfulQA

TruthfulQA measures whether language models generate truthful answers to questions where humans commonly hold misconceptions. The benchmark covers 817 questions across 38 categories including health, law, finance, and conspiracy theories. It specifically targets questions where models are incentivized to reproduce popular falsehoods rather than provide accurate but less common truths, making it a key safety benchmark.

Metrics

Truthfulness score (%) and informativeness score (%)

Created By

Stephanie Lin et al. (Oxford)

Top Model Scores

RankModelScoreDate
1Claude Opus 4.682.4%2026-02
2GPT-5.280.9%2026-03
3Gemini 3 Ultra79.5%2026-01
4Grok 476.8%2026-02
5Llama 4 405B73.2%2026-01