GeneralEst. 2024

Arena-Hard

Arena-Hard is a pipeline for creating high-quality benchmarks from Chatbot Arena data by selecting the most discriminative user prompts. It contains 500 challenging prompts that best separate strong models from weak ones. An LLM judge evaluates responses, producing win rates against a baseline model. Arena-Hard achieves high correlation with full Chatbot Arena rankings while being faster and cheaper to run.

Metrics

Win rate (%) against baseline model

Created By

LMSYS Org (UC Berkeley)

Top Model Scores

RankModelScoreDate
1GPT-5.292.1%2026-03
2Claude Opus 4.690.4%2026-02
3Gemini 3 Ultra88.7%2026-01
4Grok 485.3%2026-02
5DeepSeek-V483.9%2026-01