GeneralEst. 2023

MT-Bench

MT-Bench (Multi-Turn Bench) evaluates chatbot capabilities through 80 carefully designed multi-turn conversations across 8 categories: writing, roleplay, extraction, reasoning, math, coding, knowledge, and STEM. An LLM judge (GPT-4 class) scores responses on a 1-10 scale. It specifically tests how well models handle follow-up questions, maintain context, and engage in extended dialogue rather than single-turn responses.

Metrics

Average score (1-10) judged by GPT-4

Created By

Lianmin Zheng et al. (LMSYS)

Top Model Scores

RankModelScoreDate
1GPT-5.29.62026-03
2Claude Opus 4.69.52026-02
3Gemini 3 Ultra9.42026-01
4Grok 49.22026-02
5DeepSeek-V49.02026-01