VisionEst. 2020

DocVQA

DocVQA (Document Visual Question Answering) tests AI models on their ability to understand and answer questions about document images including invoices, letters, reports, forms, and tables. Models must perform optical character recognition, layout understanding, and reasoning over document structure to extract specific information. It is a critical benchmark for enterprise document processing and automation applications.

Metrics

ANLS score (Average Normalized Levenshtein Similarity, 0-1)

Created By

Minesh Mathew et al. (CVC Barcelona)

Top Model Scores

RankModelScoreDate
1GPT-5.20.9522026-03
2Gemini 3 Ultra0.9482026-01
3Claude Opus 4.60.9412026-02
4Grok 40.9232026-02
5Llama 4 405B0.9082026-01