Skip to content

Evaluation

Vision capabilities benchmark suite (upgraded Claude 3.5 Sonnet)

general VQA, visual math reasoning, science diagrams, chart interpretation, document analysis, MMMU

Multimodal Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet
Anthropic

rich detail 18 method fields published

See how this test was run