Skip to content

Evaluation

Human feedback evaluations (Claude 3.5 Sonnet)

human preference win rates vs prior Claude models on common tasks and expert domains

Core capabilities Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.5 Sonnet
Anthropic

partial detail 17 method fields published

See how this test was run