Evaluation
TAU-bench
Agentic tool use (tau-bench retail)
Agentic & computer use
Anthropic, OpenAI
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet
Anthropic
partial detail
16 method fields published
See how this test was run
Claude 3.7 Sonnet
Anthropic
rich detail
21 method fields published
See how this test was run
o3, o4-mini
OpenAI
rich detail
12 method fields published
See how this test was run
Claude Opus 4, Claude Sonnet 4
Anthropic
rich detail
23 method fields published
See how this test was run