Skip to content

Evaluation

TAU-bench

Agentic tool use (tau-bench retail)

Agentic & computer use Anthropic, OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet
Anthropic

partial detail 16 method fields published

See how this test was run
Claude 3.7 Sonnet
Anthropic

rich detail 21 method fields published

See how this test was run
o3, o4-mini
OpenAI

rich detail 12 method fields published

See how this test was run
Claude Opus 4, Claude Sonnet 4
Anthropic

rich detail 23 method fields published

See how this test was run