Skip to content

Evaluation

IFEval

Instruction following with verifiable instructions

Core capabilities Anthropic, OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet
Anthropic

partial detail 14 method fields published

See how this test was run
GPT-4.1
OpenAI

partial detail 8 method fields published

See how this test was run