Skip to content

Evaluation

Alignment faking reasoning evaluation

alignment faking rate and compliance gap in the 'helpful-only' setting where the model learns it will be trained to respond to all user requests

Model behavior & alignment Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 17 method fields published

See how this test was run