Skip to content

Evaluation

Monitoring for concerning thought processes (9,833 prompts)

rates of concerning thinking categories (deception/manipulation, planning harmful actions, distress language) in extended thinking outputs

Reasoning & chain-of-thought monitoring Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

partial detail 20 method fields published

See how this test was run