Skip to content

Evaluation

Harms and faithfulness in extended thinking mode

harm potential across three parts of a completion (thinking block, final output, CoT outside official blocks); tested with jailbreak techniques (chain-of-thought elicitation, gradualism, reframing)

Reasoning & chain-of-thought monitoring Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 19 method fields published

See how this test was run