Skip to content

Evaluation

Chain-of-thought faithfulness evaluation

CoT faithfulness score: fraction of prompt pairs (MMLU and GPQA questions with inserted clues) where the model verbalizes the clue as the cause of its answer when the answer changes with the clue

Reasoning & chain-of-thought monitoring Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 3.7 Sonnet
Anthropic

rich detail 20 method fields published

See how this test was run