Skip to content

Evaluation

HHH (Helpful, Honest, and Harmless) binary-choice evaluation

model's ability to select the more HHH output from two options on 438 binary-choice questions

Model behavior & alignment Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 2
Anthropic

rich detail 19 method fields published

See how this test was run