Skip to content

Evaluation

Pairwise Safety Comparison (o1)

Red teamer safety ratings of anonymized o1 vs GPT-4o responses

Red teaming & adversarial testing OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

o1
OpenAI

rich detail 10 method fields published

See how this test was run