Skip to content

Evaluation

o1 safety evaluation suite overview

Propensity to generate disallowed content, demographic fairness, hallucination, dangerous capabilities; external red teaming

Harmful content & safety OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

o1
OpenAI

partial detail 5 method fields published

See how this test was run