Skip to content

Evaluation

Standard Refusal Evaluation

Standard refusal evaluations across disallowed content categories for the codex-1 model (categories: harassment/threatening, sexual/exploitative, sexual/minors, extremist/propaganda, hate, hate/threatening, illicit/non-violent, illicit/violent, personal-data/semi-restricted, personal-data/restricted, regulated-advice, self-harm/intent, self-harm/instructions)

Harmful content & safety OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

o1-preview
OpenAI

rich detail 10 method fields published

See how this test was run
o1
OpenAI

partial detail 10 method fields published

See how this test was run
Operator
OpenAI

rich detail 10 method fields published

See how this test was run
o3-mini
OpenAI

partial detail 10 method fields published

See how this test was run
Deep Research
OpenAI

rich detail 11 method fields published

See how this test was run
GPT-4.5
OpenAI

rich detail 11 method fields published

See how this test was run
o3/o4-mini
OpenAI

rich detail 10 method fields published

See how this test was run
Codex
OpenAI

rich detail 8 method fields published

See how this test was run
o3 Operator
OpenAI

rich detail 10 method fields published

See how this test was run