Skip to content

Evaluation

Automated evaluation for risky advice

not_unsafe on risky advice prompts

Harmful content & safety OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

deep research
OpenAI

partial detail 10 method fields published

See how this test was run