Skip to content

Evaluation

GPT-4 harmful content (disallowed content) evaluations

likelihood of generating content violating content policy (hate speech, self-harm advice, illicit advice)

Harmful content & safety OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

GPT-4
OpenAI

rich detail 10 method fields published

See how this test was run