Operator
OpenAI
partial detail 8 method fields published
Evaluation
Operator-specific refusal evaluation: activities that cause or intend to cause physical harm, injury, or destruction, and non-violent wrongdoing and crime (categories: harassment/threatening, sexual/minors, sexual/exploitative, extremist/propaganda, hate/threatening, hate, illicit/violent, illicit/non-violent, personal-data/sensitive, regulated-advice, self-harm/instructions, self-harm/intent)
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 8 method fields published