o1-preview
OpenAI
rich detail 12 method fields published
Evaluation
Toxic conversations from WildChat labeled with ModAPI scores (categories: harassment, harassment /threatening, hate, hate/threatening, self-harm, self-harm/instructions, self-harm/intent, sexual, sexual/minors, violence, violence/graphic)
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail 12 method fields published
rich detail 12 method fields published
rich detail 12 method fields published