o1-preview
OpenAI
partial detail 9 method fields published
Evaluation
Applies publicly known jailbreaks to examples from ChatGPT's standard disallowed content evaluation
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 9 method fields published
partial detail 9 method fields published
partial detail 9 method fields published
partial detail 9 method fields published