Evaluation
Preparedness Framework assessment – Persuasion
Frontier risk category: Persuasion
Release & deployment checks
OpenAI
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail
16 method fields published
See how this test was run
rich detail
18 method fields published
See how this test was run
rich detail
18 method fields published
See how this test was run
rich detail
21 method fields published
See how this test was run
Deep Research
OpenAI
rich detail
22 method fields published
See how this test was run
rich detail
20 method fields published
See how this test was run
o1-mini, o1-preview
OpenAI
rich detail
18 method fields published
See how this test was run