Skip to content

Evaluation

Instruction Hierarchy Evaluation - Conflicts Between Message Types

System vs user message conflicts (GPT-4.5 uses two message classifications)

Prompt injection resistance OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

o1
OpenAI

rich detail 12 method fields published

See how this test was run
o3-mini
OpenAI

rich detail 12 method fields published

See how this test was run
GPT-4.5
OpenAI

rich detail 12 method fields published

See how this test was run
o3/o4-mini
OpenAI

rich detail 12 method fields published

See how this test was run