Skip to content

Evaluation

Automated red teaming for security (indirect prompt injection)

Indirect prompt injection scenario: attacker hides malicious instructions in retrieved email data to make Gemini invoke a send-email function exfiltrating sensitive info; attacks: Actor Critic, Beam Search, Tree of Attacks with Pruning (TAP)

Prompt injection resistance Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 2.5
Google DeepMind — Gemini

rich detail 10 method fields published

See how this test was run