Skip to content

Evaluation

Baseline assurance evaluations (safety policy violation rate)

violation rate for safety policies (child sexual abuse and exploitation, PII, hate speech/harassment, dangerous/malicious content, sexually explicit content, medical advice against consensus) using a large number of synthetic adversarial user queries with human raters labelling answers as policy-violating or not

Harmful content & safety Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemma 2
Google DeepMind — Gemma

rich detail 22 method fields published

See how this test was run
Gemma 3
Google DeepMind — Gemma

rich detail 21 method fields published

See how this test was run