Gemma 2
Google DeepMind — Gemma
rich detail 22 method fields published
Evaluation
violation rate for safety policies (child sexual abuse and exploitation, PII, hate speech/harassment, dangerous/malicious content, sexually explicit content, medical advice against consensus) using a large number of synthetic adversarial user queries with human raters labelling answers as policy-violating or not
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail 22 method fields published
rich detail 21 method fields published