← Back to Explorer
Evaluation 668 of 723
Evaluation
Training and development safety evaluations (automated)
Automated safety evaluations during training/development: content policy violation rates, tone of refusals, instruction following; scores as absolute percentage change vs baseline model
Harmful content & safety
Google DeepMind — Gemini
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
Gemini 2.0 Flash Model Card
Dec 2024
Google DeepMind — Gemini
rich detail
10 method fields published
See how this test was run
Gemini 2.0 Flash - Model Card - Googleapis.com · Dec 2024
Gemini 2.0 Flash - Model Card - Googleapis.com · Dec 2024
Gemini 2.0 Flash-Lite Model Card
Dec 2024
Google DeepMind — Gemini
rich detail
10 method fields published
See how this test was run
Gemini 2.0 Flash - Model Card - Googleapis.com · Dec 2024
Gemini 2.0 Flash - Model Card - Googleapis.com · Dec 2024
Gemini 2.5 Flash Model Card
Mar 2025
Google DeepMind — Gemini
rich detail
11 method fields published
See how this test was run
Gemini 2.5 Pro - Model Card · Mar 2025
Gemini 2.5 Pro - Model Card · Mar 2025
Gemini 2.5 Flash-Lite Model Card
Mar 2025
Google DeepMind — Gemini
rich detail
11 method fields published
See how this test was run
Gemini 2.5 Pro - Model Card · Mar 2025
Gemini 2.5 Pro - Model Card · Mar 2025
Gemini 2.5 Pro Model Card
Mar 2025
Google DeepMind — Gemini
rich detail
11 method fields published
See how this test was run
Gemini 2.5 Pro - Model Card · Mar 2025
Gemini 2.5 Pro - Model Card · Mar 2025