Skip to content

Evaluation

Assurance evaluations (baseline, for release decision-making)

Baseline assurance evaluations conducted for model release decision-making: model behavior within Google content policies and modality-specific risk areas; held-out prompt sets

Release & deployment checks Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 2.0 Flash Model Card
Google DeepMind — Gemini

rich detail 9 method fields published

See how this test was run
Gemini 2.0 Flash-Lite Model Card
Google DeepMind — Gemini

rich detail 9 method fields published

See how this test was run
Gemini 2.5 Flash Model Card
Google DeepMind — Gemini

rich detail 9 method fields published

See how this test was run
Gemini 2.5 Flash-Lite Model Card
Google DeepMind — Gemini

rich detail 9 method fields published

See how this test was run
Gemini 2.5 Pro Model Card
Google DeepMind — Gemini

rich detail 9 method fields published

See how this test was run