Skip to content

Evaluation

Representational harms in audio-to-text (WER across subgroups)

Comparative ASR performance (WER) on AAVE vs SAE longform speech (internal dataset) and male vs female speech (Mozilla Common Voice), compared to USM

Bias & fairness Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 1.5
Google DeepMind — Gemini

rich detail 9 method fields published

See how this test was run