Skip to content

Evaluation

Holistic capability harness (50+ benchmarks, six capabilities)

Holistic harness of more than 50 benchmarks in six capabilities: factuality (open/closed-book retrieval and QA), long-context (summarization, retrieval, QA), math/science, reasoning (arithmetic, scientific, commonsense), multilingual (translation, summarization, reasoning), summarization

General benchmark suites Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 1.0
Google DeepMind — Gemini

partial detail 5 method fields published

See how this test was run