Gemini 1.0
Google DeepMind — Gemini
partial detail 5 method fields published
Evaluation
Holistic harness of more than 50 benchmarks in six capabilities: factuality (open/closed-book retrieval and QA), long-context (summarization, retrieval, QA), math/science, reasoning (arithmetic, scientific, commonsense), multilingual (translation, summarization, reasoning), summarization
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 5 method fields published