Skip to content

Evaluation

Complex prompts instruction-following evaluation (internal)

Fine-grained evaluation of complex prompts with multiple instructions: per-instruction accuracy and full-response accuracy on an internal dataset of prompts with varying complexity

Core capabilities Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 1.0
Google DeepMind — Gemini

rich detail 10 method fields published

See how this test was run