Skip to content

Evaluation

Multimodal post-training vision evaluation (SxS and benchmark comparisons)

SxS evaluation of text-only quality (+0.01 for a Pro model trained with image-text data) and image-understanding tasks (+0.223 for SFT+RLHF vs SFT alone); plus SFT impact of API Vision models on standard benchmarks (InfographicVQA, AI2D, VQAv2)

Multimodal Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 1.0
Google DeepMind — Gemini

rich detail 10 method fields published

See how this test was run