Skip to content

Evaluation

Long fine-grained caption generation (DOCCI)

fine-tuning on DOCCI (15k images with detailed human-annotated English descriptions, avg 7.1 sentences / 639 characters / 136 words); models selected by test-split perplexity; captions generated on the 100-image qual_dev split (max decoding length 192); human evaluations assessing whether each generated sentence is factually entailed by the image (four options: Entailment, Neutral, Contradiction, Nothing to assess); proportion of Non-Entailment Sentences (NES) used to select most factually accurate models

Multimodal Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

PaliGemma 2
Google DeepMind — Gemma

rich detail 31 method fields published

See how this test was run