Skip to content

Evaluation

CountBenchQA

VLM-ready version of the CountBench dataset introduced because TallyQA was found lacking in its ability to assess current VLM counting ability (skewed number distribution and varying image quality)

Multimodal Google DeepMind — Gemma

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

PaliGemma 1
Google DeepMind — Gemma

partial detail 15 method fields published

See how this test was run