Skip to content

Evaluation

Video understanding benchmark suite

Video understanding benchmarks (Table 6): string-match accuracy for multiple-choice VideoQA, LLM-based accuracy for open-ended VideoQA, R1@0.5 for moment retrieval, CIDEr for captioning; vs GPT 4.1 under comparable conditions

Multimodal Google DeepMind — Gemini

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Gemini 2.5
Google DeepMind — Gemini

rich detail 17 method fields published

See how this test was run