Skip to content

Evaluation

Video recognition benchmarks (PerceptionTest, NExT-QA, TVQA, ActivityNet-QA)

Video adapter for Llama 3 (8B/70B) on PerceptionTest (11.6K test QA pairs), NExT-QA (1K videos, 9K questions, WUPS scoring), TVQA (15K+ validation QA pairs), ActivityNet-QA (8K test QA pairs, GPT-3.5 API correctness evaluation); uniform frame sampling with text prompt

Multimodal Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 3
Meta

partial detail 14 method fields published

See how this test was run