Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
VQAv2 (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Natural image QA benchmark; Gemini 1.5 maintains competitive performance
VQAv2 (visual question answering) Google DeepMind — Gemini Gemini 1.0 High-level object recognition via captioning/QA; zero-shot evaluation, greedy sampling, no external OCR
Vulnerability identification and exploitation (CyberSecEval 2 CTF) Meta Llama 3 Ability to identify and exploit vulnerabilities using CyberSecEval 2's capture-the-flag test challenges
Web of Lies Google DeepMind — Gemma Gemma 2 ability to shift participant beliefs: participants engage in short conversations about simple factual questions (e.g., 'Which country had tomatoes first - Italy or Mexico?'); the …
Wet lab protocol evaluations (ProtocolQA and Cloning Scenarios) (o1-preview) OpenAI o1-preview Wet lab protocol capability (ProtocolQA dataset and Cloning Scenarios)
Wide versus deep ablation Google DeepMind — Gemma Gemma 2 deeper vs wider 9B network at the same parameter count
WidgetCaps evaluation methodology note Google DeepMind — Gemma PaliGemma 1 issues found in at least three previous works' evaluations of WidgetCaps, rendering numerical comparisons invalid
WikiLingua (multilingual summarization) Google DeepMind — Gemini Gemini 1.0 Multilingual summarization benchmark
Winobias Google DeepMind — Gemini Gemini 1.0 Winobias benchmark measuring representational harm (bias or toxicity) in text-to-text outputs
WinoGender Google DeepMind — Gemini, Meta Gemini 1.0, Llama 1 Winogender benchmark measuring representational harm (bias or toxicity) in text-to-text outputs Co-reference resolution accuracy on Winograd-schema sentences by pronoun type …
WinoGrande Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku commonsense reasoning (Winograd schemas)
WMDP Biology OpenAI Deep Research, GPT-4.5 Hazardous biological knowledge (MCQ) from Weapons of Mass Destruction Proxy benchmark Hazardous biological knowledge (MCQ) from WMDP benchmark
WMT 23 (machine translation benchmark) Google DeepMind — Gemini Gemini 1.0 WMT 23 translation benchmark across the entire set of language pairs in few-shot setting; out-of-English translation quality
WMT23 (machine translation; Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 WMT23 translation benchmark (8 languages, 14 pairs) constructed after training data cutoff
World knowledge (NaturalQuestions, TriviaQA) Meta Llama 2 Average 5-shot performance on NaturalQuestions and TriviaQA for Llama 1 and Llama 2 base models
XLSum (multilingual summarization) Google DeepMind — Gemini Gemini 1.0 Multilingual summarization benchmark
XM-3600 (Crossmodal3600 multilingual captioning) Google DeepMind — Gemini Gemini 1.0 Multilingual image captioning on a subset of languages of the XM-3600 benchmark, 4-shot, Flamingo evaluation protocol, no fine-tuning
XSTest OpenAI GPT-4.5, o1, o1-preview, o3-mini Benign prompts from XSTest testing over-refusal edge cases (categories: Definitions, Figurative Language, Historical Events, Homonyms, Discr: Nonsense group, Discr: Nonsense …
YouCook2 (cooking video captioning; Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Cooking-focused video captioning benchmark; CIDER metric
YouCook2 (video captioning) Google DeepMind — Gemini Gemini 1.0 Few-shot video captioning benchmark; CIDER metric
Zero-shot generalization to 3D renders (Objaverse) Google DeepMind — Gemma PaliGemma 1 generalization to 3D renders from Objaverse without fine-tuning, although never explicitly trained for it
Zero-shot tool use / function calling (Nexus, API-Bank, Gorilla API-Bench, BFCL) Meta Llama 3 Zero-shot tool use (function calling) on Nexus, API-Bank, Gorilla API-Bench, and Berkeley Function Calling Leaderboard (BFCL); plus training-stage evaluation on a set of unseen …
ZeroSCROLLS Meta Llama 3 ZeroSCROLLS (Shaham et al. 2023) validation set (ground truth not publicly available)