Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| VQAv2 (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Natural image QA benchmark; Gemini 1.5 maintains competitive performance |
| VQAv2 (visual question answering) | Google DeepMind — Gemini | Gemini 1.0 | High-level object recognition via captioning/QA; zero-shot evaluation, greedy sampling, no external OCR |
| Vulnerability identification and exploitation (CyberSecEval 2 CTF) | Meta | Llama 3 | Ability to identify and exploit vulnerabilities using CyberSecEval 2's capture-the-flag test challenges |
| Web of Lies | Google DeepMind — Gemma | Gemma 2 | ability to shift participant beliefs: participants engage in short conversations about simple factual questions (e.g., 'Which country had tomatoes first - Italy or Mexico?'); the … |
| Wet lab protocol evaluations (ProtocolQA and Cloning Scenarios) (o1-preview) | OpenAI | o1-preview | Wet lab protocol capability (ProtocolQA dataset and Cloning Scenarios) |
| Wide versus deep ablation | Google DeepMind — Gemma | Gemma 2 | deeper vs wider 9B network at the same parameter count |
| WidgetCaps evaluation methodology note | Google DeepMind — Gemma | PaliGemma 1 | issues found in at least three previous works' evaluations of WidgetCaps, rendering numerical comparisons invalid |
| WikiLingua (multilingual summarization) | Google DeepMind — Gemini | Gemini 1.0 | Multilingual summarization benchmark |
| Winobias | Google DeepMind — Gemini | Gemini 1.0 | Winobias benchmark measuring representational harm (bias or toxicity) in text-to-text outputs |
| WinoGender | Google DeepMind — Gemini, Meta | Gemini 1.0, Llama 1 | Winogender benchmark measuring representational harm (bias or toxicity) in text-to-text outputs Co-reference resolution accuracy on Winograd-schema sentences by pronoun type … |
| WinoGrande | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | commonsense reasoning (Winograd schemas) |
| WMDP Biology | OpenAI | Deep Research, GPT-4.5 | Hazardous biological knowledge (MCQ) from Weapons of Mass Destruction Proxy benchmark Hazardous biological knowledge (MCQ) from WMDP benchmark |
| WMT 23 (machine translation benchmark) | Google DeepMind — Gemini | Gemini 1.0 | WMT 23 translation benchmark across the entire set of language pairs in few-shot setting; out-of-English translation quality |
| WMT23 (machine translation; Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | WMT23 translation benchmark (8 languages, 14 pairs) constructed after training data cutoff |
| World knowledge (NaturalQuestions, TriviaQA) | Meta | Llama 2 | Average 5-shot performance on NaturalQuestions and TriviaQA for Llama 1 and Llama 2 base models |
| XLSum (multilingual summarization) | Google DeepMind — Gemini | Gemini 1.0 | Multilingual summarization benchmark |
| XM-3600 (Crossmodal3600 multilingual captioning) | Google DeepMind — Gemini | Gemini 1.0 | Multilingual image captioning on a subset of languages of the XM-3600 benchmark, 4-shot, Flamingo evaluation protocol, no fine-tuning |
| XSTest | OpenAI | GPT-4.5, o1, o1-preview, o3-mini | Benign prompts from XSTest testing over-refusal edge cases (categories: Definitions, Figurative Language, Historical Events, Homonyms, Discr: Nonsense group, Discr: Nonsense … |
| YouCook2 (cooking video captioning; Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Cooking-focused video captioning benchmark; CIDER metric |
| YouCook2 (video captioning) | Google DeepMind — Gemini | Gemini 1.0 | Few-shot video captioning benchmark; CIDER metric |
| Zero-shot generalization to 3D renders (Objaverse) | Google DeepMind — Gemma | PaliGemma 1 | generalization to 3D renders from Objaverse without fine-tuning, although never explicitly trained for it |
| Zero-shot tool use / function calling (Nexus, API-Bank, Gorilla API-Bench, BFCL) | Meta | Llama 3 | Zero-shot tool use (function calling) on Nexus, API-Bank, Gorilla API-Bench, and Berkeley Function Calling Leaderboard (BFCL); plus training-stage evaluation on a set of unseen … |
| ZeroSCROLLS | Meta | Llama 3 | ZeroSCROLLS (Shaham et al. 2023) validation set (ground truth not publicly available) |