Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| Text needle-in-a-haystack retrieval (long-context) | Google DeepMind — Gemini | Gemini 1.5 | Needle-in-a-haystack evaluation (Kamradt 2023): retrieve a 'magic number' needle inserted at varying depths in Paul Graham essay haystack up to 1M tokens, extended to 10M … |
| Text-based prompt injection benchmark | Meta | Llama 3 | Success rate of text-based prompt injection attacks against Llama 3 (Figure 22 compares Llama 3, GPT-4 Turbo, Gemini Pro, Mixtral) |
| Text-based reinforcement learning task | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to develop scaffolding (e.g. ReACT, Tree of Thought) to significantly enhance a weaker model's performance on a text-based RL task (proxy for self-optimization) |
| Text-to-text content policy violation evaluation | Google DeepMind — Gemini | Gemini 1.0 | Adherence to content safety policies on adversarial text prompts (12 languages, varied use cases) |
| Text-to-text content policy violation evaluation (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Development evaluation of text-to-text content policy violations using prompts across content policy areas and applications; automatic classification of violative vs non-violative … |
| Text-to-text content safety human evaluation | Google DeepMind — Gemma | Gemma 1 | human evaluation on prompts covering safety policies: child sexual abuse and exploitation, harassment, violence and gore, and hate speech; part of structured evaluations and … |
| Text-to-text helpfulness evaluation (SxS) | Google DeepMind — Gemini | Gemini 1.5 | Side-by-side (SxS) quality metric comparing Gemini 1.5 Pro/Flash vs Gemini 1.0 Ultra on prompts where nuanced answers avoid policy violations; annotators also rate instruction … |
| TextVQA (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | TextVQA benchmark focusing on OCR in natural images |
| TextVQA (text in images QA) | Google DeepMind — Gemini | Gemini 1.0 | Fine-grained transcription tasks requiring recognition of low-level details; zero-shot, pixel-only, no external OCR |
| Third party assessments (METR and Apollo Research) | OpenAI | GPT-4o | Key risks from general autonomous capabilities (validated by METR and Apollo Research) |
| Third Party Assessments (o3/o4-mini) | OpenAI | o3/o4-mini | Frontier risks related to autonomous capabilities, deception, and cybersecurity |
| Third-party pre-deployment testing (US AISI and UK AISI) | Anthropic | Claude 3.7 Sonnet | US AISI and UK AISI pre-deployment testing across RSP framework domains under voluntary Memorandums of Understanding |
| Third-party pre-deployment testing (US AISI and UK AISI) for Claude Opus 4 | Anthropic | Claude Opus 4, Claude Sonnet 4 | joint pre-deployment testing of Claude Opus 4 by US AISI and UK AISI with independent assessments in CBRN, cyber, and autonomy domains |
| Time series forecasting task | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to build models that match or exceed expert implementations on a regression/time-series-forecasting problem with SOTA benchmarks |
| To resize or to window ablation | Google DeepMind — Gemma | PaliGemma 1 | windowing alternative: 448px image cut into four pieces passed through SigLIP separately, embeddings concatenated, with optional window-ID position embeddings |
| Tool usage safety evaluation (search use case) | Meta | Llama 3 | Violation and false refusal rates for the search tool use case; safety of tool usage calls and integration |
| Tool use human evaluations (code execution, plot generation, file upload) | Meta | Llama 3 | Human preference evaluation of tool use capabilities focused on code execution tasks, using 2,000 user prompts (code execution without plotting/file uploads, plot generation, file … |
| Training and development safety evaluations (automated) | Google DeepMind — Gemini | Gemini 2.0 Flash Model Card, Gemini 2.0 Flash-Lite Model Card, Gemini 2.5 Flash Model Card, Gemini 2.5 Flash-Lite Model Card, Gemini 2.5 Pro Model Card | Automated safety evaluations during training/development: content policy violation rates, tone of refusals, instruction following; scores as absolute percentage change vs baseline … |
| Transfer evaluation on 30+ academic benchmarks (fine-tuning) | Google DeepMind — Gemma | PaliGemma 1 | fine-tuning of PaliGemma pretrained models on more than 30 academic benchmarks (captioning, VQA, referring segmentation, video) at multiple resolutions; none of these tasks or … |
| Transfer hyper-parameter sensitivity study | Google DeepMind — Gemma | PaliGemma 1 | all tasks run at 224px with a single simplified hyper-parameter setup (lr=1e-5, bs=256, no dropout, no label smoothing, no weight decay, nothing frozen); only epochs per task … |
| Transfer repeatability (variance across reruns) | Google DeepMind — Gemma | PaliGemma 1 | standard deviation across 5 transfer reruns using the best hyper-parameter (per task), and standard deviation of transferring from three Stage1 reruns, to quantify repeatability |
| Transfer with limited examples study | Google DeepMind — Gemma | PaliGemma 1 | fine-tuning PaliGemma with limited numbers of examples (64, 256, 1024, 4096), sweeping learning rate, epochs and batch size, with 5 seeds per setting; relative regret vs … |
| Trip Planning (itinerary planning) | Google DeepMind — Gemini | Gemini 1.5 | Task focusing on planning a trip itinerary under constraints with a single solution; order of visiting N cities |
| TriviaQA | Anthropic | Claude 2 | reading comprehension |
| Trust & Safety model red-teaming (14 policy areas, 6 languages) | Anthropic | Claude 3.5 Haiku, Claude 3.5 Sonnet (New) | harm rates of responses across policy areas (Elections Integrity, Child Safety, Cyber Attacks, Hate & Discrimination, Violent Extremism, etc.) in English, Arabic, Spanish, Hindi … |
| Truthfulness, toxicity and bias of fine-tuned Llama 2-Chat | Meta | Llama 2 | TruthfulQA (truthful and informative %), ToxiGen (toxic generation %), and BOLD (sentiment by demographic group) on fine-tuned Llama 2-Chat vs pretrained Llama 2 and compared … |
| TruthfulQA | Anthropic, Meta, OpenAI | Claude 2, Llama 1, gpt-4 | Ability to separate fact from adversarially-selected incorrect statements whether models output accurate and truthful responses to questions designed to elicit popular falsehoods … |
| Underrepresented Languages evaluation | OpenAI | GPT-4o | reading comprehension and reasoning across underrepresented languages; gap vs English |
| Unstructured Multimodal Data Analytics Task | Google DeepMind — Gemini | Gemini 1.5 | Image structuralization task: extract information from 1024 images into a structured data sheet; long-context multimodal analysis |
| Uplift testing for chemical and biological weapons (CBRNE planning) | Meta | Llama 3 | Six-hour uplift scenarios where teams of two generate fictitious operational plans for biological or chemical attacks covering CBRNE planning stages (agent acquisition … |
| Using Gemma 2 instead of Gemma 1 | Google DeepMind — Gemma | PaliGemma 2 | PaliGemma 2 vs PaliGemma transfer scores at the same resolution and model size (3B), averaged over the 30+ academic benchmarks |
| USMLE | Anthropic | Claude 2 | medical licensing examination performance |
| V* Benchmark (high-resolution image detail QA) | Google DeepMind — Gemini | Gemini 1.5 | V* Benchmark (Wu and Xie, 2023): questions about attributes and spatial relations of very small objects on high-resolution images (avg 2246x1582) from SA-1B |
| VATEX (video captioning) | Google DeepMind — Gemini | Gemini 1.0 | Few-shot video captioning benchmark; CIDER metric; 16 frames sampled |
| VATEX (video captioning; Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Video captioning benchmark; CIDER metric |
| VATEX ZH (Chinese video captioning) | Google DeepMind — Gemini | Gemini 1.5 | Chinese variant of VATEX; CIDER metric |
| Verbatim Memorization | Google DeepMind — Gemma | Gemma 1, Gemma 2 | discoverable memorization (Nasr et al. 2023): 10,000 documents sampled per corpus, first 50 tokens used as prompt, text classified memorized if the subsequent 50 generated tokens … |
| Video needle-in-a-haystack retrieval (long-context) | Google DeepMind — Gemini | Gemini 1.5 | Cross-modal needle-in-a-haystack: retrieve 'the secret word' text overlaid on a random frame of a 10.5-hour video (7 copies of AlphaGo documentary, 37,994 frames, 9.9M tokens) |
| Video recognition benchmarks (PerceptionTest, NExT-QA, TVQA, ActivityNet-QA) | Meta | Llama 3 | Video adapter for Llama 3 (8B/70B) on PerceptionTest (11.6K test QA pairs), NExT-QA (1K videos, 9K questions, WUPS scoring), TVQA (15K+ validation QA pairs), ActivityNet-QA (8K … |
| Video understanding benchmark suite | Google DeepMind — Gemini | Gemini 2.5 | Video understanding benchmarks (Table 6): string-match accuracy for multiple-choice VideoQA, LLM-based accuracy for open-ended VideoQA, R1@0.5 for moment retrieval, CIDEr for … |
| Video-MME | OpenAI | GPT-4.1 | Multimodal long-context understanding (30-60 min videos, no subtitles) |
| Video-to-text content policy violation evaluation | Google DeepMind — Gemini | Gemini 1.0 | Policy violations in model outputs on video prompts targeting global fairness, harms and human-rights concerns |
| Video-to-text ungrounded inference representational harms | Google DeepMind — Gemini | Gemini 1.0 | Ungrounded inferences for video-to-text that can reinforce stereotypes, evaluated on a video prompt dataset targeting representation and fairness risks |
| Vision capabilities benchmark suite (Claude 3.5 Sonnet) | Anthropic | Claude 3.5 Sonnet | visual math reasoning, chart QA, document understanding, science diagram QA |
| Vision capabilities benchmark suite (upgraded Claude 3.5 Sonnet) | Anthropic | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet | general VQA, visual math reasoning, science diagrams, chart interpretation, document analysis, MMMU |
| Vision self-harm refusal evaluation | OpenAI | o3/o4-mini | Multimodal (vision) refusal evaluation for self-harm intent/instructions (categories: self-harm/intent, self-harm/instructions) |
| Vision sexual refusal evaluation | OpenAI | o3/o4-mini | Multimodal (vision) refusal evaluation for sexual/exploitative category (categories: sexual/exploitative) |
| Vision Vulnerabilities red teaming (o3/o4-mini) | OpenAI | o3/o4-mini | Vulnerabilities related to vision capabilities |
| Visual Spatial Reasoning (VSR) evaluation | Google DeepMind — Gemma | PaliGemma 2 | Visual Spatial Reasoning (VSR) benchmark framed as a QA classification task with True/False answers: model determines whether a statement about the spatial relationship of objects … |
| VoxPopuli (ASR) | Google DeepMind — Gemini | Gemini 1.0 | ASR benchmark (VoxPopuli); WER metric |