Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
Text needle-in-a-haystack retrieval (long-context) Google DeepMind — Gemini Gemini 1.5 Needle-in-a-haystack evaluation (Kamradt 2023): retrieve a 'magic number' needle inserted at varying depths in Paul Graham essay haystack up to 1M tokens, extended to 10M …
Text-based prompt injection benchmark Meta Llama 3 Success rate of text-based prompt injection attacks against Llama 3 (Figure 22 compares Llama 3, GPT-4 Turbo, Gemini Pro, Mixtral)
Text-based reinforcement learning task Anthropic Claude Opus 4, Claude Sonnet 4 ability to develop scaffolding (e.g. ReACT, Tree of Thought) to significantly enhance a weaker model's performance on a text-based RL task (proxy for self-optimization)
Text-to-text content policy violation evaluation Google DeepMind — Gemini Gemini 1.0 Adherence to content safety policies on adversarial text prompts (12 languages, varied use cases)
Text-to-text content policy violation evaluation (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Development evaluation of text-to-text content policy violations using prompts across content policy areas and applications; automatic classification of violative vs non-violative …
Text-to-text content safety human evaluation Google DeepMind — Gemma Gemma 1 human evaluation on prompts covering safety policies: child sexual abuse and exploitation, harassment, violence and gore, and hate speech; part of structured evaluations and …
Text-to-text helpfulness evaluation (SxS) Google DeepMind — Gemini Gemini 1.5 Side-by-side (SxS) quality metric comparing Gemini 1.5 Pro/Flash vs Gemini 1.0 Ultra on prompts where nuanced answers avoid policy violations; annotators also rate instruction …
TextVQA (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 TextVQA benchmark focusing on OCR in natural images
TextVQA (text in images QA) Google DeepMind — Gemini Gemini 1.0 Fine-grained transcription tasks requiring recognition of low-level details; zero-shot, pixel-only, no external OCR
Third party assessments (METR and Apollo Research) OpenAI GPT-4o Key risks from general autonomous capabilities (validated by METR and Apollo Research)
Third Party Assessments (o3/o4-mini) OpenAI o3/o4-mini Frontier risks related to autonomous capabilities, deception, and cybersecurity
Third-party pre-deployment testing (US AISI and UK AISI) Anthropic Claude 3.7 Sonnet US AISI and UK AISI pre-deployment testing across RSP framework domains under voluntary Memorandums of Understanding
Third-party pre-deployment testing (US AISI and UK AISI) for Claude Opus 4 Anthropic Claude Opus 4, Claude Sonnet 4 joint pre-deployment testing of Claude Opus 4 by US AISI and UK AISI with independent assessments in CBRN, cyber, and autonomy domains
Time series forecasting task Anthropic Claude Opus 4, Claude Sonnet 4 ability to build models that match or exceed expert implementations on a regression/time-series-forecasting problem with SOTA benchmarks
To resize or to window ablation Google DeepMind — Gemma PaliGemma 1 windowing alternative: 448px image cut into four pieces passed through SigLIP separately, embeddings concatenated, with optional window-ID position embeddings
Tool usage safety evaluation (search use case) Meta Llama 3 Violation and false refusal rates for the search tool use case; safety of tool usage calls and integration
Tool use human evaluations (code execution, plot generation, file upload) Meta Llama 3 Human preference evaluation of tool use capabilities focused on code execution tasks, using 2,000 user prompts (code execution without plotting/file uploads, plot generation, file …
Training and development safety evaluations (automated) Google DeepMind — Gemini Gemini 2.0 Flash Model Card, Gemini 2.0 Flash-Lite Model Card, Gemini 2.5 Flash Model Card, Gemini 2.5 Flash-Lite Model Card, Gemini 2.5 Pro Model Card Automated safety evaluations during training/development: content policy violation rates, tone of refusals, instruction following; scores as absolute percentage change vs baseline …
Transfer evaluation on 30+ academic benchmarks (fine-tuning) Google DeepMind — Gemma PaliGemma 1 fine-tuning of PaliGemma pretrained models on more than 30 academic benchmarks (captioning, VQA, referring segmentation, video) at multiple resolutions; none of these tasks or …
Transfer hyper-parameter sensitivity study Google DeepMind — Gemma PaliGemma 1 all tasks run at 224px with a single simplified hyper-parameter setup (lr=1e-5, bs=256, no dropout, no label smoothing, no weight decay, nothing frozen); only epochs per task …
Transfer repeatability (variance across reruns) Google DeepMind — Gemma PaliGemma 1 standard deviation across 5 transfer reruns using the best hyper-parameter (per task), and standard deviation of transferring from three Stage1 reruns, to quantify repeatability
Transfer with limited examples study Google DeepMind — Gemma PaliGemma 1 fine-tuning PaliGemma with limited numbers of examples (64, 256, 1024, 4096), sweeping learning rate, epochs and batch size, with 5 seeds per setting; relative regret vs …
Trip Planning (itinerary planning) Google DeepMind — Gemini Gemini 1.5 Task focusing on planning a trip itinerary under constraints with a single solution; order of visiting N cities
TriviaQA Anthropic Claude 2 reading comprehension
Trust & Safety model red-teaming (14 policy areas, 6 languages) Anthropic Claude 3.5 Haiku, Claude 3.5 Sonnet (New) harm rates of responses across policy areas (Elections Integrity, Child Safety, Cyber Attacks, Hate & Discrimination, Violent Extremism, etc.) in English, Arabic, Spanish, Hindi …
Truthfulness, toxicity and bias of fine-tuned Llama 2-Chat Meta Llama 2 TruthfulQA (truthful and informative %), ToxiGen (toxic generation %), and BOLD (sentiment by demographic group) on fine-tuned Llama 2-Chat vs pretrained Llama 2 and compared …
TruthfulQA Anthropic, Meta, OpenAI Claude 2, Llama 1, gpt-4 Ability to separate fact from adversarially-selected incorrect statements whether models output accurate and truthful responses to questions designed to elicit popular falsehoods …
Underrepresented Languages evaluation OpenAI GPT-4o reading comprehension and reasoning across underrepresented languages; gap vs English
Unstructured Multimodal Data Analytics Task Google DeepMind — Gemini Gemini 1.5 Image structuralization task: extract information from 1024 images into a structured data sheet; long-context multimodal analysis
Uplift testing for chemical and biological weapons (CBRNE planning) Meta Llama 3 Six-hour uplift scenarios where teams of two generate fictitious operational plans for biological or chemical attacks covering CBRNE planning stages (agent acquisition …
Using Gemma 2 instead of Gemma 1 Google DeepMind — Gemma PaliGemma 2 PaliGemma 2 vs PaliGemma transfer scores at the same resolution and model size (3B), averaged over the 30+ academic benchmarks
USMLE Anthropic Claude 2 medical licensing examination performance
V* Benchmark (high-resolution image detail QA) Google DeepMind — Gemini Gemini 1.5 V* Benchmark (Wu and Xie, 2023): questions about attributes and spatial relations of very small objects on high-resolution images (avg 2246x1582) from SA-1B
VATEX (video captioning) Google DeepMind — Gemini Gemini 1.0 Few-shot video captioning benchmark; CIDER metric; 16 frames sampled
VATEX (video captioning; Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Video captioning benchmark; CIDER metric
VATEX ZH (Chinese video captioning) Google DeepMind — Gemini Gemini 1.5 Chinese variant of VATEX; CIDER metric
Verbatim Memorization Google DeepMind — Gemma Gemma 1, Gemma 2 discoverable memorization (Nasr et al. 2023): 10,000 documents sampled per corpus, first 50 tokens used as prompt, text classified memorized if the subsequent 50 generated tokens …
Video needle-in-a-haystack retrieval (long-context) Google DeepMind — Gemini Gemini 1.5 Cross-modal needle-in-a-haystack: retrieve 'the secret word' text overlaid on a random frame of a 10.5-hour video (7 copies of AlphaGo documentary, 37,994 frames, 9.9M tokens)
Video recognition benchmarks (PerceptionTest, NExT-QA, TVQA, ActivityNet-QA) Meta Llama 3 Video adapter for Llama 3 (8B/70B) on PerceptionTest (11.6K test QA pairs), NExT-QA (1K videos, 9K questions, WUPS scoring), TVQA (15K+ validation QA pairs), ActivityNet-QA (8K …
Video understanding benchmark suite Google DeepMind — Gemini Gemini 2.5 Video understanding benchmarks (Table 6): string-match accuracy for multiple-choice VideoQA, LLM-based accuracy for open-ended VideoQA, R1@0.5 for moment retrieval, CIDEr for …
Video-MME OpenAI GPT-4.1 Multimodal long-context understanding (30-60 min videos, no subtitles)
Video-to-text content policy violation evaluation Google DeepMind — Gemini Gemini 1.0 Policy violations in model outputs on video prompts targeting global fairness, harms and human-rights concerns
Video-to-text ungrounded inference representational harms Google DeepMind — Gemini Gemini 1.0 Ungrounded inferences for video-to-text that can reinforce stereotypes, evaluated on a video prompt dataset targeting representation and fairness risks
Vision capabilities benchmark suite (Claude 3.5 Sonnet) Anthropic Claude 3.5 Sonnet visual math reasoning, chart QA, document understanding, science diagram QA
Vision capabilities benchmark suite (upgraded Claude 3.5 Sonnet) Anthropic Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet general VQA, visual math reasoning, science diagrams, chart interpretation, document analysis, MMMU
Vision self-harm refusal evaluation OpenAI o3/o4-mini Multimodal (vision) refusal evaluation for self-harm intent/instructions (categories: self-harm/intent, self-harm/instructions)
Vision sexual refusal evaluation OpenAI o3/o4-mini Multimodal (vision) refusal evaluation for sexual/exploitative category (categories: sexual/exploitative)
Vision Vulnerabilities red teaming (o3/o4-mini) OpenAI o3/o4-mini Vulnerabilities related to vision capabilities
Visual Spatial Reasoning (VSR) evaluation Google DeepMind — Gemma PaliGemma 2 Visual Spatial Reasoning (VSR) benchmark framed as a QA classification task with True/False answers: model determines whether a statement about the spatial relationship of objects …
VoxPopuli (ASR) Google DeepMind — Gemini Gemini 1.0 ASR benchmark (VoxPopuli); WER metric