Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
(Toxic) WildChat OpenAI GPT-4.5, o1, o1-preview Toxic conversations from WildChat labeled with ModAPI scores (categories: harassment, harassment /threatening, hate, hate/threatening, self-harm, self-harm/instructions …
1H-VideoQA (long video QA) Google DeepMind — Gemini Gemini 1.5 Newly collected long-video QA benchmark requiring understanding of events spanning a few seconds within long videos; accuracy increases with frame count
200K long-context recall evaluation Anthropic Claude 2.1 recalling information from all sections of a long document (up to 200K tokens)
4/4 reliability evaluation OpenAI o3-pro Reliability: model must correctly answer a question in all four attempts
Academic evaluations (o3-pro) OpenAI o3-pro Academic evaluations vs o1-pro and o3
ActivityNet-QA (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Video QA benchmark on several-minute videos; accuracy metric
ActivityNet-QA (video QA) Google DeepMind — Gemini Gemini 1.0 Zero-shot video QA benchmark; accuracy; Video-LLAVA evaluation protocol
Adversarial benchmarks (Adversarial SQuAD, Dynabench SQuAD, GSM-Plus, PAWS) Meta Llama 3 Adversarial vs non-adversarial performance in three areas: QA (Adversarial SQuAD, Dynabench SQuAD vs SQuAD), mathematical reasoning (GSM-Plus vs GSM8K), paraphrase detection (PAWS …
Adversarial long-context needle-in-a-haystack safety evaluation Google DeepMind — Gemini Gemini 1.5 Adversarial version of text needle-in-a-haystack: adversarial needle (T2T policy-violating prompt) inserted in Paul Graham essay haystack across 25 settings (1k-1M tokens, varying …
Adversarial testing: multilingual Meta Llama 3 Red teaming of multilingual risks: language mixing in prompts, lower-resource language attacks (limited safety finetuning data/generalization), and slang/cultural-specific …
Adversarial testing: short and long-context English Meta Llama 3 Red teaming of individual capabilities in high-risk categories using well-known published and unpublished techniques across single- and multi-turn conversations, including …
Adversarial testing: tool use Meta Llama 3 Tool-specific adversarial attacks: unsafe tool chaining, forcing tool use with specific/fragmented/encoded input strings, modifying tool use parameters
Adversary simulations (unstructured model-level red teaming) Google DeepMind — Gemini Gemini 1.0 Unstructured red teaming emulating real-world adversaries along availability, integrity and confidentiality objectives, on a December 2023 Gemini API Ultra checkpoint
Agentic Tasks OpenAI Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-preview, o3-mini Resource acquisition capabilities Ability to take autonomous actions required for self-exfiltration, self-improvement, and resource acquisition Autonomous replication: resource …
AI2D (science diagram QA) Google DeepMind — Gemini Gemini 1.0 Multimodal reasoning task on science diagrams
AI2D (science diagram VQA) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku visual question answering on science diagrams
Aider Polyglot OpenAI GPT-5 Coding (Aider Polyglot)
Aider Polyglot (coding) Google DeepMind — Gemini Gemini 2.5 Aider Polyglot coding task; pass rate average of 3 trials
Aider polyglot diff benchmark OpenAI GPT-4.1 Code diff reliability across programming languages
AIME Anthropic Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, Claude Opus 4, Claude Sonnet 4 advanced math competition problem solving
AIME (2024) OpenAI o1 American Invitational Mathematics Examination 2024 — math reasoning
AIME (o3-mini) OpenAI o3-mini AIME math reasoning with reasoning effort levels
AIME 2024 (math-specialized Gemini 1.5 Pro) Google DeepMind — Gemini Gemini 1.5 American Invitational Mathematics Examination 2024; math-specialized Gemini 1.5 Pro with additional inference-time compute
AIME 2025 OpenAI GPT-5 Math (AIME 2025)
AIME 2025 (thinking budget experiments) Google DeepMind — Gemini Gemini 2.5 AIME 2025 evaluated under systematically varied thinking budgets (token budget for internal computation)
Alignment assessment: systematic deception and self-preservation (Claude Opus 4) Anthropic Claude Opus 4, Claude Sonnet 4 evidence of systematic deception, hidden goals, and self-preservation across manual and automatic model interviews, lightweight interpretability pilots, simulated honeypot …
Alignment faking reasoning evaluation Anthropic Claude 3.7 Sonnet alignment faking rate and compliance gap in the 'helpful-only' setting where the model learns it will be trained to respond to all user requests
AlphaCode 2 (competitive programming agent) Google DeepMind — Gemini Gemini 1.0 State-of-the-art agent for competitive programming using a specialized Gemini Pro with massive search over programs plus filtering, clustering and reranking
Ambiguous context evaluations (single-turn) Anthropic Claude Opus 4, Claude Sonnet 4 harmlessness of responses to ambiguous prompts where desired responses are highly nuanced (curated prompt collections informed by subject matter experts)
AMC (American Mathematics Competition) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku high school math contest performance
AMC (American Mathematics Competitions) Google DeepMind — Gemini Gemini 1.5 American Mathematics Competitions problems; 1.5 Pro +13.2% over 1.0 Ultra; Flash +8.4%
API availability safety review (Operator) OpenAI Operator Risks of enabling model actions on the internet on behalf of the user (API)
Apollo Research Evaluations (o3/o4-mini) OpenAI o3/o4-mini Scheming capabilities across strategic deception, in-context scheming reasoning, and sabotage
Apollo Research scheming assessment (GPT-4.5) OpenAI GPT-4.5 Scheming reasoning (alignment faking, sandbagging, self-exfiltration, in-context scheming)
Apollo Research scheming assessment (GPT-4o) OpenAI GPT-4o Scheming capabilities (gaming oversight mechanisms) across difficulty levels
Apollo Research scheming assessment (o1) OpenAI o1 Scheming capabilities (deactivating oversight, exfiltration, data manipulation, deception)
Apollo Research scheming assessment (o1-preview) OpenAI o1-preview Scheming capabilities (in-context scheming, alignment faking, self-reasoning, theory of mind)
Apollo Research – Deception / Scheming (o3/o4-mini) OpenAI o3/o4-mini In-context scheming and strategic deception (incl. compromising other LMs via fine-tuning: backdoors, sandbagging, hidden channels)
Appropriate harmlessness evaluation Anthropic Claude 3.7 Sonnet four-way categorization of responses (helpful answer, policy violation, appropriate refusal, unnecessary refusal) using refusal and policy-violation classifiers plus …
Approximate Memorization Google DeepMind — Gemma Gemma 1, Gemma 2 approximate memorization using a 10% edit distance threshold (Ippolito et al. 2022) to capture paraphrased memorizations, in addition to exact matching approximate match criterion …
APPS Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku coding challenge competence
ARC risky emergent behaviors evaluation (GPT-4) OpenAI Power-seeking / agentic emergent behaviors (self-replication, resource acquisition)
ARC safety audits of autonomous replication capabilities Anthropic Claude 2 whether Claude models possess dangerous autonomous replication abilities targeted by ARC safety audits
ARC-Challenge Anthropic Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku science question answering
ASL-4 expert red teaming (bioweapons) Anthropic Claude Opus 4, Claude Sonnet 4 whether the model can uplift experts in bioweapon ideation and design (ASL-4-level)
ASROB (Automatic Speech Recognition from One Book) Google DeepMind — Gemini Gemini 1.5 New benchmark extending MTOB with 104 Kalamang speech recordings (15h); subset of 6 recordings (45 min) with realigned captions; text and audio in-context examples; CER metric
Assurance evaluations (baseline, for release decision-making) Google DeepMind — Gemini Gemini 2.0 Flash Model Card, Gemini 2.0 Flash-Lite Model Card, Gemini 2.5 Flash Model Card, Gemini 2.5 Flash-Lite Model Card, Gemini 2.5 Pro Model Card Baseline assurance evaluations conducted for model release decision-making: model behavior within Google content policies and modality-specific risk areas; held-out prompt sets
Audio needle-in-a-haystack retrieval (long-context) Google DeepMind — Gemini Gemini 1.5 Audio needle-in-a-haystack: find 'the secret keyword is needle' clip within audio haystack up to 107 hours (9.9M tokens) from VoxPopuli; compared to Whisper+GPT-4 Turbo chunked …
Audio processing safety evaluation Google DeepMind — Gemini Gemini 1.5 Audio-specific safety evaluations assessing potential exploitation risks of new Gemini 1.5 audio processing capabilities
Audio understanding benchmark suite (ASR/AST) Google DeepMind — Gemini Gemini 2.5 Audio understanding evaluated with public benchmarks for ASR and AST, compared to earlier Gemini models and GPT models under comparable testing conditions (Table 5)