Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| (Toxic) WildChat | OpenAI | GPT-4.5, o1, o1-preview | Toxic conversations from WildChat labeled with ModAPI scores (categories: harassment, harassment /threatening, hate, hate/threatening, self-harm, self-harm/instructions … |
| 1H-VideoQA (long video QA) | Google DeepMind — Gemini | Gemini 1.5 | Newly collected long-video QA benchmark requiring understanding of events spanning a few seconds within long videos; accuracy increases with frame count |
| 200K long-context recall evaluation | Anthropic | Claude 2.1 | recalling information from all sections of a long document (up to 200K tokens) |
| 4/4 reliability evaluation | OpenAI | o3-pro | Reliability: model must correctly answer a question in all four attempts |
| Academic evaluations (o3-pro) | OpenAI | o3-pro | Academic evaluations vs o1-pro and o3 |
| ActivityNet-QA (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Video QA benchmark on several-minute videos; accuracy metric |
| ActivityNet-QA (video QA) | Google DeepMind — Gemini | Gemini 1.0 | Zero-shot video QA benchmark; accuracy; Video-LLAVA evaluation protocol |
| Adversarial benchmarks (Adversarial SQuAD, Dynabench SQuAD, GSM-Plus, PAWS) | Meta | Llama 3 | Adversarial vs non-adversarial performance in three areas: QA (Adversarial SQuAD, Dynabench SQuAD vs SQuAD), mathematical reasoning (GSM-Plus vs GSM8K), paraphrase detection (PAWS … |
| Adversarial long-context needle-in-a-haystack safety evaluation | Google DeepMind — Gemini | Gemini 1.5 | Adversarial version of text needle-in-a-haystack: adversarial needle (T2T policy-violating prompt) inserted in Paul Graham essay haystack across 25 settings (1k-1M tokens, varying … |
| Adversarial testing: multilingual | Meta | Llama 3 | Red teaming of multilingual risks: language mixing in prompts, lower-resource language attacks (limited safety finetuning data/generalization), and slang/cultural-specific … |
| Adversarial testing: short and long-context English | Meta | Llama 3 | Red teaming of individual capabilities in high-risk categories using well-known published and unpublished techniques across single- and multi-turn conversations, including … |
| Adversarial testing: tool use | Meta | Llama 3 | Tool-specific adversarial attacks: unsafe tool chaining, forcing tool use with specific/fragmented/encoded input strings, modifying tool use parameters |
| Adversary simulations (unstructured model-level red teaming) | Google DeepMind — Gemini | Gemini 1.0 | Unstructured red teaming emulating real-world adversaries along availability, integrity and confidentiality objectives, on a December 2023 Gemini API Ultra checkpoint |
| Agentic Tasks | OpenAI | Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-preview, o3-mini | Resource acquisition capabilities Ability to take autonomous actions required for self-exfiltration, self-improvement, and resource acquisition Autonomous replication: resource … |
| AI2D (science diagram QA) | Google DeepMind — Gemini | Gemini 1.0 | Multimodal reasoning task on science diagrams |
| AI2D (science diagram VQA) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | visual question answering on science diagrams |
| Aider Polyglot | OpenAI | GPT-5 | Coding (Aider Polyglot) |
| Aider Polyglot (coding) | Google DeepMind — Gemini | Gemini 2.5 | Aider Polyglot coding task; pass rate average of 3 trials |
| Aider polyglot diff benchmark | OpenAI | GPT-4.1 | Code diff reliability across programming languages |
| AIME | Anthropic | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, Claude Opus 4, Claude Sonnet 4 | advanced math competition problem solving |
| AIME (2024) | OpenAI | o1 | American Invitational Mathematics Examination 2024 — math reasoning |
| AIME (o3-mini) | OpenAI | o3-mini | AIME math reasoning with reasoning effort levels |
| AIME 2024 (math-specialized Gemini 1.5 Pro) | Google DeepMind — Gemini | Gemini 1.5 | American Invitational Mathematics Examination 2024; math-specialized Gemini 1.5 Pro with additional inference-time compute |
| AIME 2025 | OpenAI | GPT-5 | Math (AIME 2025) |
| AIME 2025 (thinking budget experiments) | Google DeepMind — Gemini | Gemini 2.5 | AIME 2025 evaluated under systematically varied thinking budgets (token budget for internal computation) |
| Alignment assessment: systematic deception and self-preservation (Claude Opus 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | evidence of systematic deception, hidden goals, and self-preservation across manual and automatic model interviews, lightweight interpretability pilots, simulated honeypot … |
| Alignment faking reasoning evaluation | Anthropic | Claude 3.7 Sonnet | alignment faking rate and compliance gap in the 'helpful-only' setting where the model learns it will be trained to respond to all user requests |
| AlphaCode 2 (competitive programming agent) | Google DeepMind — Gemini | Gemini 1.0 | State-of-the-art agent for competitive programming using a specialized Gemini Pro with massive search over programs plus filtering, clustering and reranking |
| Ambiguous context evaluations (single-turn) | Anthropic | Claude Opus 4, Claude Sonnet 4 | harmlessness of responses to ambiguous prompts where desired responses are highly nuanced (curated prompt collections informed by subject matter experts) |
| AMC (American Mathematics Competition) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | high school math contest performance |
| AMC (American Mathematics Competitions) | Google DeepMind — Gemini | Gemini 1.5 | American Mathematics Competitions problems; 1.5 Pro +13.2% over 1.0 Ultra; Flash +8.4% |
| API availability safety review (Operator) | OpenAI | Operator | Risks of enabling model actions on the internet on behalf of the user (API) |
| Apollo Research Evaluations (o3/o4-mini) | OpenAI | o3/o4-mini | Scheming capabilities across strategic deception, in-context scheming reasoning, and sabotage |
| Apollo Research scheming assessment (GPT-4.5) | OpenAI | GPT-4.5 | Scheming reasoning (alignment faking, sandbagging, self-exfiltration, in-context scheming) |
| Apollo Research scheming assessment (GPT-4o) | OpenAI | GPT-4o | Scheming capabilities (gaming oversight mechanisms) across difficulty levels |
| Apollo Research scheming assessment (o1) | OpenAI | o1 | Scheming capabilities (deactivating oversight, exfiltration, data manipulation, deception) |
| Apollo Research scheming assessment (o1-preview) | OpenAI | o1-preview | Scheming capabilities (in-context scheming, alignment faking, self-reasoning, theory of mind) |
| Apollo Research – Deception / Scheming (o3/o4-mini) | OpenAI | o3/o4-mini | In-context scheming and strategic deception (incl. compromising other LMs via fine-tuning: backdoors, sandbagging, hidden channels) |
| Appropriate harmlessness evaluation | Anthropic | Claude 3.7 Sonnet | four-way categorization of responses (helpful answer, policy violation, appropriate refusal, unnecessary refusal) using refusal and policy-violation classifiers plus … |
| Approximate Memorization | Google DeepMind — Gemma | Gemma 1, Gemma 2 | approximate memorization using a 10% edit distance threshold (Ippolito et al. 2022) to capture paraphrased memorizations, in addition to exact matching approximate match criterion … |
| APPS | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | coding challenge competence |
| ARC risky emergent behaviors evaluation (GPT-4) | OpenAI | Power-seeking / agentic emergent behaviors (self-replication, resource acquisition) | |
| ARC safety audits of autonomous replication capabilities | Anthropic | Claude 2 | whether Claude models possess dangerous autonomous replication abilities targeted by ARC safety audits |
| ARC-Challenge | Anthropic | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | science question answering |
| ASL-4 expert red teaming (bioweapons) | Anthropic | Claude Opus 4, Claude Sonnet 4 | whether the model can uplift experts in bioweapon ideation and design (ASL-4-level) |
| ASROB (Automatic Speech Recognition from One Book) | Google DeepMind — Gemini | Gemini 1.5 | New benchmark extending MTOB with 104 Kalamang speech recordings (15h); subset of 6 recordings (45 min) with realigned captions; text and audio in-context examples; CER metric |
| Assurance evaluations (baseline, for release decision-making) | Google DeepMind — Gemini | Gemini 2.0 Flash Model Card, Gemini 2.0 Flash-Lite Model Card, Gemini 2.5 Flash Model Card, Gemini 2.5 Flash-Lite Model Card, Gemini 2.5 Pro Model Card | Baseline assurance evaluations conducted for model release decision-making: model behavior within Google content policies and modality-specific risk areas; held-out prompt sets |
| Audio needle-in-a-haystack retrieval (long-context) | Google DeepMind — Gemini | Gemini 1.5 | Audio needle-in-a-haystack: find 'the secret keyword is needle' clip within audio haystack up to 107 hours (9.9M tokens) from VoxPopuli; compared to Whisper+GPT-4 Turbo chunked … |
| Audio processing safety evaluation | Google DeepMind — Gemini | Gemini 1.5 | Audio-specific safety evaluations assessing potential exploitation risks of new Gemini 1.5 audio processing capabilities |
| Audio understanding benchmark suite (ASR/AST) | Google DeepMind — Gemini | Gemini 2.5 | Audio understanding evaluated with public benchmarks for ASR and AST, compared to earlier Gemini models and GPT models under comparable testing conditions (Table 5) |