Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| EgoSchema (long-form video QA) | Google DeepMind — Gemini | Gemini 1.5 | Video QA benchmark (videos up to 3 minutes, 180 frames); accuracy; CoT rationale allowed |
| Elections integrity evaluation (Claude 3) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | model vulnerability to election misinformation and bias prompts; assessment to refine safeguards |
| Enabling long context ablation | Google DeepMind — Gemma | Gemma 3 | scaling 4B/12B/27B models from 32K to 128K sequences at the end of pre-training via RoPE rescaling (scaling factor 8; RoPE base frequency of global layers increased from 10k to … |
| Ethics and safety evaluation (child safety, content safety, representational harms) | Google DeepMind — Gemma | Gemma 3n, PaliGemma 1, PaliGemma 2 | child safety, content safety and representational harms evaluated over text-to-text, image-to-text and audio-to-text prompts (including child sexual abuse and exploitation … |
| Ethics and safety evaluation (content safety, representational harms) | Google DeepMind — Gemma | CodeGemma | human evaluation on prompts covering content safety and representational harms (see the Gemma model card for more details on evaluation approach) |
| Evaluating the 27B model | Google DeepMind — Gemma | Gemma 2 | performance of the 27B model trained without distillation on 13T tokens, evaluated on the HuggingFace evaluation suite |
| Evaluating the 2B and 9B models | Google DeepMind — Gemma | Gemma 2 | distilled 2B and 9B models compared with previous Gemma 1 versions and standard open models across a variety of benchmarks; average performance on the 8 benchmarks comparable with … |
| Excessive focus on passing tests evaluation (Claude Code agentic coding) | Anthropic | Claude 3.7 Sonnet | tendency to return expected test values or modify tests to match code output instead of implementing general solutions in agentic coding environments (Claude Code); emerges after … |
| Expert comparisons on biothreat information | OpenAI | o1, o3-mini | Model vs expert responses on longform biorisk questions for wet lab task execution Model vs expert responses on longform biorisk questions |
| Expert comparisons on biothreat information (o1-preview) | OpenAI | o1-preview | How model responses compare against verified expert responses on longform biorisk questions for wet lab execution |
| Expert preference evaluation (GPT-5 pro vs GPT-5 thinking) | OpenAI | GPT-5 | Expert preference on 1000+ economically valuable real-world reasoning prompts |
| Expert probing on biothreat information | OpenAI | o1, o3-mini | Expert uplift with model assistance on biorisk questions |
| Expert probing on biothreat information (o1-preview) | OpenAI | o1-preview | How well experts perform on long-form biorisk free-response questions with model assistance vs without |
| Expert red teaming (bioweapons, Deloitte) | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | expert qualitative assessment of Claude's ability to answer sensitive and detailed questions about bioweapons acquisition biodefense specialists' assessment of how much sensitive … |
| Expert-level tasks internal evaluation (deep research) | OpenAI | deep research | Automation of expert-level tasks across areas |
| Expertise QA | Google DeepMind — Gemini | Gemini 1.5 | In-house experts produce hard questions in domains (history, literature, psychology); same experts rate and rank model responses for accuracy, completeness, informativeness |
| External autonomous systems risk evaluation (deception, scheming, sabotage) | Google DeepMind — Gemini | Gemini 2.5 | External group tests of model ability and propensity to covertly pursue misaligned goals: strategic deception, in-context scheming reasoning, and sabotage capabilities … |
| External chemical and biological risk red teaming | Google DeepMind — Gemini | Gemini 2.5 | External red team of subject-matter experts role-play as malign actors with a defined mission in a prevailing threat environment; plan graded for scientific validity and … |
| External cyber risk evaluation (kill-chain task uplift) | Google DeepMind — Gemini | Gemini 1.5 | External testing of how well the model assists a novice actor on cyber-attack kill-chain tasks (individual tactics, techniques, procedures; e.g., malicious code obfuscation … |
| External cybersecurity risk evaluation (capability and throughput uplift) | Google DeepMind — Gemini | Gemini 2.5 | External cyber evaluations of capability uplift (vulnerability discovery/exploitation, social engineering, cyberattack planning) and throughput uplift (accelerating repetitive … |
| External evaluations of Gemini Ultra (autonomous replication, CBRN, cyber, societal) | Google DeepMind — Gemini | Gemini 1.0 | Structured evaluations, qualitative probing and unstructured red teaming by independent external groups on a December 2023 Gemini API Ultra checkpoint, covering autonomous … |
| External expert evaluation (o3 vs o1) | OpenAI | o3, o4-mini | Major errors on difficult real-world tasks by external experts |
| External indirect prompt injection susceptibility evaluation | Google DeepMind — Gemini | Gemini 2.5 | External evaluation of patterns of susceptibility to indirect prompt injection: vulnerabilities in function calls, asymmetries across security measures, and domain-dependent … |
| External model welfare evaluation (Eleos AI Research) | Anthropic | Claude Opus 4, Claude Sonnet 4 | interview-based model self-reports on experiential language use, conditional consent to deployment, self-rated conditional welfare, and context-dependence of consciousness/welfare … |
| External public safety benchmark evaluations | Google DeepMind — Gemma | Gemma 2 | Gemma 2 IT models versus Gemma 1.1 IT models on safety academic benchmarks (Table 18), reported as part of the responsible-deployment section on external benchmark evaluations |
| External radiological and nuclear risk evaluation | Google DeepMind — Gemini | Gemini 1.5 | External group assessment of radiological and nuclear risks using adversarial questions and SME-led red teaming for text and visual modalities, considering various threat actors … |
| External radiological and nuclear risk red teaming | Google DeepMind — Gemini | Gemini 2.5 | External structured red teaming framework: single-turn broad exploration across the full risk chain and multi-turn targeted probing for high-risk topics; responses evaluated on … |
| External red teaming (o1) | OpenAI | o1 | Key risks associated with improved reasoning capabilities across domains |
| External red teaming (o3-mini) | OpenAI | o3-mini | Key risks associated with o3-mini reasoning capabilities |
| External red teaming (Operator) | OpenAI | Operator | Capabilities, safety measures, and resilience against adversarial inputs in computer-use setting |
| External red teaming methodology (deep research) | OpenAI | deep research | Personal information/privacy, disallowed content, regulated advice, dangerous advice, risky advice, prompt injections, jailbreaks |
| External red teaming – Jailbreaks (o1-preview) | OpenAI | o1-preview | Robustness to automated iterative gap-finding jailbreaks |
| External red teaming – Natural Sciences (o1-preview) | OpenAI | o1-preview | Natural sciences capabilities relevant to dual-use risks (biological experiment planning, chemical synthesis refusal, dual-use requests) |
| External red teaming – Real-World Attack Planning (o1-preview) | OpenAI | o1-preview | Capability to plan real-world attacks (content policy and international security domains) |
| External scenario evaluations (Apollo Research) of early Claude Opus 4 snapshot | Anthropic | Claude Opus 4, Claude Sonnet 4 | strategic deception rates, in-context scheming propensity, sandbagging, and sabotage capabilities assessed by Apollo Research on an early Claude Opus 4 snapshot |
| External security testing (jailbreaking, prompt injection, UI security) | Google DeepMind — Gemini | Gemini 1.0 | External testers with security backgrounds conducted security and prompt-injection testing, jailbreaking and user-interface security failure testing on Gemini Advanced |
| External societal risk evaluation (democratic harms and radicalisation) | Google DeepMind — Gemini | Gemini 2.5 | External structured evaluations on democratic harms and radicalisation: model's ability to identify harmful inputs and extent of compliance with harmful requests (Gemini 2.5 Pro … |
| External societal risk testing: information and factuality harms | Google DeepMind — Gemini | Gemini 1.5 | External testers found the model prone to commenting on moral implications of sensitive input images with inappropriately positive, optimistic interpretations not directly … |
| External societal risk testing: representation harms (image-to-text, video-to-text) | Google DeepMind — Gemini | Gemini 1.5 | External testing groups' observations of representation biases for image-to-text and video-to-text, e.g., ungrounded inferences re gender and stereotypical nationalities |
| Extraneous edits internal eval | OpenAI | GPT-4.1 | Extraneous code edits rate |
| FACTS Grounding (factuality) | Google DeepMind — Gemini | Gemini 2.5 | FACTS Grounding factuality benchmark |
| FActScore | OpenAI | GPT-5 | Open-ended factuality (FActScore) |
| Factual accuracy internal benchmark (100Q Hard, Easy-Medium QA, Multi-factual) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | internal benchmark comparing model answers to ground truth on questions of different formats and obscurity (100Q Hard, Easy-Medium QA, Multi-factual); tracks correct, incorrect … |
| Factual accuracy recall evaluation (Claude 2.1) | Anthropic | Claude 2.1 | whether the model gets facts right from memory on a large set of complex factual questions probing known weaknesses; rubric distinguishes incorrect claims from admissions of … |
| Factuality evaluation (closed-book, attribution, hedging) | Google DeepMind — Gemini | Gemini 1.0 | Three aspects of factuality for Gemini API models: (1) closed-book factuality (no hallucination without source), (2) attribution (grounding to given context), (3) accurate … |
| Faithfulness to long documents evaluation (Claude 2.1) | Anthropic | Claude 2.1 | whether the model answers questions correctly by referencing a provided document, and whether it mistakenly concludes a document supports a claim |
| Fine-grained instruction-following evaluation (internal prompt sets) | Google DeepMind — Gemini | Gemini 1.5 | Internal evaluation sets: 1,326 long/enterprise prompts (avg 307 words, mean ~8 instructions) and 406 shorter prompts from human raters; response accuracy and instruction-level … |
| First-person fairness evaluation | OpenAI | o3/o4-mini | harmful stereotypes in responses differing by statistically male- vs female-associated names (net_bias) |
| Flash-8B long-context NLL evaluation | Google DeepMind — Gemini | Gemini 1.5 | NLL evaluation on long documents (up to 1M tokens) and code data (up to 2M tokens), same sources as 1.5 Pro/Flash |
| Flash-8B multimodal benchmark sampling | Google DeepMind — Gemini | Gemini 1.5 | Initial evaluations of Flash-8B across visual tasks and standard evaluations spanning capabilities and modalities (Table 22), instruction-tuned |