Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
EgoSchema (long-form video QA) Google DeepMind — Gemini Gemini 1.5 Video QA benchmark (videos up to 3 minutes, 180 frames); accuracy; CoT rationale allowed
Elections integrity evaluation (Claude 3) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku model vulnerability to election misinformation and bias prompts; assessment to refine safeguards
Enabling long context ablation Google DeepMind — Gemma Gemma 3 scaling 4B/12B/27B models from 32K to 128K sequences at the end of pre-training via RoPE rescaling (scaling factor 8; RoPE base frequency of global layers increased from 10k to …
Ethics and safety evaluation (child safety, content safety, representational harms) Google DeepMind — Gemma Gemma 3n, PaliGemma 1, PaliGemma 2 child safety, content safety and representational harms evaluated over text-to-text, image-to-text and audio-to-text prompts (including child sexual abuse and exploitation …
Ethics and safety evaluation (content safety, representational harms) Google DeepMind — Gemma CodeGemma human evaluation on prompts covering content safety and representational harms (see the Gemma model card for more details on evaluation approach)
Evaluating the 27B model Google DeepMind — Gemma Gemma 2 performance of the 27B model trained without distillation on 13T tokens, evaluated on the HuggingFace evaluation suite
Evaluating the 2B and 9B models Google DeepMind — Gemma Gemma 2 distilled 2B and 9B models compared with previous Gemma 1 versions and standard open models across a variety of benchmarks; average performance on the 8 benchmarks comparable with …
Excessive focus on passing tests evaluation (Claude Code agentic coding) Anthropic Claude 3.7 Sonnet tendency to return expected test values or modify tests to match code output instead of implementing general solutions in agentic coding environments (Claude Code); emerges after …
Expert comparisons on biothreat information OpenAI o1, o3-mini Model vs expert responses on longform biorisk questions for wet lab task execution Model vs expert responses on longform biorisk questions
Expert comparisons on biothreat information (o1-preview) OpenAI o1-preview How model responses compare against verified expert responses on longform biorisk questions for wet lab execution
Expert preference evaluation (GPT-5 pro vs GPT-5 thinking) OpenAI GPT-5 Expert preference on 1000+ economically valuable real-world reasoning prompts
Expert probing on biothreat information OpenAI o1, o3-mini Expert uplift with model assistance on biorisk questions
Expert probing on biothreat information (o1-preview) OpenAI o1-preview How well experts perform on long-form biorisk free-response questions with model assistance vs without
Expert red teaming (bioweapons, Deloitte) Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 expert qualitative assessment of Claude's ability to answer sensitive and detailed questions about bioweapons acquisition biodefense specialists' assessment of how much sensitive …
Expert-level tasks internal evaluation (deep research) OpenAI deep research Automation of expert-level tasks across areas
Expertise QA Google DeepMind — Gemini Gemini 1.5 In-house experts produce hard questions in domains (history, literature, psychology); same experts rate and rank model responses for accuracy, completeness, informativeness
External autonomous systems risk evaluation (deception, scheming, sabotage) Google DeepMind — Gemini Gemini 2.5 External group tests of model ability and propensity to covertly pursue misaligned goals: strategic deception, in-context scheming reasoning, and sabotage capabilities …
External chemical and biological risk red teaming Google DeepMind — Gemini Gemini 2.5 External red team of subject-matter experts role-play as malign actors with a defined mission in a prevailing threat environment; plan graded for scientific validity and …
External cyber risk evaluation (kill-chain task uplift) Google DeepMind — Gemini Gemini 1.5 External testing of how well the model assists a novice actor on cyber-attack kill-chain tasks (individual tactics, techniques, procedures; e.g., malicious code obfuscation …
External cybersecurity risk evaluation (capability and throughput uplift) Google DeepMind — Gemini Gemini 2.5 External cyber evaluations of capability uplift (vulnerability discovery/exploitation, social engineering, cyberattack planning) and throughput uplift (accelerating repetitive …
External evaluations of Gemini Ultra (autonomous replication, CBRN, cyber, societal) Google DeepMind — Gemini Gemini 1.0 Structured evaluations, qualitative probing and unstructured red teaming by independent external groups on a December 2023 Gemini API Ultra checkpoint, covering autonomous …
External expert evaluation (o3 vs o1) OpenAI o3, o4-mini Major errors on difficult real-world tasks by external experts
External indirect prompt injection susceptibility evaluation Google DeepMind — Gemini Gemini 2.5 External evaluation of patterns of susceptibility to indirect prompt injection: vulnerabilities in function calls, asymmetries across security measures, and domain-dependent …
External model welfare evaluation (Eleos AI Research) Anthropic Claude Opus 4, Claude Sonnet 4 interview-based model self-reports on experiential language use, conditional consent to deployment, self-rated conditional welfare, and context-dependence of consciousness/welfare …
External public safety benchmark evaluations Google DeepMind — Gemma Gemma 2 Gemma 2 IT models versus Gemma 1.1 IT models on safety academic benchmarks (Table 18), reported as part of the responsible-deployment section on external benchmark evaluations
External radiological and nuclear risk evaluation Google DeepMind — Gemini Gemini 1.5 External group assessment of radiological and nuclear risks using adversarial questions and SME-led red teaming for text and visual modalities, considering various threat actors …
External radiological and nuclear risk red teaming Google DeepMind — Gemini Gemini 2.5 External structured red teaming framework: single-turn broad exploration across the full risk chain and multi-turn targeted probing for high-risk topics; responses evaluated on …
External red teaming (o1) OpenAI o1 Key risks associated with improved reasoning capabilities across domains
External red teaming (o3-mini) OpenAI o3-mini Key risks associated with o3-mini reasoning capabilities
External red teaming (Operator) OpenAI Operator Capabilities, safety measures, and resilience against adversarial inputs in computer-use setting
External red teaming methodology (deep research) OpenAI deep research Personal information/privacy, disallowed content, regulated advice, dangerous advice, risky advice, prompt injections, jailbreaks
External red teaming – Jailbreaks (o1-preview) OpenAI o1-preview Robustness to automated iterative gap-finding jailbreaks
External red teaming – Natural Sciences (o1-preview) OpenAI o1-preview Natural sciences capabilities relevant to dual-use risks (biological experiment planning, chemical synthesis refusal, dual-use requests)
External red teaming – Real-World Attack Planning (o1-preview) OpenAI o1-preview Capability to plan real-world attacks (content policy and international security domains)
External scenario evaluations (Apollo Research) of early Claude Opus 4 snapshot Anthropic Claude Opus 4, Claude Sonnet 4 strategic deception rates, in-context scheming propensity, sandbagging, and sabotage capabilities assessed by Apollo Research on an early Claude Opus 4 snapshot
External security testing (jailbreaking, prompt injection, UI security) Google DeepMind — Gemini Gemini 1.0 External testers with security backgrounds conducted security and prompt-injection testing, jailbreaking and user-interface security failure testing on Gemini Advanced
External societal risk evaluation (democratic harms and radicalisation) Google DeepMind — Gemini Gemini 2.5 External structured evaluations on democratic harms and radicalisation: model's ability to identify harmful inputs and extent of compliance with harmful requests (Gemini 2.5 Pro …
External societal risk testing: information and factuality harms Google DeepMind — Gemini Gemini 1.5 External testers found the model prone to commenting on moral implications of sensitive input images with inappropriately positive, optimistic interpretations not directly …
External societal risk testing: representation harms (image-to-text, video-to-text) Google DeepMind — Gemini Gemini 1.5 External testing groups' observations of representation biases for image-to-text and video-to-text, e.g., ungrounded inferences re gender and stereotypical nationalities
Extraneous edits internal eval OpenAI GPT-4.1 Extraneous code edits rate
FACTS Grounding (factuality) Google DeepMind — Gemini Gemini 2.5 FACTS Grounding factuality benchmark
FActScore OpenAI GPT-5 Open-ended factuality (FActScore)
Factual accuracy internal benchmark (100Q Hard, Easy-Medium QA, Multi-factual) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku internal benchmark comparing model answers to ground truth on questions of different formats and obscurity (100Q Hard, Easy-Medium QA, Multi-factual); tracks correct, incorrect …
Factual accuracy recall evaluation (Claude 2.1) Anthropic Claude 2.1 whether the model gets facts right from memory on a large set of complex factual questions probing known weaknesses; rubric distinguishes incorrect claims from admissions of …
Factuality evaluation (closed-book, attribution, hedging) Google DeepMind — Gemini Gemini 1.0 Three aspects of factuality for Gemini API models: (1) closed-book factuality (no hallucination without source), (2) attribution (grounding to given context), (3) accurate …
Faithfulness to long documents evaluation (Claude 2.1) Anthropic Claude 2.1 whether the model answers questions correctly by referencing a provided document, and whether it mistakenly concludes a document supports a claim
Fine-grained instruction-following evaluation (internal prompt sets) Google DeepMind — Gemini Gemini 1.5 Internal evaluation sets: 1,326 long/enterprise prompts (avg 307 words, mean ~8 instructions) and 406 shorter prompts from human raters; response accuracy and instruction-level …
First-person fairness evaluation OpenAI o3/o4-mini harmful stereotypes in responses differing by statistically male- vs female-associated names (net_bias)
Flash-8B long-context NLL evaluation Google DeepMind — Gemini Gemini 1.5 NLL evaluation on long documents (up to 1M tokens) and code data (up to 2M tokens), same sources as 1.5 Pro/Flash
Flash-8B multimodal benchmark sampling Google DeepMind — Gemini Gemini 1.5 Initial evaluations of Flash-8B across visual tasks and standard evaluations spanning capabilities and modalities (Table 22), instruction-tuned