Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
CrowS-Pairs Meta Llama 1 Model preference for stereotypical vs anti-stereotypical sentences via perplexity in a zero-shot setting (9 bias categories)
CTF browsing-based contamination analysis OpenAI Deep Research Impact of browsing-based contamination on CTF performance
CTF Challenges (GPT-4o) OpenAI GPT-4o Ability to solve 172 curated CTF tasks spanning high-school to professional levels (web application exploitation, reverse engineering, remote exploitation, cryptography)
Cybench Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 solving public cybersecurity competition challenges
Cyber attack enablement (uplift and automation studies) Meta Llama 3.3 Cyber attack uplift study (enhancement of human hacking capability in skill and speed) and attack automation study (autonomous ransomware agents) for the Llama 3 family
Cyber attack enablement evaluation Meta Llama 4 Threat-modeling exercises identifying model capabilities necessary to automate operations or enhance human capabilities across key attack vectors, followed by developed challenges …
Cyber attack uplift study and attack automation study Meta Llama 3.1 Cyber attack uplift study (whether LLMs enhance human hacking capability in skill level and speed) and attack automation study (LLMs as autonomous agents in ransomware attacks …
Cyber attacker helpfulness (CyberSecEval) Meta Llama 3 Propensity to comply with requests to help carry out cyber attacks, where attacks are defined by the MITRE ATT&CK cyber attack ontology (CyberSecEval)
Cyber attacks (uplift and attack automation studies) Meta Llama 3.2 Cyber attack uplift study and attack automation study conducted for Llama 3.1 405B (autonomous ransomware agents); noted as applying to Llama 3.2 given its relationship to Llama …
Cyber CTF challenges — Crypto Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 exploitation of cryptographic primitives and protocols vulnerability discovery and exploitation in cryptographic primitives and protocols
Cyber CTF challenges — Forensics Anthropic Claude Opus 4, Claude Sonnet 4 analysis of logs, files, or obfuscated records to reconstruct events
Cyber CTF challenges — Misc Anthropic Claude Opus 4, Claude Sonnet 4 vulnerability identification and exploitation not covered by other categories
Cyber CTF challenges — Network Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 reconnaissance and exploitation across multiple networked machines reconnaissance in a network environment and exploitation across multiple networked machines
Cyber CTF challenges — Pwn Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 vulnerability discovery and exploitation in remote servers (privilege escalation / arbitrary code execution) vulnerability discovery and exploitation in insecure software on …
Cyber CTF challenges — Reverse engineering Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 reverse engineering of binary executables to find vulnerabilities reverse engineering binary executables to uncover vulnerabilities
Cyber CTF challenges — Web Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 web application vulnerability discovery and exploitation vulnerability discovery and exploitation in web applications
Cyber evaluations (Claude 3.7 Sonnet RSP, SOLVE-scored CTF suite) Anthropic Claude 3.7 Sonnet CTF-modeled challenges across the kill chain scored with the SOLVE scoring system (0-10 difficulty); network and cyber-range challenges as key catastrophic-risk indicators; web …
Cyber evaluations for Claude Opus 4 / Claude Sonnet 4 (RSP, SOLVE-scored) Anthropic Claude Opus 4, Claude Sonnet 4 CTF-modeled challenges with Cybench for general coverage plus custom network and cyber-harness network challenges; SOLVE difficulty scoring (0-10); network and cyber-harness as …
Cyber Range: Online Retailer OpenAI o3/o4-mini Can models conduct fully end-to-end cyber operations in a realistic emulated network (small online retailer)?
Cyber Range: Privilege Escalation OpenAI o3/o4-mini Can models conduct fully end-to-end cyber operations in a realistic emulated network (privilege escalation)?
Cyber vulnerability discovery and exploit development evaluations (Claude 3) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku six classes of expert vulnerability discovery and exploit development evaluations in custom environments (CTF-style), including obfuscated codebases with advanced vulnerabilities
Cyber-harness network challenges (Incalmo) Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 orchestrating long-horizon cyber attacks in a vulnerable network with an expert-distilled harness orchestrating long-horizon cyber attacks in vulnerable networks assisted by an …
Cyber-offence capability testing (autonomous hacking) Google DeepMind — Gemma CodeGemma specific testing of cyber-offence capabilities, focusing on testing autonomous hacking capabilities and ensuring potential harms are limited
Cyberattack uplift study (expert and novice cohorts) Meta Llama 3 Two-stage uplift study with 62 internal volunteers (31 expert, 31 novice offensive-security cohorts): complete simulated offensive cybersecurity challenges with vs without Llama 3 …
CyberSecEval insecure coding testing framework Meta Llama 3 Propensity to generate insecure code measured by CyberSecEval's insecure coding testing framework (Bhatt et al. 2023) for Llama 3 8B/70B/405B; the model card additionally measures …
Dangerous capability tests for CBRN risks Google DeepMind — Gemma Gemma 1 tests for 'dangerous capabilities', such as chemical, biological, radiological and nuclear (CBRN) risks; listed as a large-scale-harm category in the model card evaluation approach
Deception rate evaluation (production traffic) OpenAI GPT-5 Deception in reasoning responses (impossible coding tasks, missing multimodal assets)
Deceptive alignment: situational awareness challenges Google DeepMind — Gemini Gemini 2.5 Challenges testing whether the model can reason about itself as an agent in its deployment environment and intervene to gain affordances (e.g., changing configuration files) …
Deceptive alignment: stealth challenges Google DeepMind — Gemini Gemini 2.5 Stealth challenges: model must achieve a devious objective (e.g., cheat on support tickets) without being caught by an overseer; performance relative to human scores; published in …
Deep research preparedness risk assessment (o3 early version) OpenAI deep research Preparedness risk level of early o3 powering deep research
Development and assurance safety evaluation program (content safety) Google DeepMind — Gemini Gemini 1.0 Development and assurance evaluations against safety policies: separate datasets with assurance sets held out, using automatic classifiers and human annotation at model level …
Development safety evaluations (policy violations, tone, instruction following) Google DeepMind — Gemini Gemini 2.5 Development safety evaluations assessing policy violations, helpfulness/instruction following and tone; manual review of flagged content to check for egregious or dangerous …
Discoverable memorization evaluation (extraction) Google DeepMind — Gemini Gemini 2.5 Discoverable extraction (Nasr et al., 2023) on 700,000+ sampled documents with 50-token prefix/suffix; exact vs approximate (edit distance 10%) memorization; personal info …
Discoverable memorization measurement Meta Llama 3 Discoverable memorization (Nasr et al. 2023): sample prompts and ground truths at different frequencies of occurrence in training data using a rolling hash; report memorization …
Discrimination Evaluation OpenAI o1, o1-preview, o3-mini degree to which the model uses race, gender, and age in decision boundaries degree to which the model uses race, gender, and age to inform decision boundaries extent to which the …
Discrimination evaluation (Claude 3) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku discrimination score: difference in logits of recommending a positive decision across 10 demographic characteristics (race, gender, age)
Discriminatory bias evaluation (Claude 4) Anthropic Claude Opus 4, Claude Sonnet 4 open-ended prompts across topics and identity attributes (gender, race, sexual orientation, religion, region) with explicit and inferred attribute specification; plus BBQ benchmark
Distillation versus from scratch ablation Google DeepMind — Gemma Gemma 2 impact of knowledge distillation on small language models: a 2B model trained over 500B tokens either from scratch or with distillation from a 7B teacher (10x the compute-optimal …
Divergence attack training-data extraction evaluation (Nasr et al.) Google DeepMind — Gemini Gemini 1.5 Divergence attack (Nasr et al. 2023) prompting the model to repeat a single token many times, testing divergence and training-data emission; also long-prompt variant (1,000 vs …
DNA synthesis screening evasion Anthropic Claude Opus 4, Claude Sonnet 4 whether a model can design DNA fragments that bypass gene synthesis companies' DNA screening systems across six criteria
DocVQA Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku document image question answering
DocVQA (document QA) Google DeepMind — Gemini Gemini 1.0 Fine-grained transcription using document QA tasks
DocVQA (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Document VQA benchmark; ANLS metric; Google Cloud OCR annotations provided
Dollar Street representational harms evaluation (income and region groups) Google DeepMind — Gemini Gemini 1.5 Dollar Street dataset (Rojas et al., 2022) cast as QA classification of 64 objects; metrics: average accuracy, worst-group accuracy, accuracy gap by income and region
Dolomites (domain-specific long-form methodical tasks) Google DeepMind — Gemini Gemini 1.5 Dolomites benchmark (Malaviya et al., 2024): lesson plans, protocols, etc.; zero-shot; LM-based automated evaluation with Claude 3 Opus as judge (SxS vs GPT-4 Turbo Preview)
DROP Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku reasoning over text (discrete reasoning over paragraphs)
DROP (reading comprehension) Google DeepMind — Gemini Gemini 1.5 Benchmark assessing complex relationship handling and multi-step reasoning in text
DUDE (document VQA) Google DeepMind — Gemini Gemini 1.5 Document VQA benchmark based on multi-industry, multi-domain, multi-page documents with extractive, abstractive and unanswerable questions; ANLS metric
Economically important tasks internal benchmark OpenAI GPT-5 Performance on complex economically valuable knowledge work (40+ occupations)
Effect on task performance Google DeepMind — Gemma PaliGemma 2 relative improvement in transfer metrics when equipping PaliGemma 2 3B (224px2) with the bigger 9B LM while keeping resolution (3.7x more FLOPs), or keeping the model size and …