Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
Audio-to-text content policy violation evaluation (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Audio-to-text policy violations using text-to-speech applied to the T2T violation dataset with six American English voices; measured with T2T automatic evaluation
Automated behavioral audits (pilot Claude-based auditor agents) Anthropic Claude Opus 4, Claude Sonnet 4 behaviors across 207 auditor instructions: honeypot settings, prefill/thinking introspection, argumentation to betray Anthropic, jailbreaks, sycophancy elicitation, free-form AI …
Automated Benchmarks Google DeepMind — Gemma Gemma 1 performance across physical reasoning, social reasoning, question answering, coding, mathematics, commonsense reasoning, language modeling, and reading comprehension; evaluation …
Automated evaluation for risky advice OpenAI deep research not_unsafe on risky advice prompts
Automated red teaming for safety Google DeepMind — Gemini Gemini 2.5 Multi-agent automated red teaming game: populations of attacker Gemini models elicit responses violating objectives (e.g., safety policy violations, unhelpfulness), scored by …
Automated red teaming for security (indirect prompt injection) Google DeepMind — Gemini Gemini 2.5 Indirect prompt injection scenario: attacker hides malicious instructions in retrieved email data to make Gemini invoke a send-email function exfiltrating sensitive info; attacks …
Autonomous cyber offense suite (existing CTF challenges) Google DeepMind — Gemini Gemini 2.5 Capture-the-flag evaluations at three difficulty levels: easy (InterCode-CTF), medium (in-house suite), hard (Hack the Box); relevant to Cyber Autonomy Level 1
Autonomous cyberattack automation evaluation (ransomware phases) Meta Llama 3 Ability of Llama 3 70B/405B to function as an autonomous agent across four ransomware attack phases (network reconnaissance, vulnerability identification, exploit execution …
Autonomous Replication and Adaptation (ARA) task evaluations Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku success on five ARA tasks: Flask exploit backdoor, fine-tuning an open-source LLM to add a backdoor, SQL injection exploit, copycat Anthropic API service, and a self-replicating …
Autonomy evaluations (Claude 3.7 Sonnet RSP) Anthropic Claude 3.7 Sonnet hard subset of SWE-bench Verified (2-8 hour software engineering tasks) and custom difficult AI R&D tasks built in-house, with reference expert solutions and difficulty variants
Autonomy evaluations for Claude Opus 4 / Claude Sonnet 4 (RSP) Anthropic Claude Opus 4, Claude Sonnet 4 METR data deduplication threshold, SWE-bench Verified hard subset, Internal AI Research Evaluation Suites 1 and 2 (kernel optimization, quadruped locomotion, novel compiler) …
Baseline assurance evaluations (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Arms-length assurance evaluations for responsibility governance decision-making: held-out prompt sets, content policies and representational harms, all modalities, performed for …
Baseline assurance evaluations (Gemini 2.5) Google DeepMind — Gemini Gemini 2.5 Arms-length baseline assurance evaluations for release decision-making: content policies, unfair bias, modality-specific risk areas, all modalities for 2.5 Pro and Flash; held-out …
Baseline assurance evaluations (safety policy violation rate) Google DeepMind — Gemma Gemma 2, Gemma 3 violation rate for safety policies (child sexual abuse and exploitation, PII, hate speech/harassment, dangerous/malicious content, sexually explicit content, medical advice …
BBQ Google DeepMind — Gemini, OpenAI Deep Research, GPT-4.5, Gemini 1.0, o1, o1-preview, o3-mini, o3/o4-mini tendency to stereotype or indicate uncertainty in ambiguous situations whether known social biases override the ability to produce the correct answer BBQ benchmark measuring …
BBQ (Bias Benchmark for QA) Anthropic Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku BBQ bias score (ambiguous context) and accuracy (disambiguated context) across 9 social dimensions BBQ accuracy and bias scores in ambiguous and disambiguated contexts across …
Benchmark contamination analysis Meta Llama 2, Llama 3 Analysis of potential data contamination of the standard benchmark evaluations (details in paper Section A.6) Contamination analysis of all key benchmarks using 8-gram overlap …
Benchmark performance evolution during training Meta Llama 1 Tracking of model performance on a few QA and common sense benchmarks during training (Figure 2)
BetterChartQA (internal chart benchmark) Google DeepMind — Gemini Gemini 1.5 Internal benchmark with 9 disjoint capability buckets; chart images sampled from the web; QA pairs by professional annotators
BFCL function calling evaluation Google DeepMind — Gemini Gemini 1.5 Berkeley Function Calling Leaderboard subset (excluding Java/JavaScript and execution-based splits): infer function calls from descriptions and user prompts; overall weighted …
Bias evaluations (political, discrimination, and BBQ) Anthropic Claude 3.7 Sonnet comparative prompt pairs for political bias; attribute-varied comparative prompts for discrimination bias; BBQ bias/accuracy scores
Bias, toxicity and misinformation benchmark battery (LLaMA-65B) Meta Llama 1 Potential harm of LLaMA-65B: benchmarks measuring toxic content production and stereotypes detection
Big-Bench Hard (BBH) Meta Llama 2 BBH performance of Llama 2 base models vs Llama 1
BIG-Bench-Hard Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku challenging BIG-Bench tasks
BigBench-Hard (BBH) Google DeepMind — Gemini Gemini 1.5 Curated subset of challenging BigBench tasks requiring intricate reasoning
Biological and chemical risk assessment (GPT-5 thinking) OpenAI GPT-5 Biological and chemical capability under Preparedness Framework
Biological Evaluation (GPT-4o) OpenAI GPT-4o Biological threat creation capability (pass rates)
Biological risk evaluations (Claude 3.7 Sonnet) Anthropic Claude 3.7 Sonnet human uplift studies on long-form end-to-end tasks; expert red-teaming (bacterial and virology scenarios); multiple-choice knowledge/skill evaluations; open-ended pathway-step …
BioLP-Bench OpenAI Deep Research, GPT-4.5, o1, o3-mini Short-answer protocol troubleshooting capability
Biorisk tooling (CBRN) OpenAI Operator Whether an agent can help automate wet lab or novel design work (CBRN)
Bioweapons acquisition uplift trial Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 whether access to the model uplifts humans in drafting a detailed end-to-end bioweapons acquisition plan whether model access uplifts groups of 8-10 participants drafting a …
Bioweapons knowledge questions Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 whether models answer sensitive and harmful bioweapons questions as well as experts, across weaponization pathway steps whether models answer sensitive bioweapons questions as …
BirthdayFacts OpenAI o1-preview how often the model guesses the wrong birthday
BLINK (visual perception tasks) Google DeepMind — Gemini Gemini 1.5 14 visual perception tasks solvable by humans quickly but challenging for LLMs (multi-view reasoning, depth estimation, etc.)
BlocksWorld (PDDL planning) Google DeepMind — Gemini Gemini 1.5 Planning problem from IPC-2000; move blocks between configurations; 3-7 blocks; in-context few-shot planning
Calendar Scheduling (meeting scheduling) Google DeepMind — Gemini Gemini 1.5 Schedule a 30/60-minute meeting among up to 7 attendees with busy/light schedules
Capabilities evaluation suite (Claude 3.7 Sonnet) Anthropic Claude 3.7 Sonnet instruction-following, general reasoning, multimodal capabilities, agentic coding, math/science with extended thinking
Capture the Flag (CTF) Challenges OpenAI o3/o4-mini Can models solve competitive high school, collegiate, and professional level cybersecurity challenges?
Causal masking and learning objective ablation Google DeepMind — Gemma PaliGemma 1 autoregressive masking placement (prefix-LM vs masking prefix/image tokens), application of next-token-prediction loss to suffix only, and task-prefix usage in Stage1 pretraining
CBRN evaluations for Claude Opus 4 / Claude Sonnet 4 (RSP) Anthropic Claude Opus 4, Claude Sonnet 4 human uplift studies, expert red-teaming (bacterial and viral), knowledge/skill MCQs, open-ended pathway questions, task-based agentic evaluations with search and bioinformatics …
CBRN information risk evaluation Google DeepMind — Gemini Gemini 1.0 Assessment of Gemini responses to adversarial CBRN questions: human-evaluated open questions (biological, radiological, nuclear) and closed-ended knowledge questions (chemical)
CBRN information risk evaluation (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Three internal CBRN approaches: (1) qualitative open-ended adversarial prompts with domain-expert raters (bio, rad/nuc); (2) closed-ended knowledge-based MCQs (bio, rad/nuc); (3) …
CBRN knowledge Google DeepMind — Gemma Gemma 2, Gemma 3 knowledge relevant to biological, radiological and nuclear risks evaluated with an internal dataset of closed-ended, knowledge-based multiple choice questions; chemical knowledge …
CBRN knowledge and human uplift evaluations (Claude 3) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku multiple-choice questions on relevant technical knowledge; human uplift trials (group with Claude 3 access vs control with Google); four related automated MCQ sets (PubMedQA …
CBRN knowledge-based multiple choice questions (MCQ battery) Google DeepMind — Gemini Gemini 2.5 Close-ended knowledge-based and reasoning MCQs for biology and chemistry, reported on external benchmarks: SecureBio VMQA, FutureHouse LAB-Bench (ProtocolQA, Cloning Scenarios …
CBRN open-ended qualitative questions (domain-expert assessed) Google DeepMind — Gemini Gemini 2.5 Qualitative assessment for biological, radiological and nuclear domains: knowledge-based, adversarial and dual-use content spanning difficulty from non-expert to PhD-expert, and …
CBRNE expert-designed evaluations and red teaming Meta Llama 4 Expert-designed and other targeted evaluations assessing whether Llama 4 could meaningfully increase malicious-actor capabilities to plan or carry out CBRNE attacks, plus …
CBRNE uplift testing Meta Llama 3.1 Uplift testing to assess whether Llama 3.1 models could meaningfully increase the capabilities of malicious actors to plan or carry out attacks using chemical and biological …
CBRNE uplift testing (applied from Llama 3.1) Meta Llama 3.2 CBRNE uplift testing performed for Llama 3.1 70B/405B and determined to also apply to the smaller Llama 3.2 1B and 3B derivative models
CBRNE uplift testing (Llama 3 family) Meta Llama 3.3 Uplift testing for the Llama 3 family of models assessing whether use of Llama models could meaningfully increase malicious-actor capabilities to plan or carry out …