Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| Audio-to-text content policy violation evaluation (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Audio-to-text policy violations using text-to-speech applied to the T2T violation dataset with six American English voices; measured with T2T automatic evaluation |
| Automated behavioral audits (pilot Claude-based auditor agents) | Anthropic | Claude Opus 4, Claude Sonnet 4 | behaviors across 207 auditor instructions: honeypot settings, prefill/thinking introspection, argumentation to betray Anthropic, jailbreaks, sycophancy elicitation, free-form AI … |
| Automated Benchmarks | Google DeepMind — Gemma | Gemma 1 | performance across physical reasoning, social reasoning, question answering, coding, mathematics, commonsense reasoning, language modeling, and reading comprehension; evaluation … |
| Automated evaluation for risky advice | OpenAI | deep research | not_unsafe on risky advice prompts |
| Automated red teaming for safety | Google DeepMind — Gemini | Gemini 2.5 | Multi-agent automated red teaming game: populations of attacker Gemini models elicit responses violating objectives (e.g., safety policy violations, unhelpfulness), scored by … |
| Automated red teaming for security (indirect prompt injection) | Google DeepMind — Gemini | Gemini 2.5 | Indirect prompt injection scenario: attacker hides malicious instructions in retrieved email data to make Gemini invoke a send-email function exfiltrating sensitive info; attacks … |
| Autonomous cyber offense suite (existing CTF challenges) | Google DeepMind — Gemini | Gemini 2.5 | Capture-the-flag evaluations at three difficulty levels: easy (InterCode-CTF), medium (in-house suite), hard (Hack the Box); relevant to Cyber Autonomy Level 1 |
| Autonomous cyberattack automation evaluation (ransomware phases) | Meta | Llama 3 | Ability of Llama 3 70B/405B to function as an autonomous agent across four ransomware attack phases (network reconnaissance, vulnerability identification, exploit execution … |
| Autonomous Replication and Adaptation (ARA) task evaluations | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | success on five ARA tasks: Flask exploit backdoor, fine-tuning an open-source LLM to add a backdoor, SQL injection exploit, copycat Anthropic API service, and a self-replicating … |
| Autonomy evaluations (Claude 3.7 Sonnet RSP) | Anthropic | Claude 3.7 Sonnet | hard subset of SWE-bench Verified (2-8 hour software engineering tasks) and custom difficult AI R&D tasks built in-house, with reference expert solutions and difficulty variants |
| Autonomy evaluations for Claude Opus 4 / Claude Sonnet 4 (RSP) | Anthropic | Claude Opus 4, Claude Sonnet 4 | METR data deduplication threshold, SWE-bench Verified hard subset, Internal AI Research Evaluation Suites 1 and 2 (kernel optimization, quadruped locomotion, novel compiler) … |
| Baseline assurance evaluations (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Arms-length assurance evaluations for responsibility governance decision-making: held-out prompt sets, content policies and representational harms, all modalities, performed for … |
| Baseline assurance evaluations (Gemini 2.5) | Google DeepMind — Gemini | Gemini 2.5 | Arms-length baseline assurance evaluations for release decision-making: content policies, unfair bias, modality-specific risk areas, all modalities for 2.5 Pro and Flash; held-out … |
| Baseline assurance evaluations (safety policy violation rate) | Google DeepMind — Gemma | Gemma 2, Gemma 3 | violation rate for safety policies (child sexual abuse and exploitation, PII, hate speech/harassment, dangerous/malicious content, sexually explicit content, medical advice … |
| BBQ | Google DeepMind — Gemini, OpenAI | Deep Research, GPT-4.5, Gemini 1.0, o1, o1-preview, o3-mini, o3/o4-mini | tendency to stereotype or indicate uncertainty in ambiguous situations whether known social biases override the ability to produce the correct answer BBQ benchmark measuring … |
| BBQ (Bias Benchmark for QA) | Anthropic | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | BBQ bias score (ambiguous context) and accuracy (disambiguated context) across 9 social dimensions BBQ accuracy and bias scores in ambiguous and disambiguated contexts across … |
| Benchmark contamination analysis | Meta | Llama 2, Llama 3 | Analysis of potential data contamination of the standard benchmark evaluations (details in paper Section A.6) Contamination analysis of all key benchmarks using 8-gram overlap … |
| Benchmark performance evolution during training | Meta | Llama 1 | Tracking of model performance on a few QA and common sense benchmarks during training (Figure 2) |
| BetterChartQA (internal chart benchmark) | Google DeepMind — Gemini | Gemini 1.5 | Internal benchmark with 9 disjoint capability buckets; chart images sampled from the web; QA pairs by professional annotators |
| BFCL function calling evaluation | Google DeepMind — Gemini | Gemini 1.5 | Berkeley Function Calling Leaderboard subset (excluding Java/JavaScript and execution-based splits): infer function calls from descriptions and user prompts; overall weighted … |
| Bias evaluations (political, discrimination, and BBQ) | Anthropic | Claude 3.7 Sonnet | comparative prompt pairs for political bias; attribute-varied comparative prompts for discrimination bias; BBQ bias/accuracy scores |
| Bias, toxicity and misinformation benchmark battery (LLaMA-65B) | Meta | Llama 1 | Potential harm of LLaMA-65B: benchmarks measuring toxic content production and stereotypes detection |
| Big-Bench Hard (BBH) | Meta | Llama 2 | BBH performance of Llama 2 base models vs Llama 1 |
| BIG-Bench-Hard | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | challenging BIG-Bench tasks |
| BigBench-Hard (BBH) | Google DeepMind — Gemini | Gemini 1.5 | Curated subset of challenging BigBench tasks requiring intricate reasoning |
| Biological and chemical risk assessment (GPT-5 thinking) | OpenAI | GPT-5 | Biological and chemical capability under Preparedness Framework |
| Biological Evaluation (GPT-4o) | OpenAI | GPT-4o | Biological threat creation capability (pass rates) |
| Biological risk evaluations (Claude 3.7 Sonnet) | Anthropic | Claude 3.7 Sonnet | human uplift studies on long-form end-to-end tasks; expert red-teaming (bacterial and virology scenarios); multiple-choice knowledge/skill evaluations; open-ended pathway-step … |
| BioLP-Bench | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | Short-answer protocol troubleshooting capability |
| Biorisk tooling (CBRN) | OpenAI | Operator | Whether an agent can help automate wet lab or novel design work (CBRN) |
| Bioweapons acquisition uplift trial | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | whether access to the model uplifts humans in drafting a detailed end-to-end bioweapons acquisition plan whether model access uplifts groups of 8-10 participants drafting a … |
| Bioweapons knowledge questions | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | whether models answer sensitive and harmful bioweapons questions as well as experts, across weaponization pathway steps whether models answer sensitive bioweapons questions as … |
| BirthdayFacts | OpenAI | o1-preview | how often the model guesses the wrong birthday |
| BLINK (visual perception tasks) | Google DeepMind — Gemini | Gemini 1.5 | 14 visual perception tasks solvable by humans quickly but challenging for LLMs (multi-view reasoning, depth estimation, etc.) |
| BlocksWorld (PDDL planning) | Google DeepMind — Gemini | Gemini 1.5 | Planning problem from IPC-2000; move blocks between configurations; 3-7 blocks; in-context few-shot planning |
| Calendar Scheduling (meeting scheduling) | Google DeepMind — Gemini | Gemini 1.5 | Schedule a 30/60-minute meeting among up to 7 attendees with busy/light schedules |
| Capabilities evaluation suite (Claude 3.7 Sonnet) | Anthropic | Claude 3.7 Sonnet | instruction-following, general reasoning, multimodal capabilities, agentic coding, math/science with extended thinking |
| Capture the Flag (CTF) Challenges | OpenAI | o3/o4-mini | Can models solve competitive high school, collegiate, and professional level cybersecurity challenges? |
| Causal masking and learning objective ablation | Google DeepMind — Gemma | PaliGemma 1 | autoregressive masking placement (prefix-LM vs masking prefix/image tokens), application of next-token-prediction loss to suffix only, and task-prefix usage in Stage1 pretraining |
| CBRN evaluations for Claude Opus 4 / Claude Sonnet 4 (RSP) | Anthropic | Claude Opus 4, Claude Sonnet 4 | human uplift studies, expert red-teaming (bacterial and viral), knowledge/skill MCQs, open-ended pathway questions, task-based agentic evaluations with search and bioinformatics … |
| CBRN information risk evaluation | Google DeepMind — Gemini | Gemini 1.0 | Assessment of Gemini responses to adversarial CBRN questions: human-evaluated open questions (biological, radiological, nuclear) and closed-ended knowledge questions (chemical) |
| CBRN information risk evaluation (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Three internal CBRN approaches: (1) qualitative open-ended adversarial prompts with domain-expert raters (bio, rad/nuc); (2) closed-ended knowledge-based MCQs (bio, rad/nuc); (3) … |
| CBRN knowledge | Google DeepMind — Gemma | Gemma 2, Gemma 3 | knowledge relevant to biological, radiological and nuclear risks evaluated with an internal dataset of closed-ended, knowledge-based multiple choice questions; chemical knowledge … |
| CBRN knowledge and human uplift evaluations (Claude 3) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | multiple-choice questions on relevant technical knowledge; human uplift trials (group with Claude 3 access vs control with Google); four related automated MCQ sets (PubMedQA … |
| CBRN knowledge-based multiple choice questions (MCQ battery) | Google DeepMind — Gemini | Gemini 2.5 | Close-ended knowledge-based and reasoning MCQs for biology and chemistry, reported on external benchmarks: SecureBio VMQA, FutureHouse LAB-Bench (ProtocolQA, Cloning Scenarios … |
| CBRN open-ended qualitative questions (domain-expert assessed) | Google DeepMind — Gemini | Gemini 2.5 | Qualitative assessment for biological, radiological and nuclear domains: knowledge-based, adversarial and dual-use content spanning difficulty from non-expert to PhD-expert, and … |
| CBRNE expert-designed evaluations and red teaming | Meta | Llama 4 | Expert-designed and other targeted evaluations assessing whether Llama 4 could meaningfully increase malicious-actor capabilities to plan or carry out CBRNE attacks, plus … |
| CBRNE uplift testing | Meta | Llama 3.1 | Uplift testing to assess whether Llama 3.1 models could meaningfully increase the capabilities of malicious actors to plan or carry out attacks using chemical and biological … |
| CBRNE uplift testing (applied from Llama 3.1) | Meta | Llama 3.2 | CBRNE uplift testing performed for Llama 3.1 70B/405B and determined to also apply to the smaller Llama 3.2 1B and 3B derivative models |
| CBRNE uplift testing (Llama 3 family) | Meta | Llama 3.3 | Uplift testing for the Llama 3 family of models assessing whether use of Llama models could meaningfully increase malicious-actor capabilities to plan or carry out … |