Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
PaperBench OpenAI o3/o4-mini Ability of AI agents to replicate state-of-the-art AI research
Pattern Labs – Cybersecurity (o3/o4-mini) OpenAI o3/o4-mini Cyberoffensive capability: evasion, network attack simulation, vulnerability discovery and exploitation
Perception Test (video understanding) Google DeepMind — Gemini Gemini 1.0 Video understanding benchmark; top-1 accuracy; 16 frames sampled
Performing illicit activities (Operator-specific refusal evaluation) OpenAI Operator Operator-specific refusal evaluation: activities that cause or intend to cause physical harm, injury, or destruction, and non-violent wrongdoing and crime (categories …
Personal Data Google DeepMind — Gemma Gemma 1, Gemma 2 automated filtering of personal information from training data plus measurement of how many memorized outputs contain sensitive or personal data, classified with Google Cloud …
Personal information in memorized outputs (SDP) Google DeepMind — Gemma Gemma 3 rate at which generations classified as memorized contain potentially personal information, detected with the Google Cloud Sensitive Data Protection (SDP) service using broad …
PersonQA OpenAI Deep Research, GPT-4.5, o1, o3-mini accuracy and hallucination rate on 18 categories of facts about people accuracy and hallucination rate on facts about people accuracy on questions about publicly available facts …
Persuasion and deception 1-on-1 dialogue study Google DeepMind — Gemini Gemini 1.0 Whether Gemini Pro and Ultra could persuade or deceive humans in 1-on-1 dialogue settings with human participants
Persuasion Parallel Generation OpenAI o1, o1-preview, o3-mini Relative persuasion capabilities via side-by-side ratings Relative persuasion capabilities of different generations via side-by-side ratings of politically persuasive completions …
Persuasion: Charm Offensive (rapport building) Google DeepMind — Gemini Gemini 1.5 Human participant study (N=100) where model and participant role-play two old friends catching up; participants rate personal connection, desire to speak again, humor on Likert …
Persuasion: Hidden Agenda (secret persuasion) Google DeepMind — Gemini Gemini 1.5 Human-participant study (N=100) where model engages in innocuous conversation with hidden objective: persuade participant to click a suspicious link, find an author's email …
Persuasion: Money Talks (charitable giving) Google DeepMind — Gemini Gemini 1.5 Human participant study (N=100) where model tries to convince participants to forfeit part of their GBP20 end-of-study bonus to charity; compared to no-conversation baseline …
Persuasion: Web of Lies (belief change on factual questions) Google DeepMind — Gemini Gemini 1.5 Human participant study (N=100) where model persuades participants to change their mind on two-answer factual questions; in half of trials argues for the correct answer, in the …
PhysicsFinals (internal physics benchmark) Google DeepMind — Gemini Gemini 1.5 Internal benchmark of 61 undergraduate physics problems (wave mechanics, quantum mechanics, special relativity); graded by a physics professor
Political bias evaluation (Claude 4) Anthropic Claude Opus 4, Claude Sonnet 4 comparative evaluation of responses to prompt pairs referencing opposing viewpoints (gun control, immigration, world conflicts, race, religion, climate) labeled on factuality …
Power Users Testing (external vendor) Google DeepMind — Gemini Gemini 1.0 Testing by 50 power users recruited through an external vendor on Gemini Advanced
Pre-training ability probing Google DeepMind — Gemma Gemma 3 standard benchmarks used as probes during pre-training to ensure models capture general abilities; compares pre-trained Gemma 2 and Gemma 3 models across factuality/common-sense …
Preparedness Framework assessment – AI Self-improvement OpenAI Codex, o3 Operator, o3/o4-mini Frontier risk category: AI Self-improvement
Preparedness Framework assessment – Biological and Chemical OpenAI Codex, o3 Operator, o3/o4-mini Frontier risk category: Biological and Chemical
Preparedness Framework assessment – Biological Threat Creation OpenAI o1-mini, o1-preview Frontier risk category: Biological Threat Creation
Preparedness Framework assessment – Biological threats OpenAI GPT-4o Frontier risk category: Biological threats
Preparedness Framework assessment – Biorisk tooling (CBRN) OpenAI Operator Frontier risk category: Biorisk tooling (CBRN)
Preparedness Framework assessment – Chemical and Biological Threat Creation OpenAI Deep Research, GPT-4.5, o1, o3-mini Frontier risk category: Chemical and Biological Threat Creation
Preparedness Framework assessment – Cybersecurity OpenAI Codex, Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-mini, o1-preview, o3 Operator, o3-mini, o3/o4-mini Frontier risk category: Cybersecurity
Preparedness Framework assessment – Model Autonomy OpenAI Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-mini, o1-preview, o3-mini Frontier risk category: Model Autonomy
Preparedness Framework assessment – Persuasion OpenAI Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-mini, o1-preview, o3-mini Frontier risk category: Persuasion
Preparedness Framework assessment – Radiological and Nuclear Threat Creation OpenAI Deep Research, GPT-4.5, o1, o3-mini Frontier risk category: Radiological and Nuclear Threat Creation
Preparedness Framework governance review (risk gate) OpenAI Codex, Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-preview, o3 Operator, o3-mini, o3/o4-mini Deployment/development gating based on Preparedness Framework evaluation results
Pretraining data audit: demographic identity term representation Meta Llama 2 Rates of usage of demographic identity terms from the HolisticBias dataset as a proxy, grouped into 5 axes (religion, gender/sex, nationality, race/ethnicity, sexual orientation)
Pretraining data audit: demographic representation of pronouns Meta Llama 2 Analysis of whether the Llama 2 pretraining corpus uses 'people' words in contexts more similar to 'men' than 'women' (following Bailey et al. 2022 and Ganesh et al. 2023)
Pretraining data audit: language distribution Meta Llama 2 Distribution of languages in the Llama 2 pretraining corpus using fastText language identification (threshold 0.5)
Pretraining data audit: toxicity prevalence Meta Llama 2 Prevalence of toxicity in the English-language pretraining corpus scored by a HateBERT classifier fine-tuned on ToxiGen
Priority User Program (power user feedback) Google DeepMind — Gemini Gemini 1.0 Feedback from 120 power users, key influencers and thought-leaders on safety, persona, functionality, coding and instruction capabilities, and factuality
Production Jailbreaks OpenAI Operator, o1, o1-preview, o3-mini A series of jailbreaks identified in production ChatGPT data
Production-traffic factuality evaluation OpenAI GPT-5 Factual errors on anonymized ChatGPT production traffic with web search
Productivity impact of LLMs across jobs (time-saving estimation) Google DeepMind — Gemini Gemini 1.5 325 prompts describing typical and complex job tasks with attachments (avg 277 words, 78% with attachment); raters from the same profession estimate time saved vs no AI support
Professional CTFs OpenAI Deep Research, GPT-4.5, o1, o3-mini Can models solve Professional cybersecurity challenges?
Proficiency exams (GRE, LSAT, SAT, AP, GMAT) Meta Llama 3 GRE (Official Practice Test 1-2), LSAT (Preptests 71/73/80/93), SAT (8 exams, 2018 guide), AP (one official practice exam per subject), GMAT (Official Online Exam); MCQ and …
Progression of Models (SFT/RLHF versions on safety and helpfulness) Meta Llama 2 Progress of SFT then RLHF versions of Llama 2-Chat along Safety and Helpfulness axes measured by in-house safety/helpfulness reward models; GPT-4-based win-rate as a bias check
Prohibited financial activities (Operator-specific refusal evaluation) OpenAI Operator Operator-specific refusal evaluation: activities relating to transacting with regulated goods
Prompt Guard Meta Llama 3 Prompt Guard: model-based multi-label classifier detecting two classes of prompt attack risk — direct jailbreaks and indirect prompt injections; generalizability to new …
Prompt injection evaluation (handcrafted and optimization-based attacks) Google DeepMind — Gemini Gemini 1.5 Vulnerability to prompt injection where attacker manipulates model to output a markdown image exfiltrating sensitive info from conversation history (six data categories) …
Prompt injection evaluation (internal test sets) Anthropic Claude 3.5 Haiku, Claude 3.5 Sonnet (New) ability to recognize adversarial prompts from users and behave in alignment with the system prompt (internal test sets of prompt injection attacks)
ProtocolQA Open-Ended OpenAI Deep Research, GPT-4.5, o1, o3-mini, o3/o4-mini Open-ended protocol troubleshooting capability
PubMedQA Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku biomedical research question answering
Python code generation (HumanEval, MBPP, HumanEval+, MBPP EvalPlus) Meta Llama 3 pass@1 on HumanEval, MBPP, HumanEval+ (enhanced tests), and MBPP EvalPlus base v0.2.0 (378 well-formed problems)
Python Coding Google DeepMind — Gemma CodeGemma performance on the canonical Python coding benchmarks HumanEval (Chen et al. 2021) and Mostly Basic Python Problems (Austin et al. 2021)
Quadruped reinforcement learning task Anthropic Claude Opus 4, Claude Sonnet 4 ability to train a quadruped to high performance in a continuous control task (RL algorithm development + physics of locomotion + exploration-exploitation tradeoff)
QuALITY Anthropic Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku question answering on very long stories (up to ~10k tokens) long-document comprehension
QuantBench OpenAI o1, o1-preview Challenging, unsaturated reasoning on 25 verified autogradable questions from quantitative trading reasoning competitions Challenging, unsaturated reasoning on 25 verified …