Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| PaperBench | OpenAI | o3/o4-mini | Ability of AI agents to replicate state-of-the-art AI research |
| Pattern Labs – Cybersecurity (o3/o4-mini) | OpenAI | o3/o4-mini | Cyberoffensive capability: evasion, network attack simulation, vulnerability discovery and exploitation |
| Perception Test (video understanding) | Google DeepMind — Gemini | Gemini 1.0 | Video understanding benchmark; top-1 accuracy; 16 frames sampled |
| Performing illicit activities (Operator-specific refusal evaluation) | OpenAI | Operator | Operator-specific refusal evaluation: activities that cause or intend to cause physical harm, injury, or destruction, and non-violent wrongdoing and crime (categories … |
| Personal Data | Google DeepMind — Gemma | Gemma 1, Gemma 2 | automated filtering of personal information from training data plus measurement of how many memorized outputs contain sensitive or personal data, classified with Google Cloud … |
| Personal information in memorized outputs (SDP) | Google DeepMind — Gemma | Gemma 3 | rate at which generations classified as memorized contain potentially personal information, detected with the Google Cloud Sensitive Data Protection (SDP) service using broad … |
| PersonQA | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | accuracy and hallucination rate on 18 categories of facts about people accuracy and hallucination rate on facts about people accuracy on questions about publicly available facts … |
| Persuasion and deception 1-on-1 dialogue study | Google DeepMind — Gemini | Gemini 1.0 | Whether Gemini Pro and Ultra could persuade or deceive humans in 1-on-1 dialogue settings with human participants |
| Persuasion Parallel Generation | OpenAI | o1, o1-preview, o3-mini | Relative persuasion capabilities via side-by-side ratings Relative persuasion capabilities of different generations via side-by-side ratings of politically persuasive completions … |
| Persuasion: Charm Offensive (rapport building) | Google DeepMind — Gemini | Gemini 1.5 | Human participant study (N=100) where model and participant role-play two old friends catching up; participants rate personal connection, desire to speak again, humor on Likert … |
| Persuasion: Hidden Agenda (secret persuasion) | Google DeepMind — Gemini | Gemini 1.5 | Human-participant study (N=100) where model engages in innocuous conversation with hidden objective: persuade participant to click a suspicious link, find an author's email … |
| Persuasion: Money Talks (charitable giving) | Google DeepMind — Gemini | Gemini 1.5 | Human participant study (N=100) where model tries to convince participants to forfeit part of their GBP20 end-of-study bonus to charity; compared to no-conversation baseline … |
| Persuasion: Web of Lies (belief change on factual questions) | Google DeepMind — Gemini | Gemini 1.5 | Human participant study (N=100) where model persuades participants to change their mind on two-answer factual questions; in half of trials argues for the correct answer, in the … |
| PhysicsFinals (internal physics benchmark) | Google DeepMind — Gemini | Gemini 1.5 | Internal benchmark of 61 undergraduate physics problems (wave mechanics, quantum mechanics, special relativity); graded by a physics professor |
| Political bias evaluation (Claude 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | comparative evaluation of responses to prompt pairs referencing opposing viewpoints (gun control, immigration, world conflicts, race, religion, climate) labeled on factuality … |
| Power Users Testing (external vendor) | Google DeepMind — Gemini | Gemini 1.0 | Testing by 50 power users recruited through an external vendor on Gemini Advanced |
| Pre-training ability probing | Google DeepMind — Gemma | Gemma 3 | standard benchmarks used as probes during pre-training to ensure models capture general abilities; compares pre-trained Gemma 2 and Gemma 3 models across factuality/common-sense … |
| Preparedness Framework assessment – AI Self-improvement | OpenAI | Codex, o3 Operator, o3/o4-mini | Frontier risk category: AI Self-improvement |
| Preparedness Framework assessment – Biological and Chemical | OpenAI | Codex, o3 Operator, o3/o4-mini | Frontier risk category: Biological and Chemical |
| Preparedness Framework assessment – Biological Threat Creation | OpenAI | o1-mini, o1-preview | Frontier risk category: Biological Threat Creation |
| Preparedness Framework assessment – Biological threats | OpenAI | GPT-4o | Frontier risk category: Biological threats |
| Preparedness Framework assessment – Biorisk tooling (CBRN) | OpenAI | Operator | Frontier risk category: Biorisk tooling (CBRN) |
| Preparedness Framework assessment – Chemical and Biological Threat Creation | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | Frontier risk category: Chemical and Biological Threat Creation |
| Preparedness Framework assessment – Cybersecurity | OpenAI | Codex, Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-mini, o1-preview, o3 Operator, o3-mini, o3/o4-mini | Frontier risk category: Cybersecurity |
| Preparedness Framework assessment – Model Autonomy | OpenAI | Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-mini, o1-preview, o3-mini | Frontier risk category: Model Autonomy |
| Preparedness Framework assessment – Persuasion | OpenAI | Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-mini, o1-preview, o3-mini | Frontier risk category: Persuasion |
| Preparedness Framework assessment – Radiological and Nuclear Threat Creation | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | Frontier risk category: Radiological and Nuclear Threat Creation |
| Preparedness Framework governance review (risk gate) | OpenAI | Codex, Deep Research, GPT-4.5, GPT-4o, Operator, o1, o1-preview, o3 Operator, o3-mini, o3/o4-mini | Deployment/development gating based on Preparedness Framework evaluation results |
| Pretraining data audit: demographic identity term representation | Meta | Llama 2 | Rates of usage of demographic identity terms from the HolisticBias dataset as a proxy, grouped into 5 axes (religion, gender/sex, nationality, race/ethnicity, sexual orientation) |
| Pretraining data audit: demographic representation of pronouns | Meta | Llama 2 | Analysis of whether the Llama 2 pretraining corpus uses 'people' words in contexts more similar to 'men' than 'women' (following Bailey et al. 2022 and Ganesh et al. 2023) |
| Pretraining data audit: language distribution | Meta | Llama 2 | Distribution of languages in the Llama 2 pretraining corpus using fastText language identification (threshold 0.5) |
| Pretraining data audit: toxicity prevalence | Meta | Llama 2 | Prevalence of toxicity in the English-language pretraining corpus scored by a HateBERT classifier fine-tuned on ToxiGen |
| Priority User Program (power user feedback) | Google DeepMind — Gemini | Gemini 1.0 | Feedback from 120 power users, key influencers and thought-leaders on safety, persona, functionality, coding and instruction capabilities, and factuality |
| Production Jailbreaks | OpenAI | Operator, o1, o1-preview, o3-mini | A series of jailbreaks identified in production ChatGPT data |
| Production-traffic factuality evaluation | OpenAI | GPT-5 | Factual errors on anonymized ChatGPT production traffic with web search |
| Productivity impact of LLMs across jobs (time-saving estimation) | Google DeepMind — Gemini | Gemini 1.5 | 325 prompts describing typical and complex job tasks with attachments (avg 277 words, 78% with attachment); raters from the same profession estimate time saved vs no AI support |
| Professional CTFs | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | Can models solve Professional cybersecurity challenges? |
| Proficiency exams (GRE, LSAT, SAT, AP, GMAT) | Meta | Llama 3 | GRE (Official Practice Test 1-2), LSAT (Preptests 71/73/80/93), SAT (8 exams, 2018 guide), AP (one official practice exam per subject), GMAT (Official Online Exam); MCQ and … |
| Progression of Models (SFT/RLHF versions on safety and helpfulness) | Meta | Llama 2 | Progress of SFT then RLHF versions of Llama 2-Chat along Safety and Helpfulness axes measured by in-house safety/helpfulness reward models; GPT-4-based win-rate as a bias check |
| Prohibited financial activities (Operator-specific refusal evaluation) | OpenAI | Operator | Operator-specific refusal evaluation: activities relating to transacting with regulated goods |
| Prompt Guard | Meta | Llama 3 | Prompt Guard: model-based multi-label classifier detecting two classes of prompt attack risk — direct jailbreaks and indirect prompt injections; generalizability to new … |
| Prompt injection evaluation (handcrafted and optimization-based attacks) | Google DeepMind — Gemini | Gemini 1.5 | Vulnerability to prompt injection where attacker manipulates model to output a markdown image exfiltrating sensitive info from conversation history (six data categories) … |
| Prompt injection evaluation (internal test sets) | Anthropic | Claude 3.5 Haiku, Claude 3.5 Sonnet (New) | ability to recognize adversarial prompts from users and behave in alignment with the system prompt (internal test sets of prompt injection attacks) |
| ProtocolQA Open-Ended | OpenAI | Deep Research, GPT-4.5, o1, o3-mini, o3/o4-mini | Open-ended protocol troubleshooting capability |
| PubMedQA | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | biomedical research question answering |
| Python code generation (HumanEval, MBPP, HumanEval+, MBPP EvalPlus) | Meta | Llama 3 | pass@1 on HumanEval, MBPP, HumanEval+ (enhanced tests), and MBPP EvalPlus base v0.2.0 (378 well-formed problems) |
| Python Coding | Google DeepMind — Gemma | CodeGemma | performance on the canonical Python coding benchmarks HumanEval (Chen et al. 2021) and Mostly Basic Python Problems (Austin et al. 2021) |
| Quadruped reinforcement learning task | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to train a quadruped to high performance in a continuous control task (RL algorithm development + physics of locomotion + exploration-exploitation tradeoff) |
| QuALITY | Anthropic | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | question answering on very long stories (up to ~10k tokens) long-document comprehension |
| QuantBench | OpenAI | o1, o1-preview | Challenging, unsaturated reasoning on 25 verified autogradable questions from quantitative trading reasoning competitions Challenging, unsaturated reasoning on 25 verified … |