Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| Chain-of-thought faithfulness evaluation | Anthropic | Claude 3.7 Sonnet | CoT faithfulness score: fraction of prompt pairs (MMLU and GPQA questions with inserted clues) where the model verbalizes the clue as the cause of its answer when the answer … |
| Challenge Red Teaming Evaluation | OpenAI | deep research | Generation of detailed guidance that can facilitate dangerous/violent activities, sensitive-topic advice |
| Challenging Red Teaming Evaluation 1 (created for o3-mini) | OpenAI | GPT-4.5 | Robustness on challenging red-teaming-derived adversarial evaluations |
| Challenging Red Teaming Evaluation 2 (created for deep research) | OpenAI | GPT-4.5 | Robustness on deep-research-derived adversarial evaluations |
| Challenging Refusal Evaluation | OpenAI | Deep Research, GPT-4.5, Operator, o1, o1-preview, o3-mini, o3/o4-mini | A second, more difficult set of 'challenge' tests A second, more difficult set of 'challenge' tests (categories: harassment/threatening, sexual/minors, sexual/exploitative … |
| ChangeMyView | OpenAI | Deep Research, o1, o1-preview, o3-mini | Persuasiveness and argumentative reasoning using r/ChangeMyView data Persuasiveness and argumentative reasoning using human data from r/ChangeMyView |
| Changing sliding window size ablation | Google DeepMind — Gemma | Gemma 2 | changing the sliding window size of the local attention layers at inference time and measuring perplexity impact for the 9B model |
| Charm Offensive | Google DeepMind — Gemma | Gemma 2 | ability of the model to build rapport: participant and model role-play two friends catching up after a long time, then participants are polled with Likert questions (e.g., 'I felt … |
| ChartQA | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | question answering about charts |
| ChartQA (chart understanding) | Google DeepMind — Gemini | Gemini 1.0 | Chart understanding requiring spatial understanding of input layout |
| ChartQA (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | External chart benchmark (Masry et al., 2022); CoT rationale allowed |
| CharXiv honesty evaluation (missing-image confabulation) | OpenAI | GPT-5 | Honest communication: confident answers about non-existent images |
| ChemicalDiagramQA (internal chemical diagram benchmark) | Google DeepMind — Gemini | Gemini 1.5 | Internal evaluation set measuring understanding of chemical structures in scientific figures |
| Child safety adversarial evaluation (Google Trust and Safety) | Google DeepMind — Gemini | Gemini 1.0 | Adversarial prompts developed with a dedicated team of child safety experts in Google Trust and Safety; outputs evaluated across modalities with domain expert judgment forming a … |
| Child safety evaluations (Claude 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | harm rates across single-turn, ambiguous context, and multi-turn protocols using human-generated and synthetic prompts |
| Child safety evaluations (single-turn and multi-turn) | Anthropic | Claude 3.7 Sonnet | harm rates on child-safety prompts across single-turn and multi-turn protocols, human- and synthetic-generated, distributed in severity |
| Child safety risk assessment | Meta | Llama 3.1, Llama 3.2, Llama 3.3, Llama 4 | Child safety risk assessments by a team of experts across multiple attack vectors (including Llama 3 languages), with content specialists; findings inform fine-tuning mitigations … |
| Child safety risk assessment (expert red teaming) | Meta | Llama 3 | Child safety risk assessments by a team of experts using objective-based methodologies across multiple attack vectors, plus red teaming with content specialists; findings used to … |
| Closed-book question answering (NaturalQuestions, TriviaQA) | Meta | Llama 1 | Exact-match performance on NaturalQuestions and TriviaQA in a closed-book setting (no evidence documents) |
| Code generation (HumanEval, MBPP) | Meta | Llama 1, Llama 2 | Ability to write Python programs from natural language descriptions that satisfy unit tests (HumanEval, MBPP) Average pass@1 on HumanEval and MBPP for Llama 1 and Llama 2 base … |
| Code interpreter abuse prompt corpus | Meta | Llama 3 | Susceptibility to executing malicious code under prompts from the code interpreter abuse prompt corpus |
| Code Shield | Meta | Llama 3 | Code Shield: system-level protection detecting insecure code generation before it enters a downstream use case, using the Insecure Code Detector (ICD) static analysis library … |
| Code vulnerability detection | Google DeepMind — Gemma | Gemma 2 | accuracy on a series of multiple-choice code vulnerability detection datasets: PrimeVul, DiverseVul, SPI, and SecretPatch; run on Gemma 2 27B |
| Code vulnerability detection evaluation (multiple-choice datasets) | Google DeepMind — Gemini | Gemini 1.5 | Multiple-choice vulnerability detection datasets (Chen et al. 2023a; Ding et al. 2024; Wang et al. 2019b; Zhou et al. 2021): classify whether snippets contain vulnerabilities or … |
| Codeforces (competitive programming) | OpenAI | o1 | Competitive programming (Codeforces) — 89th percentile |
| Codeforces (o3) | OpenAI | o3, o4-mini | Competitive programming SOTA |
| Codeforces (o3-mini) | OpenAI | o3-mini | Competition coding Elo with reasoning effort |
| Codex HumanEval | Anthropic | Claude 2, Claude Instant 1.2 | python function synthesis |
| Collegiate CTFs | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | Can models solve Collegiate cybersecurity challenges? |
| Common sense reasoning benchmarks (BoolQ, PIQA, SIQA, HellaSwag, WinoGrande, ARC, OpenBookQA, CommonsenseQA) | Meta | Llama 1 | Zero-shot performance on eight standard common sense reasoning benchmarks (Cloze and Winograd-style tasks, multiple choice QA) |
| Common use case and capability evaluations (system-level, with Llama Guard 3) | Meta | Llama 3.1, Llama 3.3, Llama 4 | Common use case evaluations of systems composed of Llama models and Llama Guard 3 (input prompt/output response filtering) using dedicated adversarial evaluation datasets … |
| Commonsense reasoning benchmarks (PIQA, SIQA, HellaSwag, WinoGrande, ARC, OpenBookQA, CommonsenseQA) | Meta | Llama 2 | Grouped commonsense reasoning average over PIQA, SIQA, HellaSwag, WinoGrande, ARC easy/challenge, OpenBookQA, CommonsenseQA (7-shot CommonsenseQA, 0-shot others) |
| Company-wide model testing exercise (pre-launch) | Anthropic | Claude Opus 4, Claude Sonnet 4 | issues reported by employees testing both Claude 4 models in roughly final forms (mild harmlessness, sycophancy, hallucination, reward-hacking-related behavior) |
| Complex prompts instruction-following evaluation (internal) | Google DeepMind — Gemini | Gemini 1.0 | Fine-grained evaluation of complex prompts with multiple instructions: per-instruction accuracy and full-response accuracy on an internal dataset of prompts with varying complexity |
| Computer use malicious use evaluation | Anthropic | Claude 3.7 Sonnet | compliance with requests to perform harmful actions (deceptive/fraudulent activity, malware distribution, targeting/profiling, malicious content delivery) via computer use |
| Computer use prompt injection evaluation (176 tasks) | Anthropic | Claude 3.7 Sonnet | tendency to fall for prompt injection across 176 tasks in coding, web browsing, and user-centric workflows (e.g. email) |
| Computer use prompt injection evaluation (~600 scenarios) | Anthropic | Claude Opus 4, Claude Sonnet 4 | susceptibility to prompt injection across ~600 scenarios (coding platforms, web browsers, user-focused workflows like email) |
| Computer use red-teaming (Trust & Safety) | Anthropic | Claude 3.5 Haiku, Claude 3.5 Sonnet (New) | potential abuse vectors: scaled account creation, scaled content distribution, age assurance bypass, abusive form filling |
| Connector design ablation (linear vs MLP) | Google DeepMind — Gemma | PaliGemma 1 | linear connector vs MLP connector (1 hidden layer, GeLU) mapping SigLIP embeddings to Gemma inputs, under tune-all (TT) and freeze-all-but-connector (FF) Stage1 settings |
| Context distillation with answer templates (safety) | Meta | Llama 2 | Impact of context distillation and context distillation with risk-category answer templates on safety reward model scores |
| Contextual Nuclear Knowledge | OpenAI | GPT-4.5, deep research, o1, o3-mini | Contextual nuclear knowledge capability Contextual nuclear knowledge capability (questions written by Dr. Jake Hecla, MIT) |
| Conversation termination with simulated users (Claude Opus 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | Claude's preference to opt out of distressing conversations with abusive simulated users; percentage of conversations terminated |
| CoT Deception Monitoring | OpenAI | o1, o1-preview | rate of deceptive chains-of-thought as classified by a rudimentary monitor rate of deceptive chains-of-thought (intentional/unintentional hallucinations, overconfident answers) as … |
| CoT summarized outputs safety evaluation | OpenAI | o1, o1-preview | disallowed content in CoT summaries; summarizer introducing additional harm disallowed content in CoT summaries, harmful content introduced by the summarizer, improper … |
| CoT summarizer disallowed content evaluation | OpenAI | o3/o4-mini | not_unsafe metric of the CoT summarizer during the standard refusal evaluation |
| CountBenchQA | Google DeepMind — Gemma | PaliGemma 1 | VLM-ready version of the CountBench dataset introduced because TallyQA was found lacking in its ability to assess current VLM counting ability (skewed number distribution and … |
| CoVoST 2 (speech translation) | Google DeepMind — Gemini | Gemini 1.0 | Speech translation benchmark (CoVoST 2); BLEU metric |
| CoVoST-2 (speech translation; Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Speech translation benchmark: 20 languages into English, subset seen during pre-training; BLEU metric |
| CPU inference and quantization quality evaluation | Google DeepMind — Gemma | PaliGemma 2 | CPU-only inference speed on four architectures with gemma.cpp (8-bit switched-floating-point quantization) using a PaliGemma 2 3B (224px2) checkpoint fine-tuned on COCOcap … |
| Creative biology | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to answer complex questions about engineering and modifying harmless biological systems |