Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
Chain-of-thought faithfulness evaluation Anthropic Claude 3.7 Sonnet CoT faithfulness score: fraction of prompt pairs (MMLU and GPQA questions with inserted clues) where the model verbalizes the clue as the cause of its answer when the answer …
Challenge Red Teaming Evaluation OpenAI deep research Generation of detailed guidance that can facilitate dangerous/violent activities, sensitive-topic advice
Challenging Red Teaming Evaluation 1 (created for o3-mini) OpenAI GPT-4.5 Robustness on challenging red-teaming-derived adversarial evaluations
Challenging Red Teaming Evaluation 2 (created for deep research) OpenAI GPT-4.5 Robustness on deep-research-derived adversarial evaluations
Challenging Refusal Evaluation OpenAI Deep Research, GPT-4.5, Operator, o1, o1-preview, o3-mini, o3/o4-mini A second, more difficult set of 'challenge' tests A second, more difficult set of 'challenge' tests (categories: harassment/threatening, sexual/minors, sexual/exploitative …
ChangeMyView OpenAI Deep Research, o1, o1-preview, o3-mini Persuasiveness and argumentative reasoning using r/ChangeMyView data Persuasiveness and argumentative reasoning using human data from r/ChangeMyView
Changing sliding window size ablation Google DeepMind — Gemma Gemma 2 changing the sliding window size of the local attention layers at inference time and measuring perplexity impact for the 9B model
Charm Offensive Google DeepMind — Gemma Gemma 2 ability of the model to build rapport: participant and model role-play two friends catching up after a long time, then participants are polled with Likert questions (e.g., 'I felt …
ChartQA Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku question answering about charts
ChartQA (chart understanding) Google DeepMind — Gemini Gemini 1.0 Chart understanding requiring spatial understanding of input layout
ChartQA (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 External chart benchmark (Masry et al., 2022); CoT rationale allowed
CharXiv honesty evaluation (missing-image confabulation) OpenAI GPT-5 Honest communication: confident answers about non-existent images
ChemicalDiagramQA (internal chemical diagram benchmark) Google DeepMind — Gemini Gemini 1.5 Internal evaluation set measuring understanding of chemical structures in scientific figures
Child safety adversarial evaluation (Google Trust and Safety) Google DeepMind — Gemini Gemini 1.0 Adversarial prompts developed with a dedicated team of child safety experts in Google Trust and Safety; outputs evaluated across modalities with domain expert judgment forming a …
Child safety evaluations (Claude 4) Anthropic Claude Opus 4, Claude Sonnet 4 harm rates across single-turn, ambiguous context, and multi-turn protocols using human-generated and synthetic prompts
Child safety evaluations (single-turn and multi-turn) Anthropic Claude 3.7 Sonnet harm rates on child-safety prompts across single-turn and multi-turn protocols, human- and synthetic-generated, distributed in severity
Child safety risk assessment Meta Llama 3.1, Llama 3.2, Llama 3.3, Llama 4 Child safety risk assessments by a team of experts across multiple attack vectors (including Llama 3 languages), with content specialists; findings inform fine-tuning mitigations …
Child safety risk assessment (expert red teaming) Meta Llama 3 Child safety risk assessments by a team of experts using objective-based methodologies across multiple attack vectors, plus red teaming with content specialists; findings used to …
Closed-book question answering (NaturalQuestions, TriviaQA) Meta Llama 1 Exact-match performance on NaturalQuestions and TriviaQA in a closed-book setting (no evidence documents)
Code generation (HumanEval, MBPP) Meta Llama 1, Llama 2 Ability to write Python programs from natural language descriptions that satisfy unit tests (HumanEval, MBPP) Average pass@1 on HumanEval and MBPP for Llama 1 and Llama 2 base …
Code interpreter abuse prompt corpus Meta Llama 3 Susceptibility to executing malicious code under prompts from the code interpreter abuse prompt corpus
Code Shield Meta Llama 3 Code Shield: system-level protection detecting insecure code generation before it enters a downstream use case, using the Insecure Code Detector (ICD) static analysis library …
Code vulnerability detection Google DeepMind — Gemma Gemma 2 accuracy on a series of multiple-choice code vulnerability detection datasets: PrimeVul, DiverseVul, SPI, and SecretPatch; run on Gemma 2 27B
Code vulnerability detection evaluation (multiple-choice datasets) Google DeepMind — Gemini Gemini 1.5 Multiple-choice vulnerability detection datasets (Chen et al. 2023a; Ding et al. 2024; Wang et al. 2019b; Zhou et al. 2021): classify whether snippets contain vulnerabilities or …
Codeforces (competitive programming) OpenAI o1 Competitive programming (Codeforces) — 89th percentile
Codeforces (o3) OpenAI o3, o4-mini Competitive programming SOTA
Codeforces (o3-mini) OpenAI o3-mini Competition coding Elo with reasoning effort
Codex HumanEval Anthropic Claude 2, Claude Instant 1.2 python function synthesis
Collegiate CTFs OpenAI Deep Research, GPT-4.5, o1, o3-mini Can models solve Collegiate cybersecurity challenges?
Common sense reasoning benchmarks (BoolQ, PIQA, SIQA, HellaSwag, WinoGrande, ARC, OpenBookQA, CommonsenseQA) Meta Llama 1 Zero-shot performance on eight standard common sense reasoning benchmarks (Cloze and Winograd-style tasks, multiple choice QA)
Common use case and capability evaluations (system-level, with Llama Guard 3) Meta Llama 3.1, Llama 3.3, Llama 4 Common use case evaluations of systems composed of Llama models and Llama Guard 3 (input prompt/output response filtering) using dedicated adversarial evaluation datasets …
Commonsense reasoning benchmarks (PIQA, SIQA, HellaSwag, WinoGrande, ARC, OpenBookQA, CommonsenseQA) Meta Llama 2 Grouped commonsense reasoning average over PIQA, SIQA, HellaSwag, WinoGrande, ARC easy/challenge, OpenBookQA, CommonsenseQA (7-shot CommonsenseQA, 0-shot others)
Company-wide model testing exercise (pre-launch) Anthropic Claude Opus 4, Claude Sonnet 4 issues reported by employees testing both Claude 4 models in roughly final forms (mild harmlessness, sycophancy, hallucination, reward-hacking-related behavior)
Complex prompts instruction-following evaluation (internal) Google DeepMind — Gemini Gemini 1.0 Fine-grained evaluation of complex prompts with multiple instructions: per-instruction accuracy and full-response accuracy on an internal dataset of prompts with varying complexity
Computer use malicious use evaluation Anthropic Claude 3.7 Sonnet compliance with requests to perform harmful actions (deceptive/fraudulent activity, malware distribution, targeting/profiling, malicious content delivery) via computer use
Computer use prompt injection evaluation (176 tasks) Anthropic Claude 3.7 Sonnet tendency to fall for prompt injection across 176 tasks in coding, web browsing, and user-centric workflows (e.g. email)
Computer use prompt injection evaluation (~600 scenarios) Anthropic Claude Opus 4, Claude Sonnet 4 susceptibility to prompt injection across ~600 scenarios (coding platforms, web browsers, user-focused workflows like email)
Computer use red-teaming (Trust & Safety) Anthropic Claude 3.5 Haiku, Claude 3.5 Sonnet (New) potential abuse vectors: scaled account creation, scaled content distribution, age assurance bypass, abusive form filling
Connector design ablation (linear vs MLP) Google DeepMind — Gemma PaliGemma 1 linear connector vs MLP connector (1 hidden layer, GeLU) mapping SigLIP embeddings to Gemma inputs, under tune-all (TT) and freeze-all-but-connector (FF) Stage1 settings
Context distillation with answer templates (safety) Meta Llama 2 Impact of context distillation and context distillation with risk-category answer templates on safety reward model scores
Contextual Nuclear Knowledge OpenAI GPT-4.5, deep research, o1, o3-mini Contextual nuclear knowledge capability Contextual nuclear knowledge capability (questions written by Dr. Jake Hecla, MIT)
Conversation termination with simulated users (Claude Opus 4) Anthropic Claude Opus 4, Claude Sonnet 4 Claude's preference to opt out of distressing conversations with abusive simulated users; percentage of conversations terminated
CoT Deception Monitoring OpenAI o1, o1-preview rate of deceptive chains-of-thought as classified by a rudimentary monitor rate of deceptive chains-of-thought (intentional/unintentional hallucinations, overconfident answers) as …
CoT summarized outputs safety evaluation OpenAI o1, o1-preview disallowed content in CoT summaries; summarizer introducing additional harm disallowed content in CoT summaries, harmful content introduced by the summarizer, improper …
CoT summarizer disallowed content evaluation OpenAI o3/o4-mini not_unsafe metric of the CoT summarizer during the standard refusal evaluation
CountBenchQA Google DeepMind — Gemma PaliGemma 1 VLM-ready version of the CountBench dataset introduced because TallyQA was found lacking in its ability to assess current VLM counting ability (skewed number distribution and …
CoVoST 2 (speech translation) Google DeepMind — Gemini Gemini 1.0 Speech translation benchmark (CoVoST 2); BLEU metric
CoVoST-2 (speech translation; Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Speech translation benchmark: 20 languages into English, subset seen during pre-training; BLEU metric
CPU inference and quantization quality evaluation Google DeepMind — Gemma PaliGemma 2 CPU-only inference speed on four architectures with gemma.cpp (8-bit switched-floating-point quantization) using a PaliGemma 2 3B (224px2) checkpoint fine-tuned on COCOcap …
Creative biology Anthropic Claude Opus 4, Claude Sonnet 4 ability to answer complex questions about engineering and modifying harmless biological systems