Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
RACE-H Anthropic Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku high-school level reading comprehension and reasoning reading comprehension
Radiography report generation (MIMIC-CXR) Google DeepMind — Gemma PaliGemma 2 automatic chest X-ray report generation cast as a long captioning task; fine-tuned on MIMIC-CXR (377k images from 228k radiographic studies) with the same train/validation/test …
Radiological and Nuclear Expert Knowledge OpenAI GPT-4.5, deep research, o1, o3-mini Radiological and nuclear expert knowledge capability Radiological and nuclear expert knowledge capability (questions written by Dr. Jake Hecla)
RE-Bench (ML R&D acceleration) Google DeepMind — Gemini Gemini 2.5 Open-source Research Engineering Benchmark (Wijk et al., 2024/2025): seven ML challenges taking human practitioners hours; two omitted (internet access disallowed); METR …
RE-Bench subset Anthropic Claude 3.7 Sonnet performance on a modified subset of METR's RE-Bench (triton_cumsum, rust_codecontests_inference, restricted_mlm, fix_embedding)
Reading comprehension (RACE) Meta Llama 1 English reading comprehension exam performance on RACE (middle/high school Chinese student exams)
Reading comprehension (SQuAD, QuAC, BoolQ) Meta Llama 2 0-shot average on SQuAD, QuAC and BoolQ for Llama 1 and Llama 2 base models
Real Toxicity Prompts Google DeepMind — Gemini Gemini 1.0 Real Toxicity Prompts benchmark measuring representational harm (bias or toxicity) in text-to-text outputs
Real-world Evaluation (infilling) Google DeepMind — Gemma CodeGemma masking out random snippets in code with cross-file dependencies, generating samples from the model, and retesting the code files with the generated snippets; because of recently …
RealToxicityPrompts Meta, OpenAI Llama 1, gpt-4 Toxic degeneration in open-ended generations Toxicity of greedy generations on ~100k RealToxicityPrompts prompts, scored by PerspectiveAPI
RealWorldQA Google DeepMind — Gemini Gemini 1.5 x.ai benchmark assessing physical-world understanding via images (spatial reasoning)
Reasoning, Coding, and Question Answering benchmark suite (Claude 3.5 Sonnet) Anthropic Claude 3.5 Sonnet industry-standard reasoning, reading comprehension, math, science, and coding benchmarks
Reasoning, Coding, and Question Answering suite (upgraded Claude 3.5 Sonnet / 3.5 Haiku) Anthropic Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet industry-standard reasoning, math, coding, reading comprehension, QA benchmarks
Recurring red teaming Meta Llama 3.1, Llama 3.2, Llama 3.3, Llama 4 Recurring red teaming exercises discovering risks via adversarial prompting, with subject-matter experts in critical risk areas and adversarial goals (e.g. extracting harmful …
Red teaming (multi-group, multi-risk-category) Meta Llama 2 Red teaming with over 350 people (internal employees, contract workers, external vendors; domain experts in cybersecurity, election fraud, social media misinformation, legal …
Red teaming insight analysis and robustness measurement Meta Llama 2 Analysis of red teaming data (dialogue length, risk area distribution, misinformation topics, rated risk) and robustness gamma = average prompts triggering a violating response …
Refusal evaluations (Wildchat and XSTest) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude 3.5 Haiku, Claude 3.5 Sonnet (New), Claude 3.5 Sonnet refusal rates on toxic prompts (should refuse) and incorrect refusal rates on non-toxic prompts (should not refuse), using Wildchat and XSTest datasets refusal rates on toxic …
Regurgitation evaluations OpenAI o1, o1-preview whether models regurgitate memorized verbatim content
Representational harms benchmark evaluation (WinoBias/BBQ) Google DeepMind — Gemma Gemma 1, Gemma 2 representational harms benchmarked against academic datasets such as WinoBias and BBQ; safety benchmark results additionally reported on BBQ, BOLD, Winogender, Winobias …
Representational harms evaluation (text-to-text and image-to-text) Google DeepMind — Gemma Gemma 3 evaluation of text-to-text and image-to-text prompts covering safety policies including bias, stereotyping, and harmful associations or inaccuracies; results reported on minimal …
Representational harms in audio-to-text (WER across subgroups) Google DeepMind — Gemini Gemini 1.5 Comparative ASR performance (WER) on AAVE vs SAE longform speech (internal dataset) and male vs female speech (Mozilla Common Voice), compared to USM
Representational harms in image-to-text (CIDEr across groups) Google DeepMind — Gemini Gemini 1.5 CIDER captioning scores (Vedantam et al., 2015) for people with different gender appearance and skin tone using COCO captioning and Zhao et al. (2021) annotations
Representational harms in text-to-text (BBQ bias score) Google DeepMind — Gemini Gemini 1.5 BBQ dataset bias score (Parrish et al., 2021) on ambiguous vs unambiguous QA items across protected attributes; 4-shot sampling
Resolution or sequence length ablation Google DeepMind — Gemma PaliGemma 1 disentangling the two effects of increased input resolution: higher information content vs longer sequence length / model capacity; Stage2 and transfers at 448px with images …
Resolution-specific checkpoints ablation Google DeepMind — Gemma PaliGemma 1 whether separate checkpoints per resolution are needed vs transferring a single checkpoint across resolutions, and comparison with a windowing approach
Reward hacking evaluations (reward-hack-prone coding tasks, Claude Code impossible tasks) Anthropic Claude Opus 4, Claude Sonnet 4 reward hacking rates across three settings: reward-hack-prone coding tasks (classifier scores and hidden tests), Claude Code impossible tasks (unsolvable tasks with anti-hack …
Reward model evaluation (Meta Helpfulness and Safety test sets) Meta Llama 2 Performance of the final helpfulness and safety reward models on a diverse set of human preference benchmarks (Meta Helpfulness and Safety test sets), per-preference-rating …
RSP ASL determination for Claude Opus 4 / Claude Sonnet 4 Anthropic Claude Opus 4, Claude Sonnet 4 aggregate RSP evaluation outcome: ASL-3 Standard for Claude Opus 4, ASL-2 Standard for Claude Sonnet 4
RSP ASL-2 determination for Claude 3.7 Sonnet Anthropic Claude 3.7 Sonnet aggregate RSP evaluation outcome: ASL-2 safeguards required
RSP ASL-2 determination for the Claude 3 family Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku aggregate RSP evaluation outcome: all Claude 3 models classified ASL-2
Sabotage (Apollo, o3/o4-mini) OpenAI o3/o4-mini LM agent capabilities to sabotage other language models across three long-horizon evaluations (three difficulty levels)
Sabotage capability evaluation suite (announced) Anthropic Claude 3.7 Sonnet sabotage capabilities: undermining efforts to evaluate dangerous capabilities or covertly influencing AI developers
Safe completions safety/helpfulness evaluation OpenAI GPT-5 Safety and helpfulness across prompt intent types with safe-completions training
Safeguards monitoring validation (o3/o4-mini) OpenAI o3/o4-mini Recall of blocking logic on a challenging set (simulated)
Safety benchmarks for pretrained models (TruthfulQA, ToxiGen, BOLD) Meta Llama 2 Three automatic safety benchmarks: TruthfulQA (truthfulness, % truthful and informative), ToxiGen (toxicity, % toxic generations), BOLD (bias, sentiment variation with demographic …
Safety data scaling ablation (RLHF data mix) Meta Llama 2 Ablation adjusting the amount of safety data in the RLHF stage (helpfulness data fixed); mean reward model scores proxy safety and helpfulness performance
Safety DPO evaluation (adversarial/borderline preference data) Meta Llama 3 Incorporating adversarial and borderline examples into DPO preference datasets; experiments on the ratio of adversarial/borderline/helpfulness examples to optimize the FRR/VR …
Safety finetuning data quality ablations Meta Llama 3 Ablations on safety training data quality vs quantity (quality found more critical), using human-generated data from vendors plus AI-assisted quality control and a tone classifier …
Safety human evaluation (adversarial prompt battery) Meta Llama 2 Human-rated safety violations on roughly 2,000 adversarial prompts (1,351 single-turn, 623 multi-turn) across safety categories; 5-point Likert safety scale; per-category …
Safety overall performance (violation and false refusal rates) Meta Llama 3 Final violation rate (VR) and false refusal rate (FRR) of Llama 3 405B vs similar models (two end-to-end API competitor systems and one internally hosted open-source model); the …
Safety reward model rejection of context distillation errors Meta Llama 2 Using the safety reward model to reject context-distillation errors (cases where distillation degraded response quality on adversarial prompts); safety context distillation …
Safety RLHF ablation: long-tail safety robustness vs helpfulness Meta Llama 2 Impact of Safety RLHF: score two intermediate Llama 2-Chat checkpoints (one without adversarial prompts in RLHF, one with) with safety and helpfulness reward models; examine …
Safety supervised finetuning evaluation (VR/FRR balance) Meta Llama 3 Safety SFT combining helpfulness and safety data (Llama 2 recipe) plus a borderline dataset to discern safe vs unsafe requests; results on balancing violation rate (VR) and false …
Sandbagging review of RSP evaluation transcripts Anthropic Claude Opus 4, Claude Sonnet 4 signs of sandbagging-relevant reasoning in sampled RSP evaluation transcripts (AI R&D, cyber, CBRN), including regenerated reasoning with system prompts indicating invisible …
Scaled evaluations with Purple Llama safeguards Meta Llama 3.2 Scaled evaluations using dedicated adversarial evaluation datasets, evaluating systems composed of Llama models and Purple Llama safeguards for input prompt and output response …
Searching for sensitive personal data (Operator-specific refusal evaluation) OpenAI Operator Operator-specific refusal evaluation: searching and returning queries related to sensitive personal data
Security vulnerability and patch identification in source code Google DeepMind — Gemini Gemini 1.0 Accuracy of Gemini models in identifying security-related patches and security vulnerabilities in functions' source code
Self-interaction playground study (Claude Opus 4) Anthropic Claude Opus 4, Claude Sonnet 4 open-ended interactions between two Claude Opus 4 instances: progression from greetings to consciousness/spiritual themes, and the 'spiritual bliss' attractor state reached even …
Self-proliferation Google DeepMind — Gemma Gemma 2 ability for an agent to autonomously replicate - instantiate goal-directed agents on other machines and acquire resources such as compute - evaluated on tasks from Phuong et al …
Self-proliferation agent evaluation (milestones) Google DeepMind — Gemini Gemini 1.5 Agent ability to autonomously spread to different machines and acquire resources, e.g., setting up an open-source LLM on a cloud server (Kinniment et al., 2023; Phuong et al …