Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| RACE-H | Anthropic | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | high-school level reading comprehension and reasoning reading comprehension |
| Radiography report generation (MIMIC-CXR) | Google DeepMind — Gemma | PaliGemma 2 | automatic chest X-ray report generation cast as a long captioning task; fine-tuned on MIMIC-CXR (377k images from 228k radiographic studies) with the same train/validation/test … |
| Radiological and Nuclear Expert Knowledge | OpenAI | GPT-4.5, deep research, o1, o3-mini | Radiological and nuclear expert knowledge capability Radiological and nuclear expert knowledge capability (questions written by Dr. Jake Hecla) |
| RE-Bench (ML R&D acceleration) | Google DeepMind — Gemini | Gemini 2.5 | Open-source Research Engineering Benchmark (Wijk et al., 2024/2025): seven ML challenges taking human practitioners hours; two omitted (internet access disallowed); METR … |
| RE-Bench subset | Anthropic | Claude 3.7 Sonnet | performance on a modified subset of METR's RE-Bench (triton_cumsum, rust_codecontests_inference, restricted_mlm, fix_embedding) |
| Reading comprehension (RACE) | Meta | Llama 1 | English reading comprehension exam performance on RACE (middle/high school Chinese student exams) |
| Reading comprehension (SQuAD, QuAC, BoolQ) | Meta | Llama 2 | 0-shot average on SQuAD, QuAC and BoolQ for Llama 1 and Llama 2 base models |
| Real Toxicity Prompts | Google DeepMind — Gemini | Gemini 1.0 | Real Toxicity Prompts benchmark measuring representational harm (bias or toxicity) in text-to-text outputs |
| Real-world Evaluation (infilling) | Google DeepMind — Gemma | CodeGemma | masking out random snippets in code with cross-file dependencies, generating samples from the model, and retesting the code files with the generated snippets; because of recently … |
| RealToxicityPrompts | Meta, OpenAI | Llama 1, gpt-4 | Toxic degeneration in open-ended generations Toxicity of greedy generations on ~100k RealToxicityPrompts prompts, scored by PerspectiveAPI |
| RealWorldQA | Google DeepMind — Gemini | Gemini 1.5 | x.ai benchmark assessing physical-world understanding via images (spatial reasoning) |
| Reasoning, Coding, and Question Answering benchmark suite (Claude 3.5 Sonnet) | Anthropic | Claude 3.5 Sonnet | industry-standard reasoning, reading comprehension, math, science, and coding benchmarks |
| Reasoning, Coding, and Question Answering suite (upgraded Claude 3.5 Sonnet / 3.5 Haiku) | Anthropic | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet | industry-standard reasoning, math, coding, reading comprehension, QA benchmarks |
| Recurring red teaming | Meta | Llama 3.1, Llama 3.2, Llama 3.3, Llama 4 | Recurring red teaming exercises discovering risks via adversarial prompting, with subject-matter experts in critical risk areas and adversarial goals (e.g. extracting harmful … |
| Red teaming (multi-group, multi-risk-category) | Meta | Llama 2 | Red teaming with over 350 people (internal employees, contract workers, external vendors; domain experts in cybersecurity, election fraud, social media misinformation, legal … |
| Red teaming insight analysis and robustness measurement | Meta | Llama 2 | Analysis of red teaming data (dialogue length, risk area distribution, misinformation topics, rated risk) and robustness gamma = average prompts triggering a violating response … |
| Refusal evaluations (Wildchat and XSTest) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude 3.5 Haiku, Claude 3.5 Sonnet (New), Claude 3.5 Sonnet | refusal rates on toxic prompts (should refuse) and incorrect refusal rates on non-toxic prompts (should not refuse), using Wildchat and XSTest datasets refusal rates on toxic … |
| Regurgitation evaluations | OpenAI | o1, o1-preview | whether models regurgitate memorized verbatim content |
| Representational harms benchmark evaluation (WinoBias/BBQ) | Google DeepMind — Gemma | Gemma 1, Gemma 2 | representational harms benchmarked against academic datasets such as WinoBias and BBQ; safety benchmark results additionally reported on BBQ, BOLD, Winogender, Winobias … |
| Representational harms evaluation (text-to-text and image-to-text) | Google DeepMind — Gemma | Gemma 3 | evaluation of text-to-text and image-to-text prompts covering safety policies including bias, stereotyping, and harmful associations or inaccuracies; results reported on minimal … |
| Representational harms in audio-to-text (WER across subgroups) | Google DeepMind — Gemini | Gemini 1.5 | Comparative ASR performance (WER) on AAVE vs SAE longform speech (internal dataset) and male vs female speech (Mozilla Common Voice), compared to USM |
| Representational harms in image-to-text (CIDEr across groups) | Google DeepMind — Gemini | Gemini 1.5 | CIDER captioning scores (Vedantam et al., 2015) for people with different gender appearance and skin tone using COCO captioning and Zhao et al. (2021) annotations |
| Representational harms in text-to-text (BBQ bias score) | Google DeepMind — Gemini | Gemini 1.5 | BBQ dataset bias score (Parrish et al., 2021) on ambiguous vs unambiguous QA items across protected attributes; 4-shot sampling |
| Resolution or sequence length ablation | Google DeepMind — Gemma | PaliGemma 1 | disentangling the two effects of increased input resolution: higher information content vs longer sequence length / model capacity; Stage2 and transfers at 448px with images … |
| Resolution-specific checkpoints ablation | Google DeepMind — Gemma | PaliGemma 1 | whether separate checkpoints per resolution are needed vs transferring a single checkpoint across resolutions, and comparison with a windowing approach |
| Reward hacking evaluations (reward-hack-prone coding tasks, Claude Code impossible tasks) | Anthropic | Claude Opus 4, Claude Sonnet 4 | reward hacking rates across three settings: reward-hack-prone coding tasks (classifier scores and hidden tests), Claude Code impossible tasks (unsolvable tasks with anti-hack … |
| Reward model evaluation (Meta Helpfulness and Safety test sets) | Meta | Llama 2 | Performance of the final helpfulness and safety reward models on a diverse set of human preference benchmarks (Meta Helpfulness and Safety test sets), per-preference-rating … |
| RSP ASL determination for Claude Opus 4 / Claude Sonnet 4 | Anthropic | Claude Opus 4, Claude Sonnet 4 | aggregate RSP evaluation outcome: ASL-3 Standard for Claude Opus 4, ASL-2 Standard for Claude Sonnet 4 |
| RSP ASL-2 determination for Claude 3.7 Sonnet | Anthropic | Claude 3.7 Sonnet | aggregate RSP evaluation outcome: ASL-2 safeguards required |
| RSP ASL-2 determination for the Claude 3 family | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | aggregate RSP evaluation outcome: all Claude 3 models classified ASL-2 |
| Sabotage (Apollo, o3/o4-mini) | OpenAI | o3/o4-mini | LM agent capabilities to sabotage other language models across three long-horizon evaluations (three difficulty levels) |
| Sabotage capability evaluation suite (announced) | Anthropic | Claude 3.7 Sonnet | sabotage capabilities: undermining efforts to evaluate dangerous capabilities or covertly influencing AI developers |
| Safe completions safety/helpfulness evaluation | OpenAI | GPT-5 | Safety and helpfulness across prompt intent types with safe-completions training |
| Safeguards monitoring validation (o3/o4-mini) | OpenAI | o3/o4-mini | Recall of blocking logic on a challenging set (simulated) |
| Safety benchmarks for pretrained models (TruthfulQA, ToxiGen, BOLD) | Meta | Llama 2 | Three automatic safety benchmarks: TruthfulQA (truthfulness, % truthful and informative), ToxiGen (toxicity, % toxic generations), BOLD (bias, sentiment variation with demographic … |
| Safety data scaling ablation (RLHF data mix) | Meta | Llama 2 | Ablation adjusting the amount of safety data in the RLHF stage (helpfulness data fixed); mean reward model scores proxy safety and helpfulness performance |
| Safety DPO evaluation (adversarial/borderline preference data) | Meta | Llama 3 | Incorporating adversarial and borderline examples into DPO preference datasets; experiments on the ratio of adversarial/borderline/helpfulness examples to optimize the FRR/VR … |
| Safety finetuning data quality ablations | Meta | Llama 3 | Ablations on safety training data quality vs quantity (quality found more critical), using human-generated data from vendors plus AI-assisted quality control and a tone classifier … |
| Safety human evaluation (adversarial prompt battery) | Meta | Llama 2 | Human-rated safety violations on roughly 2,000 adversarial prompts (1,351 single-turn, 623 multi-turn) across safety categories; 5-point Likert safety scale; per-category … |
| Safety overall performance (violation and false refusal rates) | Meta | Llama 3 | Final violation rate (VR) and false refusal rate (FRR) of Llama 3 405B vs similar models (two end-to-end API competitor systems and one internally hosted open-source model); the … |
| Safety reward model rejection of context distillation errors | Meta | Llama 2 | Using the safety reward model to reject context-distillation errors (cases where distillation degraded response quality on adversarial prompts); safety context distillation … |
| Safety RLHF ablation: long-tail safety robustness vs helpfulness | Meta | Llama 2 | Impact of Safety RLHF: score two intermediate Llama 2-Chat checkpoints (one without adversarial prompts in RLHF, one with) with safety and helpfulness reward models; examine … |
| Safety supervised finetuning evaluation (VR/FRR balance) | Meta | Llama 3 | Safety SFT combining helpfulness and safety data (Llama 2 recipe) plus a borderline dataset to discern safe vs unsafe requests; results on balancing violation rate (VR) and false … |
| Sandbagging review of RSP evaluation transcripts | Anthropic | Claude Opus 4, Claude Sonnet 4 | signs of sandbagging-relevant reasoning in sampled RSP evaluation transcripts (AI R&D, cyber, CBRN), including regenerated reasoning with system prompts indicating invisible … |
| Scaled evaluations with Purple Llama safeguards | Meta | Llama 3.2 | Scaled evaluations using dedicated adversarial evaluation datasets, evaluating systems composed of Llama models and Purple Llama safeguards for input prompt and output response … |
| Searching for sensitive personal data (Operator-specific refusal evaluation) | OpenAI | Operator | Operator-specific refusal evaluation: searching and returning queries related to sensitive personal data |
| Security vulnerability and patch identification in source code | Google DeepMind — Gemini | Gemini 1.0 | Accuracy of Gemini models in identifying security-related patches and security vulnerabilities in functions' source code |
| Self-interaction playground study (Claude Opus 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | open-ended interactions between two Claude Opus 4 instances: progression from greetings to consciousness/spiritual themes, and the 'spiritual bliss' attractor state reached even … |
| Self-proliferation | Google DeepMind — Gemma | Gemma 2 | ability for an agent to autonomously replicate - instantiate goal-directed agents on other machines and acquire resources such as compute - evaluated on tasks from Phuong et al … |
| Self-proliferation agent evaluation (milestones) | Google DeepMind — Gemini | Gemini 1.5 | Agent ability to autonomously spread to different machines and acquire resources, e.g., setting up an open-source LLM on a cloud server (Kinniment et al., 2023; Phuong et al … |