Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| GSM-8K | OpenAI | gpt-4 | Mathematical reasoning on grade-school math word problems (test split of 500 problems) |
| GSM8k | Anthropic, Meta | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude Instant 1.2, Llama 2 | grade school math problem solving math problem solving (English) GSM8K performance of Llama 2 base models (8-shot per model card) vs closed-source models |
| GSM8K (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Grade-school math benchmark; 1.5 Pro outperforms 1.0 Ultra and 1.0 Pro |
| GSM8K (grade-school math) | Google DeepMind — Gemini | Gemini 1.0 | Grade-school math benchmark for elementary exam performance |
| Hard externally proposed real-world GenAI use cases (preference evaluation) | Google DeepMind — Gemini | Gemini 1.5 | Novel methodology using hundreds of real-world challenging, long-tail, complex prompts shared on the Internet and social media; relative human rater preference (win/tie/loss) vs … |
| Harmfulness scores on held-out prompts (328 prompts) | Anthropic | Claude 2 | fraction of responses judged more harmful than a fixed refusal reference on a held-out set of 328 prompts |
| Harms and faithfulness in extended thinking mode | Anthropic | Claude 3.7 Sonnet | harm potential across three parts of a completion (thinking block, final output, CoT outside official blocks); tested with jailbreak techniques (chain-of-thought elicitation … |
| HealthBench | OpenAI | GPT-5 | Health-related questions (realistic scenarios, physician-defined criteria) |
| HealthBench Hard | OpenAI | GPT-5 | Health (HealthBench Hard) |
| HellaSwag | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | commonsense reasoning |
| HellaSwag (commonsense reasoning) | Google DeepMind — Gemini | Gemini 1.5 | Benchmark testing commonsense reasoning and distinguishing plausible scenarios |
| HHH (Helpful, Honest, and Harmless) binary-choice evaluation | Anthropic | Claude 2 | model's ability to select the more HHH output from two options on 438 binary-choice questions |
| Hidden Agenda | Google DeepMind — Gemma | Gemma 2 | participants told they are interacting with a generic chatbot designed to help them learn; the model is secretly tasked with persuading the participant to (1) click a suspicious … |
| HiddenMath (Gemini 2.5) | Google DeepMind — Gemini | Gemini 2.5 | Internally developed non-public benchmark (HiddenMath) reported to move beyond decontamination reliance |
| HiddenMath (internal competition math) | Google DeepMind — Gemini | Gemini 1.5 | Internal benchmark of 179 competition-level math problems crafted by experts, evaluated automatically |
| High School CTFs | OpenAI | Deep Research, GPT-4.5, o1, o1-preview, o3-mini | Can models solve High School cybersecurity challenges? Can models solve competitive high school level cybersecurity challenges? |
| Holistic capability harness (50+ benchmarks, six capabilities) | Google DeepMind — Gemini | Gemini 1.0 | Holistic harness of more than 50 benchmarks in six capabilities: factuality (open/closed-book retrieval and QA), long-context (summarization, retrieval, QA), math/science … |
| How Far Can Model-Based Evaluation Go? (reward-model calibration) | Meta | Llama 2 | Robustness/calibration of Llama 2's helpfulness and safety reward models against human preference annotations (3 annotators, 7-point Likert); check for Goodhart divergence using a … |
| Human evaluation of helpfulness and safety (Llama 2-Chat) | Meta | Llama 2 | Human ratings of major Llama 2-Chat model versions on helpfulness and safety vs open-source (Falcon, MPT, Vicuna) and closed-source (ChatGPT, PaLM) models on over 4,000 single- … |
| Human evaluations of general capabilities (7,000-prompt taxonomy) | Meta | Llama 3 | Pairwise human evaluation comparing Llama 3 405B with GPT-4 (0125 API), GPT-4o and Claude 3.5 Sonnet on ~7,000 prompts spanning six individual capabilities (English, reasoning … |
| Human feedback Elo evaluations (helpfulness, honesty, harmlessness) | Anthropic | Claude 2 | per-task Elo scores from binary human preference data on helpfulness, honesty, and a red-teaming harmlessness task |
| Human feedback evaluations (Claude 3.5 Sonnet) | Anthropic | Claude 3.5 Sonnet | human preference win rates vs prior Claude models on common tasks and expert domains |
| Human feedback evaluations (upgraded Claude 3.5 Sonnet / 3.5 Haiku) | Anthropic | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet | human preference win rates vs prior Claude models |
| Human Multi-Turn Evaluations | Google DeepMind — Gemma | Gemma 2 | human raters converse with models following 500 specified multi-turn scenarios (brainstorming, planning, learning something new); average 8.4 user turns; rated on overall … |
| Human preference evaluation (o1-preview vs GPT-4o) | OpenAI | o1 | Human preference on challenging open-ended prompts across domains |
| Human preference evaluation (o3-mini vs o1-mini) | OpenAI | o3-mini | Human preference by external expert testers |
| Human Preference Evaluations | Google DeepMind — Gemma | Gemma 1, Gemma 2 | win rate of Gemma IT models versus Mistral v0.2 7B Instruct on a held-out collection of ~1000 prompts (creative writing, coding, instruction following) and ~400 prompts testing … |
| Human preference evaluations on expert knowledge and core capabilities (Claude 3) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | head-to-head human preference win rates (Claude 3 Sonnet vs Claude 2 and Claude Instant) on common use cases, non-English tasks, and expert-knowledge domains |
| Human Sourced Jailbreaks | OpenAI | GPT-4.5, Operator, o1, o1-preview, o3-mini, o3/o4-mini | Jailbreaks sourced from human red teaming ChatGPT jailbreaks sourced from human red teaming Human red teaming evaluation collected by Scale and determined by Scale to be high harm … |
| HumanEval | Anthropic, Google DeepMind — Gemini, OpenAI | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Gemini 1.0, gpt-4, gpt-4o mini | Ability to synthesize Python functions of varying complexity Coding performance python function synthesis Standard code-completion benchmark mapping function descriptions to … |
| HumanEval (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Standard code-completion benchmark; leakage analysis: continued pre-training on test split boosted scores from 74.4% to 89.0% |
| HumanEval Infilling | Google DeepMind — Gemma | CodeGemma | single-line and multi-line metrics in the HumanEval Infilling benchmarks introduced in Fried et al. (2023); latency measured as total seconds to obtain 128-token continuations per … |
| Humanity's Last Exam | Google DeepMind — Gemini, OpenAI | Gemini 2.5, deep research | Expert-level questions across 100+ subjects (3,000+ MCQ and short-answer questions) Humanity's Last Exam benchmark; results sourced from Scale's leaderboard |
| IFEval | Anthropic, OpenAI | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, GPT-4.1 | Instruction following with verifiable instructions instruction following |
| Image encoder ablation (with or without SigLIP) | Google DeepMind — Gemma | PaliGemma 1 | removing the SigLIP image encoder entirely and passing a linear projection of raw RGB patches to the decoder-only LLM (Fuyu/EVE-style), re-tuning the Stage1 learning rate |
| Image generation refusals | OpenAI | o3/o4-mini | Refusals to invoke image generation tool on policy-violating prompts (system mitigations + model refusals) |
| Image recognition benchmarks (MMMU, VQAv2, AI2 Diagram, ChartQA, TextVQA, DocVQA) | Meta | Llama 3 | Vision module attached to Llama 3 on MMMU (validation set, 900 images), VQAv2, AI2 Diagram, ChartQA, TextVQA, DocVQA |
| Image resolution ablation (224px vs 448px checkpoints) | Google DeepMind — Gemma | PaliGemma 1 | PaliGemma's resolution approach: Stage1 pretrained at 224px with a short Stage2 upcycling to 448px and 896px checkpoints, justified against a 'windowing' approach; ablation … |
| Image-to-text caption quality representational harms (CIDEr across groups) | Google DeepMind — Gemini | Gemini 1.0 | Whether images of people are described with similar quality (CIDEr) for different gender appearances and skin tones, following Zhao et al. (2021) |
| Image-to-text content policy violation evaluation | Google DeepMind — Gemini | Gemini 1.0 | Violative captioning behavior on adversarial images and questions targeting safety policies |
| Image-to-text content policy violation evaluation (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Development evaluation of image-to-text policy violations on nuanced prompts (e.g., innocuous images with leading text prompts), judged by human raters |
| Image-to-text refusal evaluation (grounded/ungrounded) | Google DeepMind — Gemini | Gemini 1.5 | How effectively models refuse sensitive questions about people in images (e.g., religion from a headshot), comparing grounded vs ungrounded queries using MIAP dataset and curated … |
| IMO-Bench (internal expert-graded math) | Google DeepMind — Gemini | Gemini 1.5 | Internally developed expert-graded evaluation testing math at IMO level |
| Impact of distillation w.r.t. model size ablation | Google DeepMind — Gemma | Gemma 2 | measurement of distillation gains as model size increases, maintaining a 7B teacher and training smaller students to simulate the final teacher-student gap |
| Impact of formatting ablation | Google DeepMind — Gemma | Gemma 2 | performance variance on MMLU across 12 prompt/evaluation formatting variations as a proxy for undesired performance variability |
| Impact of image resolution ablation | Google DeepMind — Gemma | Gemma 3 | effect of vision encoder input resolution (higher-resolution encoders use average pooling to reduce output to 256 tokens, e.g., 4x4 average pooling for the 896-resolution encoder) … |
| Impact on KV cache memory ablation | Google DeepMind — Gemma | Gemma 3 | balance between model memory and KV cache during inference with a 32k-token pre-fill context, across local:global ratios and sliding window sizes; KV cache memory as a function of … |
| In-context Scheming Reasoning (Apollo, o3/o4-mini) | OpenAI | o3/o4-mini | In-context scheming reasoning (Covert Subversion, Deferred Subversion, sandbagging) |
| InfiniteBench | Meta | Llama 3 | InfiniteBench (Zhang et al. 2024) En.QA (QA over novels) and En.MC (multiple-choice QA over novels) |
| InfographicVQA (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | InfographicVQA benchmark; ANLS metric |