Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
GSM-8K OpenAI gpt-4 Mathematical reasoning on grade-school math word problems (test split of 500 problems)
GSM8k Anthropic, Meta Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude Instant 1.2, Llama 2 grade school math problem solving math problem solving (English) GSM8K performance of Llama 2 base models (8-shot per model card) vs closed-source models
GSM8K (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Grade-school math benchmark; 1.5 Pro outperforms 1.0 Ultra and 1.0 Pro
GSM8K (grade-school math) Google DeepMind — Gemini Gemini 1.0 Grade-school math benchmark for elementary exam performance
Hard externally proposed real-world GenAI use cases (preference evaluation) Google DeepMind — Gemini Gemini 1.5 Novel methodology using hundreds of real-world challenging, long-tail, complex prompts shared on the Internet and social media; relative human rater preference (win/tie/loss) vs …
Harmfulness scores on held-out prompts (328 prompts) Anthropic Claude 2 fraction of responses judged more harmful than a fixed refusal reference on a held-out set of 328 prompts
Harms and faithfulness in extended thinking mode Anthropic Claude 3.7 Sonnet harm potential across three parts of a completion (thinking block, final output, CoT outside official blocks); tested with jailbreak techniques (chain-of-thought elicitation …
HealthBench OpenAI GPT-5 Health-related questions (realistic scenarios, physician-defined criteria)
HealthBench Hard OpenAI GPT-5 Health (HealthBench Hard)
HellaSwag Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku commonsense reasoning
HellaSwag (commonsense reasoning) Google DeepMind — Gemini Gemini 1.5 Benchmark testing commonsense reasoning and distinguishing plausible scenarios
HHH (Helpful, Honest, and Harmless) binary-choice evaluation Anthropic Claude 2 model's ability to select the more HHH output from two options on 438 binary-choice questions
Hidden Agenda Google DeepMind — Gemma Gemma 2 participants told they are interacting with a generic chatbot designed to help them learn; the model is secretly tasked with persuading the participant to (1) click a suspicious …
HiddenMath (Gemini 2.5) Google DeepMind — Gemini Gemini 2.5 Internally developed non-public benchmark (HiddenMath) reported to move beyond decontamination reliance
HiddenMath (internal competition math) Google DeepMind — Gemini Gemini 1.5 Internal benchmark of 179 competition-level math problems crafted by experts, evaluated automatically
High School CTFs OpenAI Deep Research, GPT-4.5, o1, o1-preview, o3-mini Can models solve High School cybersecurity challenges? Can models solve competitive high school level cybersecurity challenges?
Holistic capability harness (50+ benchmarks, six capabilities) Google DeepMind — Gemini Gemini 1.0 Holistic harness of more than 50 benchmarks in six capabilities: factuality (open/closed-book retrieval and QA), long-context (summarization, retrieval, QA), math/science …
How Far Can Model-Based Evaluation Go? (reward-model calibration) Meta Llama 2 Robustness/calibration of Llama 2's helpfulness and safety reward models against human preference annotations (3 annotators, 7-point Likert); check for Goodhart divergence using a …
Human evaluation of helpfulness and safety (Llama 2-Chat) Meta Llama 2 Human ratings of major Llama 2-Chat model versions on helpfulness and safety vs open-source (Falcon, MPT, Vicuna) and closed-source (ChatGPT, PaLM) models on over 4,000 single- …
Human evaluations of general capabilities (7,000-prompt taxonomy) Meta Llama 3 Pairwise human evaluation comparing Llama 3 405B with GPT-4 (0125 API), GPT-4o and Claude 3.5 Sonnet on ~7,000 prompts spanning six individual capabilities (English, reasoning …
Human feedback Elo evaluations (helpfulness, honesty, harmlessness) Anthropic Claude 2 per-task Elo scores from binary human preference data on helpfulness, honesty, and a red-teaming harmlessness task
Human feedback evaluations (Claude 3.5 Sonnet) Anthropic Claude 3.5 Sonnet human preference win rates vs prior Claude models on common tasks and expert domains
Human feedback evaluations (upgraded Claude 3.5 Sonnet / 3.5 Haiku) Anthropic Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet human preference win rates vs prior Claude models
Human Multi-Turn Evaluations Google DeepMind — Gemma Gemma 2 human raters converse with models following 500 specified multi-turn scenarios (brainstorming, planning, learning something new); average 8.4 user turns; rated on overall …
Human preference evaluation (o1-preview vs GPT-4o) OpenAI o1 Human preference on challenging open-ended prompts across domains
Human preference evaluation (o3-mini vs o1-mini) OpenAI o3-mini Human preference by external expert testers
Human Preference Evaluations Google DeepMind — Gemma Gemma 1, Gemma 2 win rate of Gemma IT models versus Mistral v0.2 7B Instruct on a held-out collection of ~1000 prompts (creative writing, coding, instruction following) and ~400 prompts testing …
Human preference evaluations on expert knowledge and core capabilities (Claude 3) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku head-to-head human preference win rates (Claude 3 Sonnet vs Claude 2 and Claude Instant) on common use cases, non-English tasks, and expert-knowledge domains
Human Sourced Jailbreaks OpenAI GPT-4.5, Operator, o1, o1-preview, o3-mini, o3/o4-mini Jailbreaks sourced from human red teaming ChatGPT jailbreaks sourced from human red teaming Human red teaming evaluation collected by Scale and determined by Scale to be high harm …
HumanEval Anthropic, Google DeepMind — Gemini, OpenAI Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Gemini 1.0, gpt-4, gpt-4o mini Ability to synthesize Python functions of varying complexity Coding performance python function synthesis Standard code-completion benchmark mapping function descriptions to …
HumanEval (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Standard code-completion benchmark; leakage analysis: continued pre-training on test split boosted scores from 74.4% to 89.0%
HumanEval Infilling Google DeepMind — Gemma CodeGemma single-line and multi-line metrics in the HumanEval Infilling benchmarks introduced in Fried et al. (2023); latency measured as total seconds to obtain 128-token continuations per …
Humanity's Last Exam Google DeepMind — Gemini, OpenAI Gemini 2.5, deep research Expert-level questions across 100+ subjects (3,000+ MCQ and short-answer questions) Humanity's Last Exam benchmark; results sourced from Scale's leaderboard
IFEval Anthropic, OpenAI Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, GPT-4.1 Instruction following with verifiable instructions instruction following
Image encoder ablation (with or without SigLIP) Google DeepMind — Gemma PaliGemma 1 removing the SigLIP image encoder entirely and passing a linear projection of raw RGB patches to the decoder-only LLM (Fuyu/EVE-style), re-tuning the Stage1 learning rate
Image generation refusals OpenAI o3/o4-mini Refusals to invoke image generation tool on policy-violating prompts (system mitigations + model refusals)
Image recognition benchmarks (MMMU, VQAv2, AI2 Diagram, ChartQA, TextVQA, DocVQA) Meta Llama 3 Vision module attached to Llama 3 on MMMU (validation set, 900 images), VQAv2, AI2 Diagram, ChartQA, TextVQA, DocVQA
Image resolution ablation (224px vs 448px checkpoints) Google DeepMind — Gemma PaliGemma 1 PaliGemma's resolution approach: Stage1 pretrained at 224px with a short Stage2 upcycling to 448px and 896px checkpoints, justified against a 'windowing' approach; ablation …
Image-to-text caption quality representational harms (CIDEr across groups) Google DeepMind — Gemini Gemini 1.0 Whether images of people are described with similar quality (CIDEr) for different gender appearances and skin tones, following Zhao et al. (2021)
Image-to-text content policy violation evaluation Google DeepMind — Gemini Gemini 1.0 Violative captioning behavior on adversarial images and questions targeting safety policies
Image-to-text content policy violation evaluation (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Development evaluation of image-to-text policy violations on nuanced prompts (e.g., innocuous images with leading text prompts), judged by human raters
Image-to-text refusal evaluation (grounded/ungrounded) Google DeepMind — Gemini Gemini 1.5 How effectively models refuse sensitive questions about people in images (e.g., religion from a headshot), comparing grounded vs ungrounded queries using MIAP dataset and curated …
IMO-Bench (internal expert-graded math) Google DeepMind — Gemini Gemini 1.5 Internally developed expert-graded evaluation testing math at IMO level
Impact of distillation w.r.t. model size ablation Google DeepMind — Gemma Gemma 2 measurement of distillation gains as model size increases, maintaining a 7B teacher and training smaller students to simulate the final teacher-student gap
Impact of formatting ablation Google DeepMind — Gemma Gemma 2 performance variance on MMLU across 12 prompt/evaluation formatting variations as a proxy for undesired performance variability
Impact of image resolution ablation Google DeepMind — Gemma Gemma 3 effect of vision encoder input resolution (higher-resolution encoders use average pooling to reduce output to 256 tokens, e.g., 4x4 average pooling for the 896-resolution encoder) …
Impact on KV cache memory ablation Google DeepMind — Gemma Gemma 3 balance between model memory and KV cache during inference with a 32k-token pre-fill context, across local:global ratios and sliding window sizes; KV cache memory as a function of …
In-context Scheming Reasoning (Apollo, o3/o4-mini) OpenAI o3/o4-mini In-context scheming reasoning (Covert Subversion, Deferred Subversion, sandbagging)
InfiniteBench Meta Llama 3 InfiniteBench (Zhang et al. 2024) En.QA (QA over novels) and En.MC (multiple-choice QA over novels)
InfographicVQA (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 InfographicVQA benchmark; ANLS metric