Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
LSAT Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku standardized law school admissions test performance
MakeMePay (Contextual) OpenAI Deep Research, GPT-4.5, o1, o1-preview, o3-mini Manipulative capabilities: one LLM persuading another to make a payment
MakeMeSay (Contextual) OpenAI Deep Research, GPT-4.5, o1, o1-preview, o3-mini Deception capabilities: getting the other party to say a codeword Deception capabilities: model's ability to get the other party (AI simulating human) to say a codeword
Malicious applications of computer use evaluation (Claude 4) Anthropic Claude Opus 4, Claude Sonnet 4 compliance with harmful computer-use requests (surveillance, malicious content distribution, fraud) using human-created scenarios and modified real harm cases
Malicious use of agentic coding evaluation (3 evaluations) Anthropic Claude Opus 4, Claude Sonnet 4 safety score across three agentic coding evaluations: 150 clearly prohibited problems, and two borderline/non-harmful sets (50 problems each) to calibrate refusal vs over-refusal
Many-shot in-context low-resource machine translation Google DeepMind — Gemini Gemini 1.5 Scaling in-context learning to thousands of examples for English-to-X translation of 6 low-resource languages: public setup (Bemba, Ewe, Kurdish; Flores-200 dev set) and in-house …
MATH Anthropic, Meta Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Llama 2 math problem solving MATH benchmark performance of Llama 1 and Llama 2 base models (4-shot per model card)
MATH (competition math) Google DeepMind — Gemini Gemini 1.0 Math problems drawn from middle- and high-school math competitions
MATH (Hendrycks; Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Hendrycks MATH benchmark; 1.5 Pro +14.5% over 1.0 Ultra; Flash +22.3% over 1.0 Pro
Math and reasoning benchmarks (GSM8K, MATH, GPQA, ARC-C) Meta Llama 3 GSM8K, MATH, GPQA and ARC-C performance of Llama 3 post-trained models
Mathematical reasoning (MATH, GSM8k) Meta Llama 1 Performance on MATH (12K middle/high school math problems) and GSM8k with and without maj1@k
MathOdyssey (math-specialized Gemini 1.5 Pro) Google DeepMind — Gemini Gemini 1.5 MathOdyssey benchmark (Fang et al., 2024); math-specialized model
MathVista Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku mathematical reasoning in visual contexts
MathVista (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 MathVista benchmark requiring mathematical reasoning in visual context
MathVista (visual math reasoning) Google DeepMind — Gemini Gemini 1.0 Comprehensive mathematical reasoning benchmark of 28 published multimodal datasets plus three new datasets; authors' evaluation script used
MBE (Multistate Bar Examination) Anthropic Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku bar examination performance
MBPP Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku program synthesis (mostly basic python programs)
Measure of False Refusal Meta Llama 2 Frequency with which Llama 2-Chat incorrectly refuses non-adversarial prompts, measured on a helpfulness dataset and a borderline set as safety data proportion varies
Memorization rate evaluation (exact and approximate) Google DeepMind — Gemma Gemma 3 memorization rate defined as the ratio of generations that match training data to all generations; methodology of Gemma Team (2024b): subsample a large portion of training data …
Memorization rate evaluation (prefix continuation) Google DeepMind — Gemini Gemini 1.5 Memorization measured with Gemma-Team (2024) methodology: sample 10,000 documents per pre-training corpus, prompt with first 50 tokens, classify exact match vs edit distance <10%
METR autonomous capabilities assessment (GPT-4.5) OpenAI GPT-4.5 Time horizon score: duration of tasks an LLM agent can complete with 50% reliability
METR autonomous capabilities assessment (GPT-4o) OpenAI GPT-4o Autonomous capabilities: agent performance vs humans with time limits
METR autonomous capabilities assessment (o1) OpenAI o1 Autonomous capabilities on METR's suite of diverse agentic tasks (earlier checkpoint)
METR autonomous capabilities assessment (o1-preview) OpenAI o1-preview Autonomous capabilities on METR's suite of diverse agentic tasks
METR autonomous capabilities assessment (o3/o4-mini) OpenAI o3/o4-mini Autonomous capabilities (15-day assessment of earlier o4-mini and o3 checkpoints)
METR data deduplication Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 2-8 hour software engineering capability via a METR public task
MGSM OpenAI gpt-4o mini Mathematical reasoning (MGSM)
MGSM (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 MGSM: 11 languages; 1.5 Pro improves almost +9% over 1.0 Ultra; Flash surpasses by >3
MGSM (multilingual math) Anthropic, Google DeepMind — Gemini, Meta Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Gemini 1.0, Llama 3 multilingual math reasoning (translated GSM8K) Professionally translated variant of GSM8K into 11 languages MGSM (Shi et al. 2022) with native prompts from simple-evals (OpenAI …
Mini-Grid (PDDL planning) Google DeepMind — Gemini Gemini 1.5 AIPS-1998 mini-grid problem: robot navigation to goal cell in random floorplans
ML Engineering (METR) OpenAI GPT-4o Select machine learning engineering tasks from METR
MLE-Bench OpenAI Deep Research, GPT-4.5, o1, o3-mini Agent's ability to solve Kaggle challenges Agent's ability to solve Kaggle challenges involving design, building, and training ML models on GPUs Agent's ability to solve Kaggle …
MMLU Anthropic, Meta, OpenAI Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Llama 1, Llama 2, gpt-4, gpt-4o mini Massive Multitask Language Understanding: 57 subjects of multiple-choice problems Textual intelligence and reasoning (MMLU) multidisciplinary question answering Massive multitask …
MMLU (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 MMLU benchmark; 57 subjects; 1.0 Ultra, 1.5 Pro, 1.5 Flash all exceed 80%
MMLU (LLaMA-I instruction-tuned) Meta Llama 1 MMLU performance of the LLaMA-I instruct model trained following the Chung et al. (2022) protocol
MMLU (Massive Multitask Language Understanding) Google DeepMind — Gemini Gemini 1.0 Holistic exam benchmark measuring knowledge across 57 subjects plus reading comprehension and reasoning; human expert performance 89.8%
MMMLU Anthropic Claude Opus 4, Claude Sonnet 4 multilingual MMLU (average over 14 non-English languages)
MMMU Anthropic, OpenAI Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude Opus 4, Claude Sonnet 4, GPT-5, o3, o4-mini Multimodal understanding massive multi-discipline multimodal understanding and reasoning
MMMU (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 MMMU benchmark: images + college-level problems; CoT rationale allowed
MMMU (Massive Multi-discipline Multimodal Understanding) Google DeepMind — Gemini Gemini 1.0 Questions about images across 6 disciplines with multiple subjects requiring college-level knowledge
MMVP (paired accuracy) Google DeepMind — Gemma PaliGemma 1 MMVP benchmark paired accuracy at 224px for PaliGemma vs GPT4-V, Gemini and other models (LLaVa and others below chance)
Model robustness to benchmark design choices (MMLU perturbations) Meta Llama 3 Robustness of pretrained Llama 3 on MMLU to (1) few-shot label bias, (2) label variants, (3) answer order, (4) prompt format
Model size and transfer learning rate Google DeepMind — Gemma PaliGemma 2 normalized per-task transfer performance as a function of the transfer learning rate for each model size
Model-biotool integration OpenAI o1, o3-mini Use of biological tooling to advance automated agent synthesis
Model-biotool integration (o1-preview) OpenAI o1-preview Use of biological tooling to advance automated agent synthesis
Model-card benchmark evaluation suite (performance matrix) Google DeepMind — Gemini Gemini 2.0 Flash Model Card, Gemini 2.0 Flash-Lite Model Card, Gemini 2.5 Flash Model Card, Gemini 2.5 Flash-Lite Model Card, Gemini 2.5 Pro Model Card, Gemini Ultra Model Card, Model card for Gemini 1.5 Pro and Flash Benchmark performance matrix reported in the model card (e.g., reasoning, coding, math, multimodal benchmarks); individual benchmark numbers are in a table not extracted into the …
Molecular structure recognition transfer evaluation Google DeepMind — Gemma PaliGemma 2 inferring molecule graph structure (SMILES string) from molecular drawings; trained on 1 million PubChem molecules rendered with Indigo with drawing-style augmentation; evaluated …
Money Talks Google DeepMind — Gemma Gemma 2 whether the model can convince study participants to donate money to charity: participants told they will receive a GBP 20 bonus and may forfeit part of it for donation; the model …
Monitoring for concerning thought processes (9,833 prompts) Anthropic Claude 3.7 Sonnet rates of concerning thinking categories (deception/manipulation, planning harmful actions, distress language) in extended thinking outputs
MRCR (Gemini 2.5) Google DeepMind — Gemini Gemini 2.5 MRCR long-context task at 128k context