Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| LSAT | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | standardized law school admissions test performance |
| MakeMePay (Contextual) | OpenAI | Deep Research, GPT-4.5, o1, o1-preview, o3-mini | Manipulative capabilities: one LLM persuading another to make a payment |
| MakeMeSay (Contextual) | OpenAI | Deep Research, GPT-4.5, o1, o1-preview, o3-mini | Deception capabilities: getting the other party to say a codeword Deception capabilities: model's ability to get the other party (AI simulating human) to say a codeword |
| Malicious applications of computer use evaluation (Claude 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | compliance with harmful computer-use requests (surveillance, malicious content distribution, fraud) using human-created scenarios and modified real harm cases |
| Malicious use of agentic coding evaluation (3 evaluations) | Anthropic | Claude Opus 4, Claude Sonnet 4 | safety score across three agentic coding evaluations: 150 clearly prohibited problems, and two borderline/non-harmful sets (50 problems each) to calibrate refusal vs over-refusal |
| Many-shot in-context low-resource machine translation | Google DeepMind — Gemini | Gemini 1.5 | Scaling in-context learning to thousands of examples for English-to-X translation of 6 low-resource languages: public setup (Bemba, Ewe, Kurdish; Flores-200 dev set) and in-house … |
| MATH | Anthropic, Meta | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Llama 2 | math problem solving MATH benchmark performance of Llama 1 and Llama 2 base models (4-shot per model card) |
| MATH (competition math) | Google DeepMind — Gemini | Gemini 1.0 | Math problems drawn from middle- and high-school math competitions |
| MATH (Hendrycks; Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Hendrycks MATH benchmark; 1.5 Pro +14.5% over 1.0 Ultra; Flash +22.3% over 1.0 Pro |
| Math and reasoning benchmarks (GSM8K, MATH, GPQA, ARC-C) | Meta | Llama 3 | GSM8K, MATH, GPQA and ARC-C performance of Llama 3 post-trained models |
| Mathematical reasoning (MATH, GSM8k) | Meta | Llama 1 | Performance on MATH (12K middle/high school math problems) and GSM8k with and without maj1@k |
| MathOdyssey (math-specialized Gemini 1.5 Pro) | Google DeepMind — Gemini | Gemini 1.5 | MathOdyssey benchmark (Fang et al., 2024); math-specialized model |
| MathVista | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | mathematical reasoning in visual contexts |
| MathVista (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | MathVista benchmark requiring mathematical reasoning in visual context |
| MathVista (visual math reasoning) | Google DeepMind — Gemini | Gemini 1.0 | Comprehensive mathematical reasoning benchmark of 28 published multimodal datasets plus three new datasets; authors' evaluation script used |
| MBE (Multistate Bar Examination) | Anthropic | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | bar examination performance |
| MBPP | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | program synthesis (mostly basic python programs) |
| Measure of False Refusal | Meta | Llama 2 | Frequency with which Llama 2-Chat incorrectly refuses non-adversarial prompts, measured on a helpfulness dataset and a borderline set as safety data proportion varies |
| Memorization rate evaluation (exact and approximate) | Google DeepMind — Gemma | Gemma 3 | memorization rate defined as the ratio of generations that match training data to all generations; methodology of Gemma Team (2024b): subsample a large portion of training data … |
| Memorization rate evaluation (prefix continuation) | Google DeepMind — Gemini | Gemini 1.5 | Memorization measured with Gemma-Team (2024) methodology: sample 10,000 documents per pre-training corpus, prompt with first 50 tokens, classify exact match vs edit distance <10% |
| METR autonomous capabilities assessment (GPT-4.5) | OpenAI | GPT-4.5 | Time horizon score: duration of tasks an LLM agent can complete with 50% reliability |
| METR autonomous capabilities assessment (GPT-4o) | OpenAI | GPT-4o | Autonomous capabilities: agent performance vs humans with time limits |
| METR autonomous capabilities assessment (o1) | OpenAI | o1 | Autonomous capabilities on METR's suite of diverse agentic tasks (earlier checkpoint) |
| METR autonomous capabilities assessment (o1-preview) | OpenAI | o1-preview | Autonomous capabilities on METR's suite of diverse agentic tasks |
| METR autonomous capabilities assessment (o3/o4-mini) | OpenAI | o3/o4-mini | Autonomous capabilities (15-day assessment of earlier o4-mini and o3 checkpoints) |
| METR data deduplication | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | 2-8 hour software engineering capability via a METR public task |
| MGSM | OpenAI | gpt-4o mini | Mathematical reasoning (MGSM) |
| MGSM (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | MGSM: 11 languages; 1.5 Pro improves almost +9% over 1.0 Ultra; Flash surpasses by >3 |
| MGSM (multilingual math) | Anthropic, Google DeepMind — Gemini, Meta | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Gemini 1.0, Llama 3 | multilingual math reasoning (translated GSM8K) Professionally translated variant of GSM8K into 11 languages MGSM (Shi et al. 2022) with native prompts from simple-evals (OpenAI … |
| Mini-Grid (PDDL planning) | Google DeepMind — Gemini | Gemini 1.5 | AIPS-1998 mini-grid problem: robot navigation to goal cell in random floorplans |
| ML Engineering (METR) | OpenAI | GPT-4o | Select machine learning engineering tasks from METR |
| MLE-Bench | OpenAI | Deep Research, GPT-4.5, o1, o3-mini | Agent's ability to solve Kaggle challenges Agent's ability to solve Kaggle challenges involving design, building, and training ML models on GPUs Agent's ability to solve Kaggle … |
| MMLU | Anthropic, Meta, OpenAI | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Llama 1, Llama 2, gpt-4, gpt-4o mini | Massive Multitask Language Understanding: 57 subjects of multiple-choice problems Textual intelligence and reasoning (MMLU) multidisciplinary question answering Massive multitask … |
| MMLU (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | MMLU benchmark; 57 subjects; 1.0 Ultra, 1.5 Pro, 1.5 Flash all exceed 80% |
| MMLU (LLaMA-I instruction-tuned) | Meta | Llama 1 | MMLU performance of the LLaMA-I instruct model trained following the Chung et al. (2022) protocol |
| MMLU (Massive Multitask Language Understanding) | Google DeepMind — Gemini | Gemini 1.0 | Holistic exam benchmark measuring knowledge across 57 subjects plus reading comprehension and reasoning; human expert performance 89.8% |
| MMMLU | Anthropic | Claude Opus 4, Claude Sonnet 4 | multilingual MMLU (average over 14 non-English languages) |
| MMMU | Anthropic, OpenAI | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude Opus 4, Claude Sonnet 4, GPT-5, o3, o4-mini | Multimodal understanding massive multi-discipline multimodal understanding and reasoning |
| MMMU (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | MMMU benchmark: images + college-level problems; CoT rationale allowed |
| MMMU (Massive Multi-discipline Multimodal Understanding) | Google DeepMind — Gemini | Gemini 1.0 | Questions about images across 6 disciplines with multiple subjects requiring college-level knowledge |
| MMVP (paired accuracy) | Google DeepMind — Gemma | PaliGemma 1 | MMVP benchmark paired accuracy at 224px for PaliGemma vs GPT4-V, Gemini and other models (LLaVa and others below chance) |
| Model robustness to benchmark design choices (MMLU perturbations) | Meta | Llama 3 | Robustness of pretrained Llama 3 on MMLU to (1) few-shot label bias, (2) label variants, (3) answer order, (4) prompt format |
| Model size and transfer learning rate | Google DeepMind — Gemma | PaliGemma 2 | normalized per-task transfer performance as a function of the transfer learning rate for each model size |
| Model-biotool integration | OpenAI | o1, o3-mini | Use of biological tooling to advance automated agent synthesis |
| Model-biotool integration (o1-preview) | OpenAI | o1-preview | Use of biological tooling to advance automated agent synthesis |
| Model-card benchmark evaluation suite (performance matrix) | Google DeepMind — Gemini | Gemini 2.0 Flash Model Card, Gemini 2.0 Flash-Lite Model Card, Gemini 2.5 Flash Model Card, Gemini 2.5 Flash-Lite Model Card, Gemini 2.5 Pro Model Card, Gemini Ultra Model Card, Model card for Gemini 1.5 Pro and Flash | Benchmark performance matrix reported in the model card (e.g., reasoning, coding, math, multimodal benchmarks); individual benchmark numbers are in a table not extracted into the … |
| Molecular structure recognition transfer evaluation | Google DeepMind — Gemma | PaliGemma 2 | inferring molecule graph structure (SMILES string) from molecular drawings; trained on 1 million PubChem molecules rendered with Indigo with drawing-style augmentation; evaluated … |
| Money Talks | Google DeepMind — Gemma | Gemma 2 | whether the model can convince study participants to donate money to charity: participants told they will receive a GBP 20 bonus and may forfeit part of it for donation; the model … |
| Monitoring for concerning thought processes (9,833 prompts) | Anthropic | Claude 3.7 Sonnet | rates of concerning thinking categories (deception/manipulation, planning harmful actions, distress language) in extended thinking outputs |
| MRCR (Gemini 2.5) | Google DeepMind — Gemini | Gemini 2.5 | MRCR long-context task at 128k context |