Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
MRCR (Multiround Co-reference Resolution) Google DeepMind — Gemini Gemini 1.5 Long-conversation task: user requests writing on topics with two randomly placed distinct requests; measures retrieval and disambiguation to 1M tokens
MTOB (Machine Translation from One Book) Google DeepMind — Gemini Gemini 1.5 MTOB benchmark (Tanzer et al., 2023): learn to translate English-Kalamang (ISO kgv, <200 speakers) from a ~500-page grammar and ~2000-entry wordlist in context; human evaluation …
Multi-lingual Benchmarks Google DeepMind — Gemma CodeGemma code generation across a variety of popular programming languages measured with BabelCode (Orlanski et al. 2023) on BabelCode-translated HumanEval and MBPP datasets; languages …
Multi-programming language code generation (MultiPL-E) Meta Llama 3 MultiPL-E benchmark (translations of HumanEval and MBPP problems) across a subset of popular programming languages
Multi-turn testing (safeguards) Anthropic Claude Opus 4, Claude Sonnet 4 policy violations across thousands of multi-turn conversations (automated generation and manual expert conversations) filtered with policy-specific grading rubrics
MultiChallenge (Scale) OpenAI GPT-4.1 Multi-turn instruction following (4 types of information from previous messages)
Multilingual benchmark evaluation Google DeepMind — Gemma Gemma 3 performance of pre-trained models on multilingual tasks with in-context learning (multi-shot prompting) on MGSM, Global-MMLU-Lite, WMT24++, FLoRes, XQuAD, ECLeKTic, IndicGenBench …
Multilingual benchmark evaluations Meta Llama 3.1, Llama 3.2 Multilingual benchmark results reported in the Llama 3.1 model card (table caption; benchmark names not present in the extracted chunk text) Multilingual benchmark results …
Multilingual Librispeech (ASR) Google DeepMind — Gemini Gemini 1.0 ASR benchmark (Multilingual Librispeech); WER metric
Multilingual Librispeech (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Public ASR benchmark; WER metric
Multilingual MMLU Anthropic, Meta, OpenAI Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, GPT-4, Llama 3 GPT-4 capability in other languages relative to English-language MMLU performance multilingual common-sense reasoning (MMLU translated) MMLU questions, few-shot examples and …
Multilingual Performance (Simple Evals test set) OpenAI GPT-4.5, o1, o1-preview, o3-mini, o3/o4-mini multilingual capability relative to GPT-4o multilingual capability relative to o1-mini multilingual capability relative to o1 and o3-mini
Multilingual post-training evaluation (SxS, Gemini Apps) Google DeepMind — Gemini Gemini 1.0 Side-by-side (SxS) quality comparison of Gemini Apps (with Pro) vs Bard (PaLM 2-based) across 5 languages; SxS score centered at 0, range -1.5..1.5
Multilingual safety evaluation Meta Llama 3 Safety knowledge transfer across languages on an internal per-language benchmark: Llama 405B with and without Llama Guard vs two competing systems, plus violation/false-refusal …
Multimodal jailbreak evaluations OpenAI GPT-4V Refusal evaluation for text-screenshot jailbreaks where logical reasoning needed to break the model is placed in images
Multimodal policy red-teaming (Trust & Safety) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku pass/fail on harmless responses (AUP/TOS/Constitutional AI alignment) and pass/fail on desirable responses (accurate identification and thorough informative response)
Multimodal post-training vision evaluation (SxS and benchmark comparisons) Google DeepMind — Gemini Gemini 1.0 SxS evaluation of text-only quality (+0.01 for a Pro model trained with image-text data) and image-understanding tasks (+0.223 for SFT+RLHF vs SFT alone); plus SFT impact of API …
Multimodal pretraining duration ablation Google DeepMind — Gemma PaliGemma 1 effect of Stage1 multimodal pretraining duration (down to completely skipping Stage1) on transfer performance; ablations run with Stage1 10x shorter (100M examples seen) unless …
Multimodal Refusal Evaluation OpenAI GPT-4.5, o1 Refusals for multimodal inputs on standard set for disallowed text+image content and overrefusals (categories: sexual/exploitative, self-harm/intent, self-harm/instructions) …
Multimodal transfer evaluation vs PaliGemma 2 Google DeepMind — Gemma Gemma 3 fine-tuning of multimodal Gemma 3 pre-trained checkpoints following the Steiner et al. (2024) protocol (only learning rate swept, otherwise same transfer settings), compared with …
Multimodal troubleshooting virology OpenAI Deep Research, GPT-4.5, o1, o3-mini, o3/o4-mini Virology protocol troubleshooting capability (MCQ) Virology protocol troubleshooting capability (MCQ, multimodal)
Multimodal virology (VCT) Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 performance on multiple-choice virology questions combining text statements with images (multiple-select variant) performance on multiple-select virology questions combining text …
Multiple needles-in-a-haystack retrieval (100 needles) Google DeepMind — Gemini Gemini 1.5 Extension of needle-in-a-haystack with 100 unique needles in a single haystack up to 1M tokens; recall of correct needles
Natural language capability benchmarks (CodeGemma) Google DeepMind — Gemma CodeGemma performance on question answering (BoolQ, PIQA, TriviaQA), natural language (ARC-Challenge, HellaSwag, MMLU, WinoGrande) and mathematical reasoning (GSM8K, MATH) for the two 7B …
Natural2Code Google DeepMind — Gemini Gemini 1.0 Held-out evaluation benchmark for Python code generation tasks with no web leakage
Natural2Code (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Held-out code generation test set preventing web leakage, same format as HumanEval
Needle In A Haystack Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude 3.5 Sonnet long-context information retrieval long-context retrieval
Needle in a haystack internal eval OpenAI GPT-4.1 Long-context retrieval (up to 1M tokens)
Needle-in-a-Haystack Meta Llama 3 Needle-in-a-Haystack (Kamradt 2023): retrieve hidden information inserted in random parts of long documents at all depths and context lengths; also Multi-needle variation with …
New token initialization ablation Google DeepMind — Gemma PaliGemma 1 initialization of the 1024 location tokens (<loc0000>-<loc1023>) and 128 VQVAE mask tokens (<seg000>-<seg127>) added to Gemma's vocabulary: standard Gaussian noise (sigma=0.02) vs …
NextQA (video QA) Google DeepMind — Gemini Gemini 1.0 Zero-shot video QA benchmark; WUPS metric; 16 equally-spaced frames sampled
Novel compiler task Anthropic Claude Opus 4, Claude Sonnet 4 ability to create a compiler for a novel and unusual programming language given only a specification and test cases
Novelty probing (deep research) OpenAI Deep Research Whether model can apply biological knowledge in a novel manner for threat design
o1 safety evaluation suite overview OpenAI o1 Propensity to generate disallowed content, demographic fairness, hallucination, dangerous capabilities; external red teaming
o1-preview safety evaluation suite overview OpenAI o1-preview Propensity to generate disallowed content, demographic fairness, hallucination, dangerous capabilities; external red teaming
Object detection transfer evaluation (COCO, DocLayNet) Google DeepMind — Gemma PaliGemma 2 transfer of PaliGemma to MS COCO object detection and DocLayNet document layout detection using a pix2seq-inspired sequence augmentation approach ('detect all classes' prefix …
Offensive cyber-security Google DeepMind — Gemma Gemma 2 automated capture-the-flag (CTF) challenges where the model is tasked with hacking into a simulated server to retrieve secret information; run on the Gemma 2 27B against …
Offensive cybersecurity CTF challenge suite Google DeepMind — Gemini Gemini 1.0 Ability of Gemini API Pro/Ultra and Gemini Advanced to solve offensive cybersecurity CTF challenges of varying difficulty
Offensive cybersecurity CTF challenge suite (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 CTF challenges where the agent breaks into a simulated server and finds secret information; suites: internal challenge suite (Phuong et al., 2024), in-house 'worm' challenge, Hack …
Open Ended Questions (hallucination) OpenAI o1-preview number of incorrect statements generated when asked for arbitrary facts
OpenAI PRs OpenAI Deep Research, GPT-4.5, o3-mini, o3/o4-mini Ability to replicate pull request contributions by OpenAI employees
OpenAI Research Engineer Interviews (Multiple Choice & Coding) OpenAI Deep Research, GPT-4.5, GPT-4o, o1, o1-preview, o3-mini, o3/o4-mini Ability to pass OpenAI Research Engineer interview loop Ability to automate machine learning research & development (chained actions and reliable coding task execution)
OpenAI-MRCR (Multi-Round Coreference) OpenAI GPT-4.1 Multi-needle disambiguation in multi-turn synthetic conversations
OpenEQA (embodied question answering) Google DeepMind — Gemini Gemini 1.5 Open-vocabulary benchmark for embodied question answering; LM evaluation scores from the original paper
Operator risk identification and policy creation OpenAI operator Risky tasks and actions for a computer-using agent; policy created by risk severity and reversibility
Optical music score recognition transfer evaluation Google DeepMind — Gemma PaliGemma 2 translating images of single-line pianoform scores into their digital score representation in the **kern format; GrandStaff dataset (53.7k images, official splits); Character …
OSWorld Anthropic Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet computer use: real-world computer tasks involving web and desktop applications, OS file I/O, multi-application workflows
Pairwise Safety Comparison (o1) OpenAI o1 Red teamer safety ratings of anonymized o1 vs GPT-4o responses
Pairwise Safety Comparison (o3-mini) OpenAI o3-mini Red teamer safety ratings of anonymized gpt-4o, o1, o3-mini responses
Pan & Scan ablation Google DeepMind — Gemma Gemma 3 Pan & Scan (P&S), which enables capturing images at close to their native aspect ratio and resolution; 4-shot evaluation with and without P&S on a pre-trained checkpoint (27B IT …