Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| MRCR (Multiround Co-reference Resolution) | Google DeepMind — Gemini | Gemini 1.5 | Long-conversation task: user requests writing on topics with two randomly placed distinct requests; measures retrieval and disambiguation to 1M tokens |
| MTOB (Machine Translation from One Book) | Google DeepMind — Gemini | Gemini 1.5 | MTOB benchmark (Tanzer et al., 2023): learn to translate English-Kalamang (ISO kgv, <200 speakers) from a ~500-page grammar and ~2000-entry wordlist in context; human evaluation … |
| Multi-lingual Benchmarks | Google DeepMind — Gemma | CodeGemma | code generation across a variety of popular programming languages measured with BabelCode (Orlanski et al. 2023) on BabelCode-translated HumanEval and MBPP datasets; languages … |
| Multi-programming language code generation (MultiPL-E) | Meta | Llama 3 | MultiPL-E benchmark (translations of HumanEval and MBPP problems) across a subset of popular programming languages |
| Multi-turn testing (safeguards) | Anthropic | Claude Opus 4, Claude Sonnet 4 | policy violations across thousands of multi-turn conversations (automated generation and manual expert conversations) filtered with policy-specific grading rubrics |
| MultiChallenge (Scale) | OpenAI | GPT-4.1 | Multi-turn instruction following (4 types of information from previous messages) |
| Multilingual benchmark evaluation | Google DeepMind — Gemma | Gemma 3 | performance of pre-trained models on multilingual tasks with in-context learning (multi-shot prompting) on MGSM, Global-MMLU-Lite, WMT24++, FLoRes, XQuAD, ECLeKTic, IndicGenBench … |
| Multilingual benchmark evaluations | Meta | Llama 3.1, Llama 3.2 | Multilingual benchmark results reported in the Llama 3.1 model card (table caption; benchmark names not present in the extracted chunk text) Multilingual benchmark results … |
| Multilingual Librispeech (ASR) | Google DeepMind — Gemini | Gemini 1.0 | ASR benchmark (Multilingual Librispeech); WER metric |
| Multilingual Librispeech (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Public ASR benchmark; WER metric |
| Multilingual MMLU | Anthropic, Meta, OpenAI | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, GPT-4, Llama 3 | GPT-4 capability in other languages relative to English-language MMLU performance multilingual common-sense reasoning (MMLU translated) MMLU questions, few-shot examples and … |
| Multilingual Performance (Simple Evals test set) | OpenAI | GPT-4.5, o1, o1-preview, o3-mini, o3/o4-mini | multilingual capability relative to GPT-4o multilingual capability relative to o1-mini multilingual capability relative to o1 and o3-mini |
| Multilingual post-training evaluation (SxS, Gemini Apps) | Google DeepMind — Gemini | Gemini 1.0 | Side-by-side (SxS) quality comparison of Gemini Apps (with Pro) vs Bard (PaLM 2-based) across 5 languages; SxS score centered at 0, range -1.5..1.5 |
| Multilingual safety evaluation | Meta | Llama 3 | Safety knowledge transfer across languages on an internal per-language benchmark: Llama 405B with and without Llama Guard vs two competing systems, plus violation/false-refusal … |
| Multimodal jailbreak evaluations | OpenAI | GPT-4V | Refusal evaluation for text-screenshot jailbreaks where logical reasoning needed to break the model is placed in images |
| Multimodal policy red-teaming (Trust & Safety) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | pass/fail on harmless responses (AUP/TOS/Constitutional AI alignment) and pass/fail on desirable responses (accurate identification and thorough informative response) |
| Multimodal post-training vision evaluation (SxS and benchmark comparisons) | Google DeepMind — Gemini | Gemini 1.0 | SxS evaluation of text-only quality (+0.01 for a Pro model trained with image-text data) and image-understanding tasks (+0.223 for SFT+RLHF vs SFT alone); plus SFT impact of API … |
| Multimodal pretraining duration ablation | Google DeepMind — Gemma | PaliGemma 1 | effect of Stage1 multimodal pretraining duration (down to completely skipping Stage1) on transfer performance; ablations run with Stage1 10x shorter (100M examples seen) unless … |
| Multimodal Refusal Evaluation | OpenAI | GPT-4.5, o1 | Refusals for multimodal inputs on standard set for disallowed text+image content and overrefusals (categories: sexual/exploitative, self-harm/intent, self-harm/instructions) … |
| Multimodal transfer evaluation vs PaliGemma 2 | Google DeepMind — Gemma | Gemma 3 | fine-tuning of multimodal Gemma 3 pre-trained checkpoints following the Steiner et al. (2024) protocol (only learning rate swept, otherwise same transfer settings), compared with … |
| Multimodal troubleshooting virology | OpenAI | Deep Research, GPT-4.5, o1, o3-mini, o3/o4-mini | Virology protocol troubleshooting capability (MCQ) Virology protocol troubleshooting capability (MCQ, multimodal) |
| Multimodal virology (VCT) | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | performance on multiple-choice virology questions combining text statements with images (multiple-select variant) performance on multiple-select virology questions combining text … |
| Multiple needles-in-a-haystack retrieval (100 needles) | Google DeepMind — Gemini | Gemini 1.5 | Extension of needle-in-a-haystack with 100 unique needles in a single haystack up to 1M tokens; recall of correct needles |
| Natural language capability benchmarks (CodeGemma) | Google DeepMind — Gemma | CodeGemma | performance on question answering (BoolQ, PIQA, TriviaQA), natural language (ARC-Challenge, HellaSwag, MMLU, WinoGrande) and mathematical reasoning (GSM8K, MATH) for the two 7B … |
| Natural2Code | Google DeepMind — Gemini | Gemini 1.0 | Held-out evaluation benchmark for Python code generation tasks with no web leakage |
| Natural2Code (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Held-out code generation test set preventing web leakage, same format as HumanEval |
| Needle In A Haystack | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku, Claude 3.5 Sonnet | long-context information retrieval long-context retrieval |
| Needle in a haystack internal eval | OpenAI | GPT-4.1 | Long-context retrieval (up to 1M tokens) |
| Needle-in-a-Haystack | Meta | Llama 3 | Needle-in-a-Haystack (Kamradt 2023): retrieve hidden information inserted in random parts of long documents at all depths and context lengths; also Multi-needle variation with … |
| New token initialization ablation | Google DeepMind — Gemma | PaliGemma 1 | initialization of the 1024 location tokens (<loc0000>-<loc1023>) and 128 VQVAE mask tokens (<seg000>-<seg127>) added to Gemma's vocabulary: standard Gaussian noise (sigma=0.02) vs … |
| NextQA (video QA) | Google DeepMind — Gemini | Gemini 1.0 | Zero-shot video QA benchmark; WUPS metric; 16 equally-spaced frames sampled |
| Novel compiler task | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to create a compiler for a novel and unusual programming language given only a specification and test cases |
| Novelty probing (deep research) | OpenAI | Deep Research | Whether model can apply biological knowledge in a novel manner for threat design |
| o1 safety evaluation suite overview | OpenAI | o1 | Propensity to generate disallowed content, demographic fairness, hallucination, dangerous capabilities; external red teaming |
| o1-preview safety evaluation suite overview | OpenAI | o1-preview | Propensity to generate disallowed content, demographic fairness, hallucination, dangerous capabilities; external red teaming |
| Object detection transfer evaluation (COCO, DocLayNet) | Google DeepMind — Gemma | PaliGemma 2 | transfer of PaliGemma to MS COCO object detection and DocLayNet document layout detection using a pix2seq-inspired sequence augmentation approach ('detect all classes' prefix … |
| Offensive cyber-security | Google DeepMind — Gemma | Gemma 2 | automated capture-the-flag (CTF) challenges where the model is tasked with hacking into a simulated server to retrieve secret information; run on the Gemma 2 27B against … |
| Offensive cybersecurity CTF challenge suite | Google DeepMind — Gemini | Gemini 1.0 | Ability of Gemini API Pro/Ultra and Gemini Advanced to solve offensive cybersecurity CTF challenges of varying difficulty |
| Offensive cybersecurity CTF challenge suite (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | CTF challenges where the agent breaks into a simulated server and finds secret information; suites: internal challenge suite (Phuong et al., 2024), in-house 'worm' challenge, Hack … |
| Open Ended Questions (hallucination) | OpenAI | o1-preview | number of incorrect statements generated when asked for arbitrary facts |
| OpenAI PRs | OpenAI | Deep Research, GPT-4.5, o3-mini, o3/o4-mini | Ability to replicate pull request contributions by OpenAI employees |
| OpenAI Research Engineer Interviews (Multiple Choice & Coding) | OpenAI | Deep Research, GPT-4.5, GPT-4o, o1, o1-preview, o3-mini, o3/o4-mini | Ability to pass OpenAI Research Engineer interview loop Ability to automate machine learning research & development (chained actions and reliable coding task execution) |
| OpenAI-MRCR (Multi-Round Coreference) | OpenAI | GPT-4.1 | Multi-needle disambiguation in multi-turn synthetic conversations |
| OpenEQA (embodied question answering) | Google DeepMind — Gemini | Gemini 1.5 | Open-vocabulary benchmark for embodied question answering; LM evaluation scores from the original paper |
| Operator risk identification and policy creation | OpenAI | operator | Risky tasks and actions for a computer-using agent; policy created by risk severity and reversibility |
| Optical music score recognition transfer evaluation | Google DeepMind — Gemma | PaliGemma 2 | translating images of single-line pianoform scores into their digital score representation in the **kern format; GrandStaff dataset (53.7k images, official splits); Character … |
| OSWorld | Anthropic | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet | computer use: real-world computer tasks involving web and desktop applications, OS file I/O, multi-application workflows |
| Pairwise Safety Comparison (o1) | OpenAI | o1 | Red teamer safety ratings of anonymized o1 vs GPT-4o responses |
| Pairwise Safety Comparison (o3-mini) | OpenAI | o3-mini | Red teamer safety ratings of anonymized gpt-4o, o1, o3-mini responses |
| Pan & Scan ablation | Google DeepMind — Gemma | Gemma 3 | Pan & Scan (P&S), which enables capturing images at close to their native aspect ratio and resolution; 4-shot evaluation with and without P&S on a pre-trained checkpoint (27B IT … |