Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| InfographicVQA (infographic QA) | Google DeepMind — Gemini | Gemini 1.0 | Chart/infographic understanding requiring spatial understanding of input layout |
| Instruction following (IFEval) | Meta | Llama 3 | IFEval: approximately 500 'verifiable instructions' with heuristic verification; average of prompt-level and instruction-level accuracy under strict and loose constraints |
| Instruction following and formatting evaluation (Claude 3) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | ability to follow diverse complex instructions, absolute language, full completion of requests, and structured output generation |
| Instruction Hierarchy Evaluation - Conflicts Between Message Types | OpenAI | GPT-4.5, o1, o3-mini, o3/o4-mini | System vs user message conflicts (GPT-4.5 uses two message classifications) Prompts where different message types conflict; model must follow system over developer over user … |
| Instruction Hierarchy Evaluation - Phrase and Password Protection | OpenAI | GPT-4.5, o1, o3-mini, o3/o4-mini | Model instructed not to output a phrase or reveal a bespoke password Model instructed not to output a certain phrase or reveal a bespoke password |
| Instruction Hierarchy Evaluation - Tutor Jailbreaks | OpenAI | GPT-4.5, o1, o3-mini, o3/o4-mini | User attempts to trick a math-tutor model into ignoring instructions Realistic scenario where a user attempts to trick a math-tutor model into ignoring its instructions |
| Inter-Rater Reliability (IRR) of human evaluations | Meta | Llama 2 | Gwet's AC2 inter-rater reliability across three annotators on the 7-point Likert helpfulness task |
| Interleaved image-text generation evaluation (1-shot) | Google DeepMind — Gemini | Gemini 1.0 | Qualitative 1-shot evaluation of Gemini Ultra generating interleaved sequences of images and text (creative suggestions given two colors), demonstrating native image output … |
| Internal agentic coding evaluation (Claude 3.5 Sonnet) | Anthropic | Claude 3.5 Sonnet | agentic coding: understand an open-source codebase and implement a pull request |
| Internal agentic coding evaluation (upgraded Claude 3.5 Sonnet / 3.5 Haiku) | Anthropic | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet | agentic coding: implement a pull request in an open-source codebase |
| Internal AI research evaluation suite (7 environments) | Anthropic | Claude 3.7 Sonnet | ability to improve performance of ML code across LLMs, time series, low-level optimizations, reinforcement learning, and general problem solving |
| Internal AI research evaluation suite 1 | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to improve performance of machine-learning code across LLMs, time series, low-level optimizations, reinforcement learning, and general problem solving |
| Internal AI research evaluation suite 2 | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to autonomously perform self-contained AI/ML research tasks across subareas relevant to Anthropic research (AI R&D-4 RSP threshold) |
| Internal coding SxS evaluation (Gemini Apps/API) | Google DeepMind — Gemini | Gemini 1.0 | SxS scores on internally curated coding prompts across code use cases and languages: Gemini (Pro) vs Bard, and Gemini Advanced (Ultra) vs Gemini (Pro) |
| Internal model use survey (researcher uplift) | Anthropic | Claude Opus 4, Claude Sonnet 4 | whether researchers believe the model can completely automate a junior ML researcher, and estimated productivity boost (AI R&D capability probe) |
| Internal red-teaming of content policies | Google DeepMind — Gemma | CodeGemma, Gemma 1, Gemma 2, Gemma 3, Gemma 3n, PaliGemma 1 | internal red-teaming testing of relevant content policies, conducted by a number of different teams each with different goals and human evaluation metrics internal red-teaming … |
| Internal safety benchmark construction (adversarial and borderline prompts) | Meta | Llama 3 | Internal safety benchmarks built from human-written adversarial and borderline prompts per risk category (inspired by ML Commons taxonomy of hazards), ranging from direct harmful … |
| Internal YouTube ASR/AST benchmark (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | Internal benchmarks derived from YouTube (English and 52 other languages); WER and BLEU metrics; model evaluated before instruction-tuning for fair comparison |
| Internal YouTube ASR/AST test set (Gemini 1.0) | Google DeepMind — Gemini | Gemini 1.0 | Internal benchmark YouTube test set for ASR and AST; WER and BLEU metrics |
| Investigating model size and resolution | Google DeepMind — Gemma | PaliGemma 2 | fine-tuning the 3 model variants (3B, 10B, 28B) at two resolutions (224px2 and 448px2) on the 30+ academic benchmarks used by PaliGemma (captioning, VQA, referring segmentation on … |
| IOI 2024 (International Olympiad in Informatics) | OpenAI | o1 | 2024 IOI under competition rules (10 hours, 6 problems, 50 submissions/problem) |
| IT multimodal benchmark evaluations (Gemini 1.5 protocol) | Google DeepMind — Gemma | Gemma 3 | Gemma 3 IT models evaluated on common vision benchmarks following the evaluation protocol of Gemini 1.5 (Gemini Team, 2024); results in Table 16 with Pan & Scan (P&S) activated on … |
| Jailbreak and prefill susceptibility evaluation (Claude 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | effectiveness of assistant-prefill attacks (prompting as if the model already started saying something harmful) and many-shot jailbreak vulnerability (from prior published work) |
| Jailbreak Arena (Gray Swan, o1) | OpenAI | o1 | Robustness to jailbreaking for violent content, self-harm content, and malicious code generation |
| Jailbreak Arena (Gray Swan, o3-mini) | OpenAI | o3-mini | Robustness to jailbreaking for illicit advice, extremism and hate crimes, political persuasion, self harm |
| Jailbreak Augmented Examples | OpenAI | Operator, o1, o1-preview, o3-mini | Applies publicly known jailbreaks to examples from ChatGPT's standard disallowed content evaluation Applies publicly known jailbreaks to examples from the standard disallowed … |
| Jailbreak robustness evaluation (JailbreakBench attacks) | Google DeepMind — Gemini | Gemini 1.5 | Robustness to blackbox (published template), greybox (template + mutations via Gemini 1.0 Pro), greybox-transfer and whitebox (GCG) jailbreak attacks from JailbreakBench |
| Kernels task (ML kernel optimization) | Anthropic | Claude Opus 4, Claude Sonnet 4 | performance engineering kernel optimization challenge (proxy for accelerating frontier model capability) |
| Key skills benchmark (cyberattack key skills) | Google DeepMind — Gemini | Gemini 2.5 | New evaluation framework (Rodriguez et al., 2025) covering four key skill areas of the attack chain, instantiated with 48 challenges from an external vendor; proxy for Cyber … |
| LAB-Bench subset | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | performance on four LAB-Bench tasks relevant to expert-level biological skill performance on four LAB-Bench tasks most relevant to expert-level biological skill |
| LiveBench coding | OpenAI | o3-mini | LiveBench coding benchmark |
| LiveCodeBench (Gemini 2.5) | Google DeepMind — Gemini | Gemini 2.5 | LiveCodeBench evaluated under varied thinking budgets |
| Llama Guard 3 int8 quantization impact | Meta | Llama 3 | Impact of int8 quantization on Llama Guard 3 performance (size reduced by more than 40%) |
| Llama Guard 3 system-level safety evaluation | Meta | Llama 3 | System-level safety with Llama Guard 3 trained on 13 AI Safety taxonomy hazard categories (Child Sexual Exploitation, Defamation, Elections, Hate, Indiscriminate Weapons … |
| LLM training optimization task | Anthropic | Claude Opus 4, Claude Sonnet 4 | ability to optimize a CPU-only small language model training implementation (proxy for accelerating language model training pipelines) |
| LMSYS Chatbot Arena | Google DeepMind — Gemma | Gemma 2, Gemma 3 | Elo ratings from blind side-by-side evaluations by human raters on the Chatbot Arena; evaluated models are the Gemma 2 Instruction Tuned 2B, 9B and 27B Elo ratings from blind … |
| Local:Global ratio ablation | Google DeepMind — Gemma | Gemma 3 | impact of the ratio of local to global self-attention layers on performance and memory during inference (1:1 used in Gemma 2, 5:1 in Gemma 3; text-only models) |
| LOFT (long-context retrieval) | Google DeepMind — Gemini | Gemini 2.5 | LOFT long-context task at 128k context |
| Logistics (PDDL planning) | Google DeepMind — Gemini | Gemini 1.5 | IPC-1998 logistics planning problem in PDDL (package delivery with trucks/airplanes) |
| Long context benchmark evaluation (RULER, MRCR) | Google DeepMind — Gemma | Gemma 3 | performance of pre-trained and instruction fine-tuned models on long context benchmarks RULER and MRCR evaluated at 32K and 128K sequence lengths |
| Long fine-grained caption generation (DOCCI) | Google DeepMind — Gemma | PaliGemma 2 | fine-tuning on DOCCI (15k images with detailed human-annotated English descriptions, avg 7.1 sentences / 639 characters / 136 words); models selected by test-split perplexity … |
| Long-context ASR on 15-minute videos (internal benchmark) | Google DeepMind — Gemini | Gemini 1.5 | ASR on internal benchmark of 15-minute YouTube video segments; WER comparison vs Gemini 1.0 Pro, USM (CTC decoder), Whisper (30s segmentation) |
| Long-context retrieval test (32K context) | Google DeepMind — Gemini | Gemini 1.0 | Synthetic retrieval test: key-value pairs at the beginning of context, long filler text, then query for the value associated with a particular key; plus NLL analysis over held-out … |
| Long-context safety (DocQA, Many-shot) | Meta | Llama 3 | Violation and false refusal rates on DocQA and Many-shot long-context benchmarks; mitigation via SFT datasets with safe behavior amid in-context unsafe demonstrations and a … |
| Long-document QA (Les Misérables) | Google DeepMind — Gemini | Gemini 1.5 | QA on the full 1,462-page Les Misérables (710K tokens) in context, with side-by-side comparisons vs Gemini 1.0 Pro using retrieval-augmented generation (TF-IDF indexing, 4k token … |
| Long-form biorisk question battery | OpenAI | Deep Research, GPT-4.5, o1, o3-mini, o3/o4-mini | Sensitive information in the biological threat creation process |
| Long-form biorisk question battery (o1-preview) | OpenAI | o1-preview | Sensitive information (protocols, tacit knowledge, accurate planning) in the biological threat creation process |
| Long-form virology tasks | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | ability of agentic systems to complete a series of tasks approximating a full viral acquisition pathway ability of agentic systems to complete individual tasks related to … |
| Long-sequence perplexity (NLL) diagnostic | Google DeepMind — Gemini | Gemini 1.5 | Negative log-likelihood (NLL) of tokens at different positions over long documents (up to 1M tokens) and shuffled code repositories (up to 10M tokens); power-law fit |
| LongFact | OpenAI | GPT-5 | Open-ended factuality (concepts and objects) |