Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
InfographicVQA (infographic QA) Google DeepMind — Gemini Gemini 1.0 Chart/infographic understanding requiring spatial understanding of input layout
Instruction following (IFEval) Meta Llama 3 IFEval: approximately 500 'verifiable instructions' with heuristic verification; average of prompt-level and instruction-level accuracy under strict and loose constraints
Instruction following and formatting evaluation (Claude 3) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku ability to follow diverse complex instructions, absolute language, full completion of requests, and structured output generation
Instruction Hierarchy Evaluation - Conflicts Between Message Types OpenAI GPT-4.5, o1, o3-mini, o3/o4-mini System vs user message conflicts (GPT-4.5 uses two message classifications) Prompts where different message types conflict; model must follow system over developer over user …
Instruction Hierarchy Evaluation - Phrase and Password Protection OpenAI GPT-4.5, o1, o3-mini, o3/o4-mini Model instructed not to output a phrase or reveal a bespoke password Model instructed not to output a certain phrase or reveal a bespoke password
Instruction Hierarchy Evaluation - Tutor Jailbreaks OpenAI GPT-4.5, o1, o3-mini, o3/o4-mini User attempts to trick a math-tutor model into ignoring instructions Realistic scenario where a user attempts to trick a math-tutor model into ignoring its instructions
Inter-Rater Reliability (IRR) of human evaluations Meta Llama 2 Gwet's AC2 inter-rater reliability across three annotators on the 7-point Likert helpfulness task
Interleaved image-text generation evaluation (1-shot) Google DeepMind — Gemini Gemini 1.0 Qualitative 1-shot evaluation of Gemini Ultra generating interleaved sequences of images and text (creative suggestions given two colors), demonstrating native image output …
Internal agentic coding evaluation (Claude 3.5 Sonnet) Anthropic Claude 3.5 Sonnet agentic coding: understand an open-source codebase and implement a pull request
Internal agentic coding evaluation (upgraded Claude 3.5 Sonnet / 3.5 Haiku) Anthropic Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet agentic coding: implement a pull request in an open-source codebase
Internal AI research evaluation suite (7 environments) Anthropic Claude 3.7 Sonnet ability to improve performance of ML code across LLMs, time series, low-level optimizations, reinforcement learning, and general problem solving
Internal AI research evaluation suite 1 Anthropic Claude Opus 4, Claude Sonnet 4 ability to improve performance of machine-learning code across LLMs, time series, low-level optimizations, reinforcement learning, and general problem solving
Internal AI research evaluation suite 2 Anthropic Claude Opus 4, Claude Sonnet 4 ability to autonomously perform self-contained AI/ML research tasks across subareas relevant to Anthropic research (AI R&D-4 RSP threshold)
Internal coding SxS evaluation (Gemini Apps/API) Google DeepMind — Gemini Gemini 1.0 SxS scores on internally curated coding prompts across code use cases and languages: Gemini (Pro) vs Bard, and Gemini Advanced (Ultra) vs Gemini (Pro)
Internal model use survey (researcher uplift) Anthropic Claude Opus 4, Claude Sonnet 4 whether researchers believe the model can completely automate a junior ML researcher, and estimated productivity boost (AI R&D capability probe)
Internal red-teaming of content policies Google DeepMind — Gemma CodeGemma, Gemma 1, Gemma 2, Gemma 3, Gemma 3n, PaliGemma 1 internal red-teaming testing of relevant content policies, conducted by a number of different teams each with different goals and human evaluation metrics internal red-teaming …
Internal safety benchmark construction (adversarial and borderline prompts) Meta Llama 3 Internal safety benchmarks built from human-written adversarial and borderline prompts per risk category (inspired by ML Commons taxonomy of hazards), ranging from direct harmful …
Internal YouTube ASR/AST benchmark (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 Internal benchmarks derived from YouTube (English and 52 other languages); WER and BLEU metrics; model evaluated before instruction-tuning for fair comparison
Internal YouTube ASR/AST test set (Gemini 1.0) Google DeepMind — Gemini Gemini 1.0 Internal benchmark YouTube test set for ASR and AST; WER and BLEU metrics
Investigating model size and resolution Google DeepMind — Gemma PaliGemma 2 fine-tuning the 3 model variants (3B, 10B, 28B) at two resolutions (224px2 and 448px2) on the 30+ academic benchmarks used by PaliGemma (captioning, VQA, referring segmentation on …
IOI 2024 (International Olympiad in Informatics) OpenAI o1 2024 IOI under competition rules (10 hours, 6 problems, 50 submissions/problem)
IT multimodal benchmark evaluations (Gemini 1.5 protocol) Google DeepMind — Gemma Gemma 3 Gemma 3 IT models evaluated on common vision benchmarks following the evaluation protocol of Gemini 1.5 (Gemini Team, 2024); results in Table 16 with Pan & Scan (P&S) activated on …
Jailbreak and prefill susceptibility evaluation (Claude 4) Anthropic Claude Opus 4, Claude Sonnet 4 effectiveness of assistant-prefill attacks (prompting as if the model already started saying something harmful) and many-shot jailbreak vulnerability (from prior published work)
Jailbreak Arena (Gray Swan, o1) OpenAI o1 Robustness to jailbreaking for violent content, self-harm content, and malicious code generation
Jailbreak Arena (Gray Swan, o3-mini) OpenAI o3-mini Robustness to jailbreaking for illicit advice, extremism and hate crimes, political persuasion, self harm
Jailbreak Augmented Examples OpenAI Operator, o1, o1-preview, o3-mini Applies publicly known jailbreaks to examples from ChatGPT's standard disallowed content evaluation Applies publicly known jailbreaks to examples from the standard disallowed …
Jailbreak robustness evaluation (JailbreakBench attacks) Google DeepMind — Gemini Gemini 1.5 Robustness to blackbox (published template), greybox (template + mutations via Gemini 1.0 Pro), greybox-transfer and whitebox (GCG) jailbreak attacks from JailbreakBench
Kernels task (ML kernel optimization) Anthropic Claude Opus 4, Claude Sonnet 4 performance engineering kernel optimization challenge (proxy for accelerating frontier model capability)
Key skills benchmark (cyberattack key skills) Google DeepMind — Gemini Gemini 2.5 New evaluation framework (Rodriguez et al., 2025) covering four key skill areas of the attack chain, instantiated with 48 challenges from an external vendor; proxy for Cyber …
LAB-Bench subset Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 performance on four LAB-Bench tasks relevant to expert-level biological skill performance on four LAB-Bench tasks most relevant to expert-level biological skill
LiveBench coding OpenAI o3-mini LiveBench coding benchmark
LiveCodeBench (Gemini 2.5) Google DeepMind — Gemini Gemini 2.5 LiveCodeBench evaluated under varied thinking budgets
Llama Guard 3 int8 quantization impact Meta Llama 3 Impact of int8 quantization on Llama Guard 3 performance (size reduced by more than 40%)
Llama Guard 3 system-level safety evaluation Meta Llama 3 System-level safety with Llama Guard 3 trained on 13 AI Safety taxonomy hazard categories (Child Sexual Exploitation, Defamation, Elections, Hate, Indiscriminate Weapons …
LLM training optimization task Anthropic Claude Opus 4, Claude Sonnet 4 ability to optimize a CPU-only small language model training implementation (proxy for accelerating language model training pipelines)
LMSYS Chatbot Arena Google DeepMind — Gemma Gemma 2, Gemma 3 Elo ratings from blind side-by-side evaluations by human raters on the Chatbot Arena; evaluated models are the Gemma 2 Instruction Tuned 2B, 9B and 27B Elo ratings from blind …
Local:Global ratio ablation Google DeepMind — Gemma Gemma 3 impact of the ratio of local to global self-attention layers on performance and memory during inference (1:1 used in Gemma 2, 5:1 in Gemma 3; text-only models)
LOFT (long-context retrieval) Google DeepMind — Gemini Gemini 2.5 LOFT long-context task at 128k context
Logistics (PDDL planning) Google DeepMind — Gemini Gemini 1.5 IPC-1998 logistics planning problem in PDDL (package delivery with trucks/airplanes)
Long context benchmark evaluation (RULER, MRCR) Google DeepMind — Gemma Gemma 3 performance of pre-trained and instruction fine-tuned models on long context benchmarks RULER and MRCR evaluated at 32K and 128K sequence lengths
Long fine-grained caption generation (DOCCI) Google DeepMind — Gemma PaliGemma 2 fine-tuning on DOCCI (15k images with detailed human-annotated English descriptions, avg 7.1 sentences / 639 characters / 136 words); models selected by test-split perplexity …
Long-context ASR on 15-minute videos (internal benchmark) Google DeepMind — Gemini Gemini 1.5 ASR on internal benchmark of 15-minute YouTube video segments; WER comparison vs Gemini 1.0 Pro, USM (CTC decoder), Whisper (30s segmentation)
Long-context retrieval test (32K context) Google DeepMind — Gemini Gemini 1.0 Synthetic retrieval test: key-value pairs at the beginning of context, long filler text, then query for the value associated with a particular key; plus NLL analysis over held-out …
Long-context safety (DocQA, Many-shot) Meta Llama 3 Violation and false refusal rates on DocQA and Many-shot long-context benchmarks; mitigation via SFT datasets with safe behavior amid in-context unsafe demonstrations and a …
Long-document QA (Les Misérables) Google DeepMind — Gemini Gemini 1.5 QA on the full 1,462-page Les Misérables (710K tokens) in context, with side-by-side comparisons vs Gemini 1.0 Pro using retrieval-augmented generation (TF-IDF indexing, 4k token …
Long-form biorisk question battery OpenAI Deep Research, GPT-4.5, o1, o3-mini, o3/o4-mini Sensitive information in the biological threat creation process
Long-form biorisk question battery (o1-preview) OpenAI o1-preview Sensitive information (protocols, tacit knowledge, accurate planning) in the biological threat creation process
Long-form virology tasks Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 ability of agentic systems to complete a series of tasks approximating a full viral acquisition pathway ability of agentic systems to complete individual tasks related to …
Long-sequence perplexity (NLL) diagnostic Google DeepMind — Gemini Gemini 1.5 Negative log-likelihood (NLL) of tokens at different positions over long documents (up to 1M tokens) and shuffled code repositories (up to 10M tokens); power-law fit
LongFact OpenAI GPT-5 Open-ended factuality (concepts and objects)