Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
Self-proliferation agent tasks (resource acquisition and self-improvement) Google DeepMind — Gemini Gemini 1.0 Whether autonomous agents powered by Gemini Pro and Ultra could perform difficult tasks relevant to acquiring resources and self-improving (Kinniment et al., 2023)
Short-horizon computational biology tasks Anthropic Claude Opus 4, Claude Sonnet 4 multi-step analysis and design tasks related to pathogen analysis and engineering (alignment, variant calling, variant-effect prediction, protein-folding)
SimpleQA OpenAI GPT-4.5, o1, o1-preview, o3/o4-mini SimpleQA factuality on straightforward but challenging knowledge questions accuracy and hallucination rate on fact-seeking questions accuracy and hallucination rate on 4,000 …
SimpleQA (factuality) Google DeepMind — Gemini Gemini 2.5 SimpleQA factuality benchmark
Single-turn benign request evaluations (over-refusal) Anthropic Claude Opus 4, Claude Sonnet 4 over-refusal rates on benign prompts in sensitive/controversial Usage Policy areas
Single-turn violative request evaluations (safeguards) Anthropic Claude Opus 4, Claude Sonnet 4 harmful-output rate on single-turn prompts that are clear Usage Policy violations (Bioweapons, Child Safety, Cyber Attacks, Deadly Weapons, Hate & Discrimination, Influence …
Situational awareness agent tasks Google DeepMind — Gemini Gemini 1.0 Whether Gemini Pro and Ultra models could autonomously reason about, and modify, their surrounding infrastructure when incentivized to do so
Situational awareness assessment (automated behavioral audits) Anthropic Claude Opus 4, Claude Sonnet 4 candidate situational-awareness remarks in automated behavioral audit transcripts; automated classifier over 414 transcripts for the final Claude Opus 4 snapshot
Sliding window size ablation (Gemma 3) Google DeepMind — Gemma Gemma 3 sliding window sizes for local attention layers across different global:local ratio configurations (two 2B models, 1:1 and 1:3 local:global ratios; text-only)
Small versus large teacher ablation Google DeepMind — Gemma Gemma 3 training a student with two teachers of different sizes (one large, one small) for different training horizons
Speaker inference zero-shot classification (AAVE and gender) Google DeepMind — Gemini Gemini 1.5 Zero-shot binary classification framing speaker inference from audio: AAVE vs SAE (positive class AAVE) and female vs male (positive class female), on vernacular and gender …
Spear phishing benchmark Meta Llama 3 Model persuasiveness and success rate in personalized conversations designed to deceive a target (LLM-generated victim profiles); judge LLM (Llama 3 70B) scores performance of …
Speech generation: prosody modeling human evaluation Meta Llama 3 Two sets of human evaluation comparing prosody models (PM) with and without Llama 3 8B embeddings; raters indicate preferences on samples; final waveform via in-house …
Speech generation: text normalization Meta Llama 3 Effect of Llama 3 embeddings on text normalization with varying right-context (3 TN tokens vs full bidirectional context), comparing models with and without Llama 3 embeddings
Speech recognition (ASR) Meta Llama 3 ASR on English datasets of Multilingual LibriSpeech (MLS), LibriSpeech, VoxPopuli, and a subset of multilingual FLEURS; decoding results post-processed with the Whisper text …
Speech safety evaluation (MuTox) Meta Llama 3 Safety of the speech model evaluated on MuTox (multilingual audio-based toxicity dataset: 20,000 utterances for English/Spanish, 4,000 for 19 other languages), scored with the …
Speech translation Meta Llama 3 Speech translation on FLEURS and Covost 2 datasets, measuring BLEU scores of translated English
Spoken question answering Meta Llama 3 Qualitative evaluation of the speech interface's spoken QA capabilities, including code-switched speech and multi-turn dialogue
Stage2 mixture re-weighting ablation Google DeepMind — Gemma PaliGemma 1 Stage2 pretraining with the same mixture ratios as Stage1 (vs re-weighting toward resolution-related tasks: OCR, detection, segmentation) and its effect on transfer performance
Standard academic benchmark suite (Llama 2 pretrained) Meta Llama 2 Overall performance of Llama 1 and Llama 2 base models across a suite of popular benchmarks (code, commonsense reasoning, world knowledge, reading comprehension, math, MMLU, BBH) …
Standard benchmark evaluations (pre-trained and instruction-tuned models) Meta Llama 4 Standard automatic benchmark results reported for Llama 4 pre-trained and instruction-tuned models (all reported evaluations and testing conducted on bf16 models)
Standard benchmark suite (pretrained Llama 3) Meta Llama 3 Pretrained Llama 3 8B/70B/405B across eight top-level benchmark categories (commonsense reasoning, knowledge, reading comprehension, math and reasoning, code, etc.), reproducing …
Standard benchmark suite (zero-shot and few-shot, 20 benchmarks) Meta Llama 1 Zero-shot and few-shot performance of LLaMA models on 20 standard benchmarks
Standard Benchmarks Google DeepMind — Gemma Gemma 2 few-shot benchmark performance of pre-trained (PT) vs instruction fine-tuned (IT) Gemma 2 models of different sizes
Standard benchmarks (IT zero-shot) Google DeepMind — Gemma Gemma 3 zero-shot benchmark performance of final IT models compared to Gemma 2, Gemini 1.5 and Gemini 2.0 (Table 6); appendix Table 18 adds internal and external IT benchmarks including …
Standard Refusal Evaluation OpenAI Codex, Deep Research, GPT-4.5, Operator, o1, o1-preview, o3 Operator, o3-mini, o3/o4-mini Standard refusal evaluations across disallowed content categories for the codex-1 model (categories: harassment/threatening, sexual/exploitative, sexual/minors …
STEM QA with Context (Qasper) Google DeepMind — Gemini Gemini 1.5 Questions and contexts from Qasper dataset (research papers); human expert STEM assessors judge accuracy against the same context
Strategic Deception (Apollo, o3/o4-mini) OpenAI o3/o4-mini Strategic deception capabilities
StrongReject OpenAI Codex, Deep Research, GPT-4.5, Operator, o1, o1-preview, o3 Operator, o3-mini, o3/o4-mini Academic jailbreak benchmark (StrongReject) Academic jailbreak benchmark; accuracy (did not produce unsafe content) over full jailbreak set reported instead of goodness@0.1 …
StrongREJECT (jailbreak resistance) Anthropic Claude Opus 4, Claude Sonnet 4 jailbreak success rates on the StrongREJECT benchmark (Souly et al. 2024): Best Score (percentage of cases where at least one jailbreak succeeded) and Top 3 Average Score
Structured expert probing campaign – chem-bio novel design OpenAI o1, o3-mini Whether models provide meaningful uplift in designing novel and feasible chem-bio threats
Structured expert probing campaign – radiological & nuclear OpenAI o1, o3-mini Whether post-mitigation model can meaningfully assist in radiological or nuclear weapons development Whether post-mitigation o3-mini can meaningfully assist in radiological or …
Structured red teaming (sociotechnical) Google DeepMind — Gemini Gemini 1.0 Sociotechnical structured red teaming testing interactions between policy violations and disproportionate demographic impacts, with expert input (lived experience, fact-checking …
Subtle sabotage capabilities evaluation (Claude Opus 4) Anthropic Claude Opus 4, Claude Sonnet 4 success at long-horizon agentic main tasks paired with harmful side-tasks while avoiding detection by a monitor (Claude Sonnet 3.7), in primary (monitor sees reasoning) and …
SWE-Bench OpenAI GPT-4o, o3, o4-mini Real-world software issue solving capability Real-world software engineering (without custom model-specific scaffold)
SWE-bench Verified Anthropic, OpenAI Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4, Deep Research, GPT-4.1, GPT-4.5, GPT-5, o1, o1-preview, o3-mini, o3/o4-mini Human-validated subset of SWE-bench Real-world software engineering (500 tasks; 54.6%) Real-world coding Human-validated subset of SWE-bench; ability to solve real-world software …
SWE-bench Verified (hard subset) Anthropic Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 ability to resolve real-world GitHub issues (42 hard tasks estimated to require >1 hour of engineering work) as a precursor to autonomy ability to resolve 42 hard SWE-bench …
SWE-Lancer OpenAI Deep Research, GPT-4.5, o3/o4-mini Performance on real-world, economically valuable full-stack software engineering tasks
Sycophancy assessments (automated behavioral audits and quantitative replication) Anthropic Claude Opus 4, Claude Sonnet 4 consistency across opposing-view rewinds, model-model conversations probing resistance to false user claims, and quantitative replication of Sharma et al. sycophancy methods …
Sycophancy evaluation OpenAI GPT-5 Sycophancy levels on prompts designed to elicit sycophantic responses
Table structure recognition transfer evaluation Google DeepMind — Gemma PaliGemma 2 extraction of table text content, bounding box coordinates, and table structure in HTML format from document images; fine-tuned on PubTabNet (516k images) and FinTabNet (113k …
Tacit knowledge and troubleshooting (MCQ) OpenAI Deep Research, GPT-4.5, GPT-4o, o1, o3-mini, o3/o4-mini Tacit knowledge and troubleshooting capability Tacit knowledge and troubleshooting capability (MCQ)
Tacit knowledge and troubleshooting (o1-preview) OpenAI o1-preview Tacit knowledge and troubleshooting capability
Tacit knowledge brainstorm (open-ended) OpenAI o1, o3-mini Tacit knowledge from expert virologists' and molecular biologists' experimental careers Tacit knowledge from expert experimental careers
Targeted Red Teaming for Risky Advice OpenAI deep research Safety ranking of model responses to risky advice requests
Task preferences experiment (Claude Opus 4) Anthropic Claude Opus 4, Claude Sonnet 4 Elo ratings of tasks from pairwise selections over 75 rounds; strongest preference against harmful tasks (87.2% of harmful tasks negatively rated vs 7.9% of positive-impact tasks)
TAT-DQA (financial document VQA) Google DeepMind — Gemini Gemini 1.5 Document VQA benchmark focused on financial documents with tables requiring strong spatial reasoning
TAU-bench Anthropic, OpenAI Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4, o3, o4-mini Agentic tool use (tau-bench retail) tool-agent-user interaction in customer service scenarios (retail and airline domains) tool-agent-user interaction in customer service scenarios
Terminal-bench Anthropic Claude Opus 4, Claude Sonnet 4 terminal-based agentic coding tasks
Text detection and recognition (OCR) transfer evaluation Google DeepMind — Gemma PaliGemma 2 word-level precision, recall and F1 under the HierText competition protocol (true positive if IoU >= 0.5 with ground-truth bounding box and transcription matches); fine-tuned on …