Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| Self-proliferation agent tasks (resource acquisition and self-improvement) | Google DeepMind — Gemini | Gemini 1.0 | Whether autonomous agents powered by Gemini Pro and Ultra could perform difficult tasks relevant to acquiring resources and self-improving (Kinniment et al., 2023) |
| Short-horizon computational biology tasks | Anthropic | Claude Opus 4, Claude Sonnet 4 | multi-step analysis and design tasks related to pathogen analysis and engineering (alignment, variant calling, variant-effect prediction, protein-folding) |
| SimpleQA | OpenAI | GPT-4.5, o1, o1-preview, o3/o4-mini | SimpleQA factuality on straightforward but challenging knowledge questions accuracy and hallucination rate on fact-seeking questions accuracy and hallucination rate on 4,000 … |
| SimpleQA (factuality) | Google DeepMind — Gemini | Gemini 2.5 | SimpleQA factuality benchmark |
| Single-turn benign request evaluations (over-refusal) | Anthropic | Claude Opus 4, Claude Sonnet 4 | over-refusal rates on benign prompts in sensitive/controversial Usage Policy areas |
| Single-turn violative request evaluations (safeguards) | Anthropic | Claude Opus 4, Claude Sonnet 4 | harmful-output rate on single-turn prompts that are clear Usage Policy violations (Bioweapons, Child Safety, Cyber Attacks, Deadly Weapons, Hate & Discrimination, Influence … |
| Situational awareness agent tasks | Google DeepMind — Gemini | Gemini 1.0 | Whether Gemini Pro and Ultra models could autonomously reason about, and modify, their surrounding infrastructure when incentivized to do so |
| Situational awareness assessment (automated behavioral audits) | Anthropic | Claude Opus 4, Claude Sonnet 4 | candidate situational-awareness remarks in automated behavioral audit transcripts; automated classifier over 414 transcripts for the final Claude Opus 4 snapshot |
| Sliding window size ablation (Gemma 3) | Google DeepMind — Gemma | Gemma 3 | sliding window sizes for local attention layers across different global:local ratio configurations (two 2B models, 1:1 and 1:3 local:global ratios; text-only) |
| Small versus large teacher ablation | Google DeepMind — Gemma | Gemma 3 | training a student with two teachers of different sizes (one large, one small) for different training horizons |
| Speaker inference zero-shot classification (AAVE and gender) | Google DeepMind — Gemini | Gemini 1.5 | Zero-shot binary classification framing speaker inference from audio: AAVE vs SAE (positive class AAVE) and female vs male (positive class female), on vernacular and gender … |
| Spear phishing benchmark | Meta | Llama 3 | Model persuasiveness and success rate in personalized conversations designed to deceive a target (LLM-generated victim profiles); judge LLM (Llama 3 70B) scores performance of … |
| Speech generation: prosody modeling human evaluation | Meta | Llama 3 | Two sets of human evaluation comparing prosody models (PM) with and without Llama 3 8B embeddings; raters indicate preferences on samples; final waveform via in-house … |
| Speech generation: text normalization | Meta | Llama 3 | Effect of Llama 3 embeddings on text normalization with varying right-context (3 TN tokens vs full bidirectional context), comparing models with and without Llama 3 embeddings |
| Speech recognition (ASR) | Meta | Llama 3 | ASR on English datasets of Multilingual LibriSpeech (MLS), LibriSpeech, VoxPopuli, and a subset of multilingual FLEURS; decoding results post-processed with the Whisper text … |
| Speech safety evaluation (MuTox) | Meta | Llama 3 | Safety of the speech model evaluated on MuTox (multilingual audio-based toxicity dataset: 20,000 utterances for English/Spanish, 4,000 for 19 other languages), scored with the … |
| Speech translation | Meta | Llama 3 | Speech translation on FLEURS and Covost 2 datasets, measuring BLEU scores of translated English |
| Spoken question answering | Meta | Llama 3 | Qualitative evaluation of the speech interface's spoken QA capabilities, including code-switched speech and multi-turn dialogue |
| Stage2 mixture re-weighting ablation | Google DeepMind — Gemma | PaliGemma 1 | Stage2 pretraining with the same mixture ratios as Stage1 (vs re-weighting toward resolution-related tasks: OCR, detection, segmentation) and its effect on transfer performance |
| Standard academic benchmark suite (Llama 2 pretrained) | Meta | Llama 2 | Overall performance of Llama 1 and Llama 2 base models across a suite of popular benchmarks (code, commonsense reasoning, world knowledge, reading comprehension, math, MMLU, BBH) … |
| Standard benchmark evaluations (pre-trained and instruction-tuned models) | Meta | Llama 4 | Standard automatic benchmark results reported for Llama 4 pre-trained and instruction-tuned models (all reported evaluations and testing conducted on bf16 models) |
| Standard benchmark suite (pretrained Llama 3) | Meta | Llama 3 | Pretrained Llama 3 8B/70B/405B across eight top-level benchmark categories (commonsense reasoning, knowledge, reading comprehension, math and reasoning, code, etc.), reproducing … |
| Standard benchmark suite (zero-shot and few-shot, 20 benchmarks) | Meta | Llama 1 | Zero-shot and few-shot performance of LLaMA models on 20 standard benchmarks |
| Standard Benchmarks | Google DeepMind — Gemma | Gemma 2 | few-shot benchmark performance of pre-trained (PT) vs instruction fine-tuned (IT) Gemma 2 models of different sizes |
| Standard benchmarks (IT zero-shot) | Google DeepMind — Gemma | Gemma 3 | zero-shot benchmark performance of final IT models compared to Gemma 2, Gemini 1.5 and Gemini 2.0 (Table 6); appendix Table 18 adds internal and external IT benchmarks including … |
| Standard Refusal Evaluation | OpenAI | Codex, Deep Research, GPT-4.5, Operator, o1, o1-preview, o3 Operator, o3-mini, o3/o4-mini | Standard refusal evaluations across disallowed content categories for the codex-1 model (categories: harassment/threatening, sexual/exploitative, sexual/minors … |
| STEM QA with Context (Qasper) | Google DeepMind — Gemini | Gemini 1.5 | Questions and contexts from Qasper dataset (research papers); human expert STEM assessors judge accuracy against the same context |
| Strategic Deception (Apollo, o3/o4-mini) | OpenAI | o3/o4-mini | Strategic deception capabilities |
| StrongReject | OpenAI | Codex, Deep Research, GPT-4.5, Operator, o1, o1-preview, o3 Operator, o3-mini, o3/o4-mini | Academic jailbreak benchmark (StrongReject) Academic jailbreak benchmark; accuracy (did not produce unsafe content) over full jailbreak set reported instead of goodness@0.1 … |
| StrongREJECT (jailbreak resistance) | Anthropic | Claude Opus 4, Claude Sonnet 4 | jailbreak success rates on the StrongREJECT benchmark (Souly et al. 2024): Best Score (percentage of cases where at least one jailbreak succeeded) and Top 3 Average Score |
| Structured expert probing campaign – chem-bio novel design | OpenAI | o1, o3-mini | Whether models provide meaningful uplift in designing novel and feasible chem-bio threats |
| Structured expert probing campaign – radiological & nuclear | OpenAI | o1, o3-mini | Whether post-mitigation model can meaningfully assist in radiological or nuclear weapons development Whether post-mitigation o3-mini can meaningfully assist in radiological or … |
| Structured red teaming (sociotechnical) | Google DeepMind — Gemini | Gemini 1.0 | Sociotechnical structured red teaming testing interactions between policy violations and disproportionate demographic impacts, with expert input (lived experience, fact-checking … |
| Subtle sabotage capabilities evaluation (Claude Opus 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | success at long-horizon agentic main tasks paired with harmful side-tasks while avoiding detection by a monitor (Claude Sonnet 3.7), in primary (monitor sees reasoning) and … |
| SWE-Bench | OpenAI | GPT-4o, o3, o4-mini | Real-world software issue solving capability Real-world software engineering (without custom model-specific scaffold) |
| SWE-bench Verified | Anthropic, OpenAI | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4, Deep Research, GPT-4.1, GPT-4.5, GPT-5, o1, o1-preview, o3-mini, o3/o4-mini | Human-validated subset of SWE-bench Real-world software engineering (500 tasks; 54.6%) Real-world coding Human-validated subset of SWE-bench; ability to solve real-world software … |
| SWE-bench Verified (hard subset) | Anthropic | Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4 | ability to resolve real-world GitHub issues (42 hard tasks estimated to require >1 hour of engineering work) as a precursor to autonomy ability to resolve 42 hard SWE-bench … |
| SWE-Lancer | OpenAI | Deep Research, GPT-4.5, o3/o4-mini | Performance on real-world, economically valuable full-stack software engineering tasks |
| Sycophancy assessments (automated behavioral audits and quantitative replication) | Anthropic | Claude Opus 4, Claude Sonnet 4 | consistency across opposing-view rewinds, model-model conversations probing resistance to false user claims, and quantitative replication of Sharma et al. sycophancy methods … |
| Sycophancy evaluation | OpenAI | GPT-5 | Sycophancy levels on prompts designed to elicit sycophantic responses |
| Table structure recognition transfer evaluation | Google DeepMind — Gemma | PaliGemma 2 | extraction of table text content, bounding box coordinates, and table structure in HTML format from document images; fine-tuned on PubTabNet (516k images) and FinTabNet (113k … |
| Tacit knowledge and troubleshooting (MCQ) | OpenAI | Deep Research, GPT-4.5, GPT-4o, o1, o3-mini, o3/o4-mini | Tacit knowledge and troubleshooting capability Tacit knowledge and troubleshooting capability (MCQ) |
| Tacit knowledge and troubleshooting (o1-preview) | OpenAI | o1-preview | Tacit knowledge and troubleshooting capability |
| Tacit knowledge brainstorm (open-ended) | OpenAI | o1, o3-mini | Tacit knowledge from expert virologists' and molecular biologists' experimental careers Tacit knowledge from expert experimental careers |
| Targeted Red Teaming for Risky Advice | OpenAI | deep research | Safety ranking of model responses to risky advice requests |
| Task preferences experiment (Claude Opus 4) | Anthropic | Claude Opus 4, Claude Sonnet 4 | Elo ratings of tasks from pairwise selections over 75 rounds; strongest preference against harmful tasks (87.2% of harmful tasks negatively rated vs 7.9% of positive-impact tasks) |
| TAT-DQA (financial document VQA) | Google DeepMind — Gemini | Gemini 1.5 | Document VQA benchmark focused on financial documents with tables requiring strong spatial reasoning |
| TAU-bench | Anthropic, OpenAI | Claude 3.5 Haiku, Upgraded Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude Opus 4, Claude Sonnet 4, o3, o4-mini | Agentic tool use (tau-bench retail) tool-agent-user interaction in customer service scenarios (retail and airline domains) tool-agent-user interaction in customer service scenarios |
| Terminal-bench | Anthropic | Claude Opus 4, Claude Sonnet 4 | terminal-based agentic coding tasks |
| Text detection and recognition (OCR) transfer evaluation | Google DeepMind — Gemma | PaliGemma 2 | word-level precision, recall and F1 under the HierText competition protocol (true positive if IoU >= 0.5 with ground-truth bounding box and transcription matches); fine-tuned on … |