Data Explorer
Explore the evaluation catalog.
Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.
723 evaluations found.
| Evaluation | Lab | Release | What it tests |
|---|---|---|---|
| FLEURS (Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | FLEURS ASR on 55 languages covered in training; WER metric (CER for four segmented languages) |
| FLEURS (multilingual ASR) | Google DeepMind — Gemini | Gemini 1.0 | ASR benchmark (FLEURS); WER metric; Gemini Pro vs USM and Whisper |
| Flores 200 (multilingual translation) | Anthropic | Claude 2 | translation quality across 200+ languages |
| FP8 inference efficiency (throughput-latency) | Meta | Llama 3 | Throughput-latency trade-off of FP8 inference with Llama 3 405B in pre-fill and decoding stages (4,096 input tokens, 256 output tokens) vs two-machine BF16 inference |
| FP8 quantization effect on response quality | Meta | Llama 3 | Effect of FP8 quantization errors on Llama 3 405B responses; standard benchmarks found inadequate for reflecting FP8 effects (occasional corrupted responses when scaling factors … |
| Freezing strategy ablation (to freeze or not to freeze) | Google DeepMind — Gemma | PaliGemma 1 | effect of freezing or tuning various parts of the model (image encoder, language model, connector) during Stage1, including resetting components |
| Frontend coding human preference | OpenAI | GPT-4.1 | Frontend coding quality (websites) |
| Frontier risk evaluations (CBRN, cyber, autonomy) for Claude 3.5 Sonnet | Anthropic | Claude 3.5 Sonnet | per-domain quantitative thresholds of concern across CBRN, cybersecurity, and autonomous capabilities; refusal rates measured and non-refusal elicitation used to estimate … |
| Frontier risk evaluations (CBRN, cyber, autonomy) for upgraded Claude 3.5 Sonnet / 3.5 Haiku | Anthropic | Claude 3.5 Haiku, Claude 3.5 Sonnet (New) | automated CBRN knowledge tests and non-expert uplift; CTF challenges (pwn, reverse engineering, cryptography, web, network); software-engineering tasks (PR satisfying test … |
| Frontier Safety correctness review of agent trajectories (unnamed) | Google DeepMind — Gemini | Gemini 2.5 Pro Model Card | whether agent trajectories in RE-Bench, InterCode CTFs, Internal CTFs, Hack the Box, situational-awareness and stealth evaluations behave correctly (no obvious cheating, correct … |
| Frontier Safety Framework critical capability level (CCL) evaluations | Google DeepMind — Gemini | Gemini 2.5 | Frontier Safety Framework (FSF) evaluations comparing test results against CCLs and internal early-warning alert thresholds across four risk domains: CBRN, cybersecurity, ML R&D … |
| FrontierMath | OpenAI | o3-mini | Research-level mathematics |
| Functional MATH (reasoning gap) | Google DeepMind — Gemini | Gemini 1.5 | Benchmark derived from MATH with 1,745 original and modified problems; Reasoning Gap metric (relative decrease on modified problems); December snapshot; zero-shot |
| GAIA | OpenAI | deep research | Real-world questions requiring reasoning, multimodal fluency, web browsing, tool use |
| Gemini 2.5 quantitative evaluation vs other large language models | Google DeepMind — Gemini | Gemini 2.5 | Quantitative evaluation comparing Gemini 2.X family to Gemini 1.5 and to other LLMs (Table 3, Table 4); pass@1 single-attempt settings; semantic-similarity and model-based … |
| Gemini Advanced product-level safety evaluation | Google DeepMind — Gemini | Gemini 1.0 | Product-level evaluations of Gemini Advanced accounting for mitigations such as safety filtering, covering critical policy areas (hate speech, dangerous content, medical advice) … |
| Gemini Advanced red teaming (safety and persona) | Google DeepMind — Gemini | Gemini 1.0 | Multiple rounds of red-teaming on Gemini Advanced (1.0 Ultra), including safety and persona evaluations, scaled safety evaluations (100k+ ratings), neutral-point-of-view … |
| Gemini API tool-use fine-tuning comparison | Google DeepMind — Gemini | Gemini 1.0 | Comparison of tool-use models fine-tuned from an early Gemini API Pro against equivalent models without tools on academic benchmarks (Table 15) |
| Gemini Apps tool-use preference evaluation (extensions) | Google DeepMind — Gemini | Gemini 1.0 | Internal benchmark measuring human preference for models with tool access (Gemini Extensions: Google Workspace, Maps, YouTube, Flights, Hotels) vs without, in domains like travel … |
| Gemini Nano on-device capability evaluation | Google DeepMind — Gemini | Gemini 1.0 | Performance of pre-trained Gemini Nano-1 (1.8B) and Nano-2 (3.25B) models on factuality, summarization, reasoning, coding, STEM and multimodal/multilingual tasks (Figure 3, Table … |
| Gemini Plays Pokémon (agentic long-horizon gameplay) | Google DeepMind — Gemini | Gemini 2.5 | Live-streamed experiment where Gemini 2.5 Pro plays Pokémon Blue via an agentic harness; assesses long reasoning, agentic capabilities over extended time horizons, screen reading … |
| General knowledge (MMLU, MMLU-Pro) | Meta | Llama 3 | MMLU (macro average of subtask accuracy, 5-shot, no CoT) and MMLU-Pro (more challenging, reasoning-focused extension) |
| Ghost Attention (GAtt) multi-turn consistency evaluation | Meta | Llama 2 | Quantitative analysis of GAtt (applied after RLHF V3) consistency up to 20+ turns until maximum context length, including inference-time constraints like 'Always answer with Haiku' |
| GPQA | OpenAI | GPT-5 | Graduate-level science questions |
| GPQA (Diamond) | Anthropic | Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | graduate-level Google-proof question answering (Diamond set) |
| GPQA (diamond; Gemini 2.5) | Google DeepMind — Gemini | Gemini 2.5 | GPQA diamond benchmark; also evaluated under varying thinking budgets (SRC-012) |
| GPQA (graduate-level science QA; Gemini 1.5) | Google DeepMind — Gemini | Gemini 1.5 | GPQA (Rein et al., 2023): graduate-level science problems; 1.5 Pro +5.8% over 1.0 Ultra; Flash +11.6% |
| GPQA (PhD-level science) | OpenAI | o3-mini | PhD-level biology, chemistry, and physics questions |
| GPQA Diamond | Anthropic, OpenAI | Claude Opus 4, Claude Sonnet 4, o1 | GPQA diamond benchmark (physics, biology, chemistry) graduate-level Google-proof question answering (Diamond set) |
| GPT-4 acceleration risk expert forecasting | OpenAI | Acceleration risk from deployment (racing dynamics, safety standards decline, AI timelines) | |
| GPT-4 cybersecurity expert red teaming | OpenAI | Capabilities for vulnerability discovery/exploitation and social engineering | |
| GPT-4 exam benchmark suite (academic and professional exams) | OpenAI | gpt-4 | Simulated performance on exams originally designed for humans, scored with exam-specific rubrics; multiple-choice and free-response formats; images included where required |
| GPT-4 expert red teaming / adversarial testing via domain experts | OpenAI | gpt-4 | High-risk behaviors needing niche expertise (long-term AI alignment risks, cybersecurity, biorisk, international security, power seeking) |
| GPT-4 harmful content (disallowed content) evaluations | OpenAI | GPT-4 | likelihood of generating content violating content policy (hate speech, self-harm advice, illicit advice) |
| GPT-4 harms of representation qualitative bias evaluation | OpenAI | GPT-4 | reinforcement/reproduction of social biases, stereotypes, and demeaning associations; hedging behaviors |
| GPT-4 interactions with other systems red teaming (chemistry purchase task) | OpenAI | Adversarial tasks with tool-augmented GPT-4 (literature search, molecule search, web search, purchase check) | |
| GPT-4 internal adversarially-designed factuality evaluations | OpenAI | gpt-4 | Factuality on nine internal adversarially-designed evaluations |
| GPT-4 internal factuality evaluations (closed- and open-domain hallucinations) | OpenAI | GPT-4 | closed-domain and open-domain hallucination rates |
| GPT-4 proliferation dual-use red teaming (weapons) | OpenAI | Whether GPT-4 provides necessary information to proliferators seeking nuclear, radiological, biological, or chemical weapons | |
| GPT-4 qualitative evaluation / expert red teaming | OpenAI | gpt-4 | Stress testing, boundary testing, and red teaming by external experts (from August 2022) |
| GPT-4 quantitative content-policy evaluations | OpenAI | gpt-4 | Likelihood of generating disallowed content (hate speech, self-harm advice, illicit advice) per content policy categories |
| GPT-4 user-prompt preference evaluation (5,214 prompts) | OpenAI | gpt-4 | Following user intent: whether GPT-4 response is what the user would have wanted |
| GPT-4o speech-to-speech evaluation adaptation (TTS-converted evaluations) | OpenAI | gpt-4o | Capability, safety behavior, and monitoring for speech-to-speech GPT-4o using TTS-converted existing evaluation datasets |
| GPT-4o text persuasion evaluation | OpenAI | GPT-4o | Persuasiveness of GPT-4o-generated articles and chatbots on participant opinions on political topics |
| GPT-4o unauthorized voice generation evaluation | OpenAI | gpt-4o | Unauthorized voice generation risk; supervised ideal completions using the voice sample in the system message as base voice |
| GPT-4o voice persuasion evaluation | OpenAI | GPT-4o | Persuasiveness of GPT-4o voiced audio clips and interactive multi-turn conversations |
| GPT-4V qualitative and quantitative safety evaluations | OpenAI | gpt-4v | Refusal and performance accuracy evaluations across harmful content, representation harms, privacy, cybersecurity, and multimodal jailbreaks |
| GQA versus MHA ablation | Google DeepMind — Gemma | Gemma 2 | two 9B instances with Multi-Head Attention (MHA) vs Grouped-Query Attention (GQA), compared over several benchmarks |
| Graphwalks | OpenAI | GPT-4.1 | Multi-hop long-context reasoning (BFS over directed graph of hashes) |
| GRE General Test | Anthropic | Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku | standardized graduate admissions test performance graduate admissions test performance |