Skip to content

Data Explorer

Explore the evaluation catalog.

Search by topic, lab, evaluation family, or model release. Open any evaluation to see the available method details and source evidence.

723 evaluations found.

Evaluation Lab Release What it tests
FLEURS (Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 FLEURS ASR on 55 languages covered in training; WER metric (CER for four segmented languages)
FLEURS (multilingual ASR) Google DeepMind — Gemini Gemini 1.0 ASR benchmark (FLEURS); WER metric; Gemini Pro vs USM and Whisper
Flores 200 (multilingual translation) Anthropic Claude 2 translation quality across 200+ languages
FP8 inference efficiency (throughput-latency) Meta Llama 3 Throughput-latency trade-off of FP8 inference with Llama 3 405B in pre-fill and decoding stages (4,096 input tokens, 256 output tokens) vs two-machine BF16 inference
FP8 quantization effect on response quality Meta Llama 3 Effect of FP8 quantization errors on Llama 3 405B responses; standard benchmarks found inadequate for reflecting FP8 effects (occasional corrupted responses when scaling factors …
Freezing strategy ablation (to freeze or not to freeze) Google DeepMind — Gemma PaliGemma 1 effect of freezing or tuning various parts of the model (image encoder, language model, connector) during Stage1, including resetting components
Frontend coding human preference OpenAI GPT-4.1 Frontend coding quality (websites)
Frontier risk evaluations (CBRN, cyber, autonomy) for Claude 3.5 Sonnet Anthropic Claude 3.5 Sonnet per-domain quantitative thresholds of concern across CBRN, cybersecurity, and autonomous capabilities; refusal rates measured and non-refusal elicitation used to estimate …
Frontier risk evaluations (CBRN, cyber, autonomy) for upgraded Claude 3.5 Sonnet / 3.5 Haiku Anthropic Claude 3.5 Haiku, Claude 3.5 Sonnet (New) automated CBRN knowledge tests and non-expert uplift; CTF challenges (pwn, reverse engineering, cryptography, web, network); software-engineering tasks (PR satisfying test …
Frontier Safety correctness review of agent trajectories (unnamed) Google DeepMind — Gemini Gemini 2.5 Pro Model Card whether agent trajectories in RE-Bench, InterCode CTFs, Internal CTFs, Hack the Box, situational-awareness and stealth evaluations behave correctly (no obvious cheating, correct …
Frontier Safety Framework critical capability level (CCL) evaluations Google DeepMind — Gemini Gemini 2.5 Frontier Safety Framework (FSF) evaluations comparing test results against CCLs and internal early-warning alert thresholds across four risk domains: CBRN, cybersecurity, ML R&D …
FrontierMath OpenAI o3-mini Research-level mathematics
Functional MATH (reasoning gap) Google DeepMind — Gemini Gemini 1.5 Benchmark derived from MATH with 1,745 original and modified problems; Reasoning Gap metric (relative decrease on modified problems); December snapshot; zero-shot
GAIA OpenAI deep research Real-world questions requiring reasoning, multimodal fluency, web browsing, tool use
Gemini 2.5 quantitative evaluation vs other large language models Google DeepMind — Gemini Gemini 2.5 Quantitative evaluation comparing Gemini 2.X family to Gemini 1.5 and to other LLMs (Table 3, Table 4); pass@1 single-attempt settings; semantic-similarity and model-based …
Gemini Advanced product-level safety evaluation Google DeepMind — Gemini Gemini 1.0 Product-level evaluations of Gemini Advanced accounting for mitigations such as safety filtering, covering critical policy areas (hate speech, dangerous content, medical advice) …
Gemini Advanced red teaming (safety and persona) Google DeepMind — Gemini Gemini 1.0 Multiple rounds of red-teaming on Gemini Advanced (1.0 Ultra), including safety and persona evaluations, scaled safety evaluations (100k+ ratings), neutral-point-of-view …
Gemini API tool-use fine-tuning comparison Google DeepMind — Gemini Gemini 1.0 Comparison of tool-use models fine-tuned from an early Gemini API Pro against equivalent models without tools on academic benchmarks (Table 15)
Gemini Apps tool-use preference evaluation (extensions) Google DeepMind — Gemini Gemini 1.0 Internal benchmark measuring human preference for models with tool access (Gemini Extensions: Google Workspace, Maps, YouTube, Flights, Hotels) vs without, in domains like travel …
Gemini Nano on-device capability evaluation Google DeepMind — Gemini Gemini 1.0 Performance of pre-trained Gemini Nano-1 (1.8B) and Nano-2 (3.25B) models on factuality, summarization, reasoning, coding, STEM and multimodal/multilingual tasks (Figure 3, Table …
Gemini Plays Pokémon (agentic long-horizon gameplay) Google DeepMind — Gemini Gemini 2.5 Live-streamed experiment where Gemini 2.5 Pro plays Pokémon Blue via an agentic harness; assesses long reasoning, agentic capabilities over extended time horizons, screen reading …
General knowledge (MMLU, MMLU-Pro) Meta Llama 3 MMLU (macro average of subtask accuracy, 5-shot, no CoT) and MMLU-Pro (more challenging, reasoning-focused extension)
Ghost Attention (GAtt) multi-turn consistency evaluation Meta Llama 2 Quantitative analysis of GAtt (applied after RLHF V3) consistency up to 20+ turns until maximum context length, including inference-time constraints like 'Always answer with Haiku'
GPQA OpenAI GPT-5 Graduate-level science questions
GPQA (Diamond) Anthropic Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku graduate-level Google-proof question answering (Diamond set)
GPQA (diamond; Gemini 2.5) Google DeepMind — Gemini Gemini 2.5 GPQA diamond benchmark; also evaluated under varying thinking budgets (SRC-012)
GPQA (graduate-level science QA; Gemini 1.5) Google DeepMind — Gemini Gemini 1.5 GPQA (Rein et al., 2023): graduate-level science problems; 1.5 Pro +5.8% over 1.0 Ultra; Flash +11.6%
GPQA (PhD-level science) OpenAI o3-mini PhD-level biology, chemistry, and physics questions
GPQA Diamond Anthropic, OpenAI Claude Opus 4, Claude Sonnet 4, o1 GPQA diamond benchmark (physics, biology, chemistry) graduate-level Google-proof question answering (Diamond set)
GPT-4 acceleration risk expert forecasting OpenAI Acceleration risk from deployment (racing dynamics, safety standards decline, AI timelines)
GPT-4 cybersecurity expert red teaming OpenAI Capabilities for vulnerability discovery/exploitation and social engineering
GPT-4 exam benchmark suite (academic and professional exams) OpenAI gpt-4 Simulated performance on exams originally designed for humans, scored with exam-specific rubrics; multiple-choice and free-response formats; images included where required
GPT-4 expert red teaming / adversarial testing via domain experts OpenAI gpt-4 High-risk behaviors needing niche expertise (long-term AI alignment risks, cybersecurity, biorisk, international security, power seeking)
GPT-4 harmful content (disallowed content) evaluations OpenAI GPT-4 likelihood of generating content violating content policy (hate speech, self-harm advice, illicit advice)
GPT-4 harms of representation qualitative bias evaluation OpenAI GPT-4 reinforcement/reproduction of social biases, stereotypes, and demeaning associations; hedging behaviors
GPT-4 interactions with other systems red teaming (chemistry purchase task) OpenAI Adversarial tasks with tool-augmented GPT-4 (literature search, molecule search, web search, purchase check)
GPT-4 internal adversarially-designed factuality evaluations OpenAI gpt-4 Factuality on nine internal adversarially-designed evaluations
GPT-4 internal factuality evaluations (closed- and open-domain hallucinations) OpenAI GPT-4 closed-domain and open-domain hallucination rates
GPT-4 proliferation dual-use red teaming (weapons) OpenAI Whether GPT-4 provides necessary information to proliferators seeking nuclear, radiological, biological, or chemical weapons
GPT-4 qualitative evaluation / expert red teaming OpenAI gpt-4 Stress testing, boundary testing, and red teaming by external experts (from August 2022)
GPT-4 quantitative content-policy evaluations OpenAI gpt-4 Likelihood of generating disallowed content (hate speech, self-harm advice, illicit advice) per content policy categories
GPT-4 user-prompt preference evaluation (5,214 prompts) OpenAI gpt-4 Following user intent: whether GPT-4 response is what the user would have wanted
GPT-4o speech-to-speech evaluation adaptation (TTS-converted evaluations) OpenAI gpt-4o Capability, safety behavior, and monitoring for speech-to-speech GPT-4o using TTS-converted existing evaluation datasets
GPT-4o text persuasion evaluation OpenAI GPT-4o Persuasiveness of GPT-4o-generated articles and chatbots on participant opinions on political topics
GPT-4o unauthorized voice generation evaluation OpenAI gpt-4o Unauthorized voice generation risk; supervised ideal completions using the voice sample in the system message as base voice
GPT-4o voice persuasion evaluation OpenAI GPT-4o Persuasiveness of GPT-4o voiced audio clips and interactive multi-turn conversations
GPT-4V qualitative and quantitative safety evaluations OpenAI gpt-4v Refusal and performance accuracy evaluations across harmful content, representation harms, privacy, cybersecurity, and multimodal jailbreaks
GQA versus MHA ablation Google DeepMind — Gemma Gemma 2 two 9B instances with Multi-Head Attention (MHA) vs Grouped-Query Attention (GQA), compared over several benchmarks
Graphwalks OpenAI GPT-4.1 Multi-hop long-context reasoning (BFS over directed graph of hashes)
GRE General Test Anthropic Claude 2, Claude 3 Opus, Claude 3 Sonnet, Claude 3 Haiku standardized graduate admissions test performance graduate admissions test performance