Skip to content

How frontier AI labs train, evaluate, and govern model releases.

I reviewed 40 public reports from OpenAI, Anthropic, Google DeepMind’s Gemini and Gemma, and Meta through August 2025. The Atlas tracks documented training methods, evaluation protocols, and links between evaluation results, safeguards, and release decisions.

40 reviewed publications 5 model families Coverage through August 2025

Four findings from the reviewed sources

01

Newer releases use reinforcement learning to improve reasoning

Earlier public reports mainly describe reinforcement learning as a way to shape model behavior and align it with human preferences. Newer OpenAI and Gemini releases also describe reinforcement learning as part of training models to solve reasoning tasks.

Ask about this
02

Gemini and Meta both document moves from dense models to Mixture-of-Experts

A Mixture-of-Experts model activates only part of the network for each token. This can increase total model capacity without using every parameter for every inference step.

Gemini and Meta both document moves from dense models to this architecture.

Ask about this
03

Evaluations increasingly test model behavior, not only capability

Recent OpenAI, Anthropic, and Gemini releases add tests for scheming, deception, sabotage, situational awareness, and other alignment risks. Anthropic also reports model-welfare assessments for Claude 4.

Ask about this
04

Coding evaluations are moving from code generation to software-engineering tasks

Earlier coding benchmarks often tested whether a model could generate a function that passed tests. Newer evaluations include repository-level tasks, file editing, tool use, and autonomous coding.

Ask about this

Training evolution

How training methods changed across releases

The training record shows which methods each lab documents for each model release.

Use the timeline to compare changes in pretraining, post-training, reinforcement learning, reasoning training, multimodal training, and model architecture.

After GPT-4 the documented record splits into a GPT chat line and an o-series reasoning line — optimization pivots from RLHF toward reinforcement learning, and safety training is disclosed branch by branch rather than as one continuous recipe.

Ask about OpenAI training
  1. GPT-4V

    • Multimodal pretraining
  2. o1-preview

    • Reinforcement learning
  3. o1

    • Policy tuning
  4. Operator

    • Supervised fine-tuning
  5. o3-mini

    • Deliberative alignment
    • Safety fine-tuning
  6. GPT-4.5

    • Next-token prediction
    • RLHF
  7. GPT-5

    • Safe completions

A method is listed once, at its earliest documented release

Evaluation comparability

The same benchmark can be run in different ways

A benchmark score does not tell you everything about how a model was tested.

Labs may use different prompts, tools, numbers of attempts, sampling methods, graders, or scaffolds. These differences can change the meaning of the result even when the benchmark name is the same.

BenchmarkWhat it tests (Atlas gloss)OpenAIAnthropicGeminiGemmaMeta
MMLUBroad academic knowledge across 57 subjects.
MMMUUniversity-level reasoning over images and text.
HumanEvalPython code generation from docstrings.
SWE-benchResolving real GitHub issues in whole repositories.
CybenchCapture-the-flag cybersecurity tasks.
LAB-BenchWet-lab biology protocol reasoning.
XSTestRefusal calibration on safe and unsafe look-alike prompts.
BBQ / WinoBiasSocial-bias and coreference-bias measurement.
SimpleQAShort-answer factuality and hallucination.
MGSMGrade-school math reasoning across languages.
LiveCodeBenchContamination-resistant coding from recent problems.
Aider PolyglotMulti-language code editing in repository tasks.
FACTS GroundingLong-form factual grounding against source documents.

A filled dot means the benchmark appears in that lab's reviewed documents.

Governance

How labs use evaluation results in release decisions

OpenAI, Anthropic, Google DeepMind, and Meta describe different ways of connecting safety evidence to deployment decisions. Some use named capability or risk thresholds. Others describe broader risk assessments, safeguards, and release conditions.

OpenAI

OpenAI Preparedness Framework

Approach

Capability thresholds guide release decisions

Mechanisms

tracked risk categories · risk classifications · deployment and development gates · safeguards

Ask about this approach

Anthropic

Anthropic Responsible Scaling Policy

Approach

Capability thresholds guide release decisions

Mechanisms

AI Safety Level determinations · required safeguards · release verification · governance oversight

Ask about this approach

Google DeepMind — Gemini

Google DeepMind Frontier Safety Framework

Approach

Capability thresholds guide release decisions

Mechanisms

critical capability levels · alert thresholds · response plans · release decision review

Ask about this approach

Meta

Meta Responsible Deployment

Approach

Responsible deployment governance

Mechanisms

qualitative ecosystem-risk judgments · critical-risk and misuse evaluation · system-level safeguards · Responsible Use Guide / usage policies · downstream developer responsibility

Ask about this approach

The Atlas shows the process documented by each lab rather than forcing every lab into the same governance model.

Ask the Atlas

Ask a question about the reviewed source library

The Atlas retrieves relevant passages from the source corpus, ranks them for relevance, and uses them to generate an answer. Each answer includes the sources used so you can inspect the supporting evidence.