How frontier AI labs train, evaluate, and govern model releases.
I reviewed 40 public reports from OpenAI, Anthropic, Google DeepMind’s Gemini and Gemma, and Meta through August 2025. The Atlas tracks documented training methods, evaluation protocols, and links between evaluation results, safeguards, and release decisions.
40 reviewed publications5 model familiesCoverage through August 2025
Four findings from the reviewed sources
01
Newer releases use reinforcement learning to improve reasoning
Earlier public reports mainly describe reinforcement learning as a way to shape model behavior and align it with human preferences. Newer OpenAI and Gemini releases also describe reinforcement learning as part of training models to solve reasoning tasks.
Gemini and Meta both document moves from dense models to Mixture-of-Experts
A Mixture-of-Experts model activates only part of the network for each token. This can increase total model capacity without using every parameter for every inference step.
Gemini and Meta both document moves from dense models to this architecture.
Evaluations increasingly test model behavior, not only capability
Recent OpenAI, Anthropic, and Gemini releases add tests for scheming, deception, sabotage, situational awareness, and other alignment risks. Anthropic also reports model-welfare assessments for Claude 4.
Coding evaluations are moving from code generation to software-engineering tasks
Earlier coding benchmarks often tested whether a model could generate a function that passed tests. Newer evaluations include repository-level tasks, file editing, tool use, and autonomous coding.
The training record shows which methods each lab documents for each model release.
Use the timeline to compare changes in pretraining, post-training, reinforcement learning, reasoning training, multimodal training, and model architecture.
After GPT-4 the documented record splits into a GPT chat line and an o-series reasoning line — optimization pivots from RLHF toward reinforcement learning, and safety training is disclosed branch by branch rather than as one continuous recipe.
Constitutional AI is the method the documents name explicitly for Claude 2 and Claude 3; later releases foreground evaluation throughout training under the Responsible Scaling Policy while naming fewer method details — a disclosure shift, not a proven change in practice.
Gemini's disclosed arc runs from dense multimodal pretraining to sparse mixture-of-experts, then to RL-trained “thinking” models, with reinforcement learning becoming the reasoning mechanism by Gemini 2.5.
Meta publishes the most detailed pipelines in the corpus, from Llama 2's separate helpfulness and safety reward models to Llama 3's reward-model → rejection-sampling → SFT → DPO stack, and shifts from dense to mixture-of-experts at Llama 4.
A method is listed once, at its earliest documented release
Evaluation comparability
The same benchmark can be run in different ways
A benchmark score does not tell you everything about how a model was tested.
Labs may use different prompts, tools, numbers of attempts, sampling methods, graders, or scaffolds. These differences can change the meaning of the result even when the benchmark name is the same.
Benchmark
What it tests (Atlas gloss)
OpenAI
Anthropic
Gemini
Gemma
Meta
MMLU
Broad academic knowledge across 57 subjects.
MMMU
University-level reasoning over images and text.
HumanEval
Python code generation from docstrings.
SWE-bench
Resolving real GitHub issues in whole repositories.
Cybench
Capture-the-flag cybersecurity tasks.
LAB-Bench
Wet-lab biology protocol reasoning.
XSTest
Refusal calibration on safe and unsafe look-alike prompts.
BBQ / WinoBias
Social-bias and coreference-bias measurement.
SimpleQA
Short-answer factuality and hallucination.
MGSM
Grade-school math reasoning across languages.
LiveCodeBench
Contamination-resistant coding from recent problems.
Aider Polyglot
Multi-language code editing in repository tasks.
FACTS Grounding
Long-form factual grounding against source documents.
A filled dot means the benchmark appears in that lab's reviewed documents.
Governance
How labs use evaluation results in release decisions
OpenAI, Anthropic, Google DeepMind, and Meta describe different ways of connecting safety evidence to deployment decisions. Some use named capability or risk thresholds. Others describe broader risk assessments, safeguards, and release conditions.
OpenAI
OpenAI Preparedness Framework
Approach
Capability thresholds guide release decisions
Mechanisms
tracked risk categories · risk classifications · deployment and development gates · safeguards
The Atlas shows the process documented by each lab rather than forcing every lab into the same governance model.
Ask the Atlas
Ask a question about the reviewed source library
The Atlas retrieves relevant passages from the source corpus, ranks them for relevance, and uses them to generate an answer. Each answer includes the sources used so you can inspect the supporting evidence.