About
How this research was built
Inside the Frontier looks at how major AI labs train models, evaluate capabilities and risks, and use evaluation results in release decisions.
Source set
40 public publications from OpenAI, Anthropic, Google Deepmind (Gemini, Gemma), and Meta
The research uses system cards, model cards, technical reports, safety reports, evaluation reports, governance frameworks, and release publications.
The documents were divided into 2,111 passages while keeping their source information attached. The analysis compares evidence by model release so changes can be traced over time.
Research
The analysis focuses on training, evaluation, and governance
Training
I tracked documented training methods across model releases and grouped similar methods so they could be compared over time.
Evaluation
I tracked what models were evaluated on and, where the sources provided enough detail, how those evaluations were run.
Governance
I reviewed how labs connect evaluation results to safety assessments, safeguards, deployment conditions, and release decisions.
Ask the Atlas
Ask questions across the same source set
Ask the Atlas searches the research corpus for relevant passages, ranks the best matches, and uses them to answer the question. Each answer shows the publications used.
Ask a questionLimits
What this research does not cover
The Atlas reflects what is documented in the publications I reviewed. Missing information does not mean that a lab did not use a method, run an evaluation, or apply a safeguard.
Some numeric evaluation tables were not included in the corpus, and the source set ends in August 2025. The project is a selected research corpus, not a complete archive of every publication from these labs.