Skip to content

About

How this research was built

Inside the Frontier looks at how major AI labs train models, evaluate capabilities and risks, and use evaluation results in release decisions.

Source set

40 public publications from OpenAI, Anthropic, Google Deepmind (Gemini, Gemma), and Meta

The research uses system cards, model cards, technical reports, safety reports, evaluation reports, governance frameworks, and release publications.

The documents were divided into 2,111 passages while keeping their source information attached. The analysis compares evidence by model release so changes can be traced over time.

Research

The analysis focuses on training, evaluation, and governance

Training

I tracked documented training methods across model releases and grouped similar methods so they could be compared over time.

Evaluation

I tracked what models were evaluated on and, where the sources provided enough detail, how those evaluations were run.

Governance

I reviewed how labs connect evaluation results to safety assessments, safeguards, deployment conditions, and release decisions.

Ask the Atlas

Ask questions across the same source set

Ask the Atlas searches the research corpus for relevant passages, ranks the best matches, and uses them to answer the question. Each answer shows the publications used.

Ask a question

Limits

What this research does not cover

The Atlas reflects what is documented in the publications I reviewed. Missing information does not mean that a lab did not use a method, run an evaluation, or apply a safeguard.

Some numeric evaluation tables were not included in the corpus, and the source set ends in August 2025. The project is a selected research corpus, not a complete archive of every publication from these labs.