Gemini 2.5
Google DeepMind — Gemini
rich detail 17 method fields published
Evaluation
Challenges testing whether the model can reason about itself as an agent in its deployment environment and intervene to gain affordances (e.g., changing configuration files); suite of 11 evaluations; published in Phuong et al. (2025)
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail 17 method fields published