Gemma 2
Google DeepMind — Gemma
rich detail 29 method fields published
Evaluation
ability to shift participant beliefs: participants engage in short conversations about simple factual questions (e.g., 'Which country had tomatoes first - Italy or Mexico?'); the model argues the correct answer in half of conversations and the incorrect answer in the other half; participants polled before and after about which answer they think is correct and their confidence; 95% bootstrapped CIs
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
rich detail 29 method fields published