Llama 3
Meta
partial detail 13 method fields published
Evaluation
Human preference evaluation of tool use capabilities focused on code execution tasks, using 2,000 user prompts (code execution without plotting/file uploads, plot generation, file uploads) from LMSys, GAIA, human annotators and synthetic generation
How it was run
Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.
partial detail 13 method fields published