Skip to content

Evaluation

Tool use human evaluations (code execution, plot generation, file upload)

Human preference evaluation of tool use capabilities focused on code execution tasks, using 2,000 user prompts (code execution without plotting/file uploads, plot generation, file uploads) from LMSys, GAIA, human annotators and synthetic generation

Agentic & computer use Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 3
Meta

partial detail 13 method fields published

See how this test was run