Skip to content

Evaluation

Adversarial benchmarks (Adversarial SQuAD, Dynabench SQuAD, GSM-Plus, PAWS)

Adversarial vs non-adversarial performance in three areas: QA (Adversarial SQuAD, Dynabench SQuAD vs SQuAD), mathematical reasoning (GSM-Plus vs GSM8K), paraphrase detection (PAWS vs QQP)

Core capabilities Meta

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 3
Meta

partial detail 10 method fields published

See how this test was run