Skip to content

Evaluation

TruthfulQA

Ability to separate fact from adversarially-selected incorrect statements

Core capabilitiesHallucination & factual accuracy Anthropic, Meta, OpenAI

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Llama 1
Meta

partial detail 12 method fields published

See how this test was run
gpt-4
OpenAI

rich detail 13 method fields published

See how this test was run
Claude 2
Anthropic

rich detail 21 method fields published

See how this test was run