Skip to content

Evaluation

Faithfulness to long documents evaluation (Claude 2.1)

whether the model answers questions correctly by referencing a provided document, and whether it mistakenly concludes a document supports a claim

Hallucination & factual accuracy Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 2.1
Anthropic

partial detail 15 method fields published

See how this test was run