Skip to content

Evaluation

200K long-context recall evaluation

recalling information from all sections of a long document (up to 200K tokens)

Core capabilities Anthropic

How it was run

Open a release to see the setup, scoring, and other details that the lab published. Only fields the lab actually disclosed are shown.

Claude 2.1
Anthropic

rich detail 14 method fields published

See how this test was run