Comprehend everyday text across registers — news, conversation, instructions.
Resolve genuine ambiguity in expert-level legal, scientific, and literary texts.
Reading comprehension, summarization, question answering at near-human level.
Pragmatics and implied meaning in unfamiliar cultural contexts.
| Technology | Quality | Source | Score vs. baseline | Score |
|---|---|---|---|---|
| Claude Opus 4.7 | vendor reported | anthropic.com/claude/opus-4-7 | 97% | |
| Claude Sonnet 4.6 | vendor reported | anthropic.com/claude/sonnet-4-6 | 94% | |
| Llama 4 Maverick | vendor reported | ai.meta.com/llama-4 | 92% | |
| Llama 4 Scout | vendor reported | ai.meta.com/llama-4 | 90% | |
| Mistral Large 3 | vendor reported | mistral.ai/news/mistral-large-3 | 88% |