AGIdex
agidex/Capabilities/Communication/Language Understanding
SUB-CAPABILITY · COMMUNICATION

Language Understanding

96%
vs. HUMAN BASELINE = 100
SolvedConf · High

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Comprehend everyday text across registers — news, conversation, instructions.

Human frontier

Resolve genuine ambiguity in expert-level legal, scientific, and literary texts.

Current state

What works

Reading comprehension, summarization, question answering at near-human level.

Key gaps

Pragmatics and implied meaning in unfamiliar cultural contexts.

Evidence

5 sources
TechnologyQualitySourceScore vs. baselineScore
Claude Opus 4.7vendor reportedanthropic.com/claude/opus-4-7
97%
Claude Sonnet 4.6vendor reportedanthropic.com/claude/sonnet-4-6
94%
Llama 4 Maverickvendor reportedai.meta.com/llama-4
92%
Llama 4 Scoutvendor reportedai.meta.com/llama-4
90%
Mistral Large 3vendor reportedmistral.ai/news/mistral-large-3
88%
Source: super.gluebenchmark.com · Last reviewed May 2026