AGIdex
agidex/Capabilities/Communication/Language Generation
SUB-CAPABILITY · COMMUNICATION

Language Generation

100%
vs. HUMAN BASELINE = 100
SolvedConf · High

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Produce fluent, on-topic text indistinguishable from a competent writer.

Human frontier

Sustain authorial voice and coherent argument across book-length work.

Current state

What works

Most short- and medium-form writing — emails, essays, code documentation.

Key gaps

Long-form structural coherence; losing thread across tens of thousands of tokens.

Evidence

3 sources
TechnologyQualitySourceScore vs. baselineScore
Claude Opus 4.7independentlmarena.ai
102%
Claude Sonnet 4.6independentlmarena.ai
98%
GPT-5.5independentlmarena.ai
104%
Source: lmsys.org/blog/2023-05-03-arena · Last reviewed May 2026