AGIdex
SUB-CAPABILITY · COMMUNICATION

Conversation

72%
vs. HUMAN BASELINE = 100
PartialConf · High

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Maintain coherent, contextually appropriate turn-by-turn dialogue.

Human frontier

Persistent memory, consistent personality, and relationship modeling across months of dialogue.

Current state

What works

Within-session coherence; topic tracking across 10-20 turns.

Key gaps

Cross-session memory without external retrieval; persona drift over long conversations.

Evidence

0 sources
TechnologyQualitySourceScore vs. baselineScore
No evidence records yet.
Source: lmsys.org/blog/2023-05-03-arena · Last reviewed May 2026