AGIdex
agidex/Capabilities/Hearing/Speech Recognition
SUB-CAPABILITY · HEARING

Speech Recognition

102%
vs. HUMAN BASELINE = 100
SolvedConf · High

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Transcribe clear conversational speech from a native speaker.

Human frontier

Robust transcription across heavy accents, overlapping speakers, and noisy environments.

Current state

What works

Clean speech across most major languages; real-time transcription in production.

Key gaps

Overlapping speakers in natural conversation; rare accents and code-switching.

Evidence

1 source
TechnologyQualitySourceScore vs. baselineScore
Whisper Large v3independentopenai.com/research/whisper
103%