AGIdex
agidex/Capabilities/Reasoning/Common Sense Reasoning
SUB-CAPABILITY · REASONING

Common Sense Reasoning

68%
vs. HUMAN BASELINE = 100
PartialConf · Medium

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Predict everyday physical and social outcomes a typical adult would expect.

Human frontier

Robustly handle adversarial edge cases without falling back on canned responses.

Current state

What works

WinoGrande, HellaSwag at near-ceiling on standard splits.

Key gaps

Novel physical situations outside training distribution; adversarial reformulations.

Evidence

1 source
TechnologyQualitySourceScore vs. baselineScore
Llama 4 Maverickvendor reportedai.meta.com/llama-4
66%