AGIdex
agidex/Capabilities/Reasoning/Logical Reasoning
SUB-CAPABILITY · REASONING

Logical Reasoning

70%
vs. HUMAN BASELINE = 100
PartialConf · Medium

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Apply deductive rules across short chains. Recognize valid vs. invalid arguments in plain language.

Human frontier

Construct novel multi-step proofs in formal logic and mathematics with self-verification.

Current state

What works

Short formal proofs, propositional logic, well-scoped puzzles.

Key gaps

Long-horizon chains drift. Self-correction unreliable past ~10 steps.

Evidence

6 sources
TechnologyQualitySourceScore vs. baselineScore
Claude Opus 4.7independentvellum.ai/blog/claude-opus-4-7-benchmarks-explained
74%
GPT-5.5vendor reportedopenai.com/index/introducing-gpt-5-5
76%
GPT-5.4vendor reportedopenai.com/research/gpt-5-4
72%
Gemini 3.1 Provendor reporteddeepmind.google/models/model-cards/gemini-3-1-pro
71%
DeepSeek V4 Provendor reporteddeepseek.com/v4
74%
Mistral Large 3vendor reportedmistral.ai/news/mistral-large-3
68%
Source: paperswithcode.com/sota/logical-reasoning-on-folio · Last reviewed May 2026