AGIdex
agidex/Capabilities/Reasoning/Multi-step Planning
SUB-CAPABILITY · REASONING

Multi-step Planning

50%
vs. HUMAN BASELINE = 100
PartialConf · Medium

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Plan a coherent sequence of actions toward a stated goal in a familiar domain.

Human frontier

Long-horizon plans with backtracking and replanning when conditions change mid-task.

Current state

What works

Short task decomposition; tool-use sequences in controlled environments.

Key gaps

Horizons beyond ~30 steps. Replanning when premises change mid-task.

Evidence

3 sources
TechnologyQualitySourceScore vs. baselineScore
Claude Opus 4.7independentvellum.ai/blog/claude-opus-4-7-benchmarks-explained
57%
GPT-5.4independentpaperswithcode.com/sota/autonomous-agents-on-gaia
56%
DeepSeek V4 Provendor reporteddeepseek.com/v4
58%
Source: paperswithcode.com/sota/autonomous-agents-on-gaia · Last reviewed May 2026