AGIdex
SUB-CAPABILITY · EMBODIED COGNITION

Task Planning

36%
vs. HUMAN BASELINE = 100
UnsolvedConf · Low

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Follow a multi-step verbal instruction to complete a household task.

Human frontier

Plan and execute long-horizon tasks autonomously in novel environments without human correction.

Current state

What works

Short-horizon tasks in structured settings; SayCan-style grounded planning with LLMs.

Key gaps

Long-horizon tasks beyond ~10 steps; recovery when early steps fail or environment changes.

Evidence

1 source
TechnologyQualitySourceScore vs. baselineScore
Figure 03vendor reportedfigure.ai/news
38%
Source: paperswithcode.com/task/task-planning · Last reviewed May 2026