AGIdex
agidex/Capabilities/Learning/Continual Learning
SUB-CAPABILITY · LEARNING

Continual Learning

22%
vs. HUMAN BASELINE = 100
UnsolvedConf · Medium

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Add a new skill over time without meaningfully degrading older ones — like a human apprentice.

Human frontier

Acquire skills over months on the job without replay or retraining on prior data.

Current state

What works

Replay-based methods on narrow benchmarks show modest forgetting reduction.

Key gaps

Catastrophic forgetting at scale. No deployed model learns post-training in production.

Evidence

0 sources
TechnologyQualitySourceScore vs. baselineScore
No evidence records yet.
Source: paperswithcode.com/task/continual-learning · Last reviewed May 2026