AGIdex
agidex/Capabilities/Decision Making/Judgment Under Uncertainty
SUB-CAPABILITY · DECISION MAKING

Judgment Under Uncertainty

46%
vs. HUMAN BASELINE = 100
PartialConf · Medium

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Reason about probabilities and update on new evidence in familiar situations.

Human frontier

Calibrated forecasts across long horizons with explicit uncertainty quantification.

Current state

What works

Short-horizon Bayesian updates; structured forecasting on known question types.

Key gaps

Calibration drift on rare events. Overconfidence on out-of-distribution scenarios.

Evidence

0 sources
TechnologyQualitySourceScore vs. baselineScore
No evidence records yet.
Source: metaculus.com/questions · Last reviewed May 2026