AGIdex
agidex/Methodology
How agidex measures progress

Methodology

Every scoring decision documented in public. This page is what separates agidex from opinion sites.

agidex tracks how close AI is to what humans can do. It does so at the capability level — not as a single countdown clock, and not as one person's opinion. Each capability is decomposed into sub-capabilities, and that is where the real work happens: each sub-capability has a score, a status, a confidence level, and a source.

1 · Anchored to human baseline

The scoring scale is anchored to human baseline = 100. A score of 100 means AI matches what most adults can do on that sub-capability. A score above 100 means AI has surpassed typical human ability — which is already true for several sub-capabilities like multilingual translation and document OCR. A score of 50 means AI is roughly halfway from "no meaningful capability" to "matches a typical adult".

Each sub-capability also has an optional human frontier — what the best human specialists have proven. Frontier is reserved as a separate axis; it is not the same as baseline. Many sub-capabilities will have AI at or above baseline long before they reach frontier.

2 · How scores flow

Evidence flows upward. Benchmarks set sub-capability scores; primary scores derive from sub-capabilities; track scores derive from primaries. We never hand-set anything above the sub-capability layer.

Benchmark evidence
arxiv · papers-with-code · leaderboards
Sub-capability scores
set manually · sourced · 0–120+
Primary capability scores
derived · mean of sub-capabilities
Track scores
mean of primaries · Cognitive · Embodied
Overall AGI index
mean of all capabilities

3 · Scoring rubric

Every score on agidex maps to one of six interpretation bands. The rubric is the universal anchor — without it, percentages feel arbitrary. With it, you should be able to read any number on the site and know exactly what it means.

RangeInterpretation
0–20Early research
20–40Narrow benchmark competence
40–60Strong benchmark · reliability gaps
60–80Human-competitive in common scenarios
80–100Comparable to typical skilled adult
100+Reliably exceeds typical human baseline

Bands above 60 require benchmark evidence. Bands above 100 mean AI reliably exceeds the typical human baseline on that sub-capability — already true for several Vision and Hearing sub-capabilities, and approaching for multilingual generation.

4 · Status — solved · partial · unsolved

  • Solved  AI is at or above typical human baseline.
  • Partial  Meaningful progress, but significant gaps remain.
  • Unsolved  Early or no meaningful progress.

A sub-capability can have score above 60 and still be marked Partial — the threshold is judgment about whether the gaps are significant, not arithmetic.

5 · Confidence levels

Confidence reflects how well-defined the human baseline is and how reliable the benchmark evidence is. It is shown publicly so visitors understand the strength of each score.

  • Conf · High  Baseline is objectively measurable; strong benchmark evidence exists.
  • Conf · Medium  Baseline is reasonably defined; some benchmark evidence exists.
  • Conf · Low  Baseline is inherently subjective; benchmark evidence is limited or contested.

Physical capabilities generally score High confidence. Cognitive and social capabilities generally score Medium or Low. This is not a weakness in the methodology — it is an honest reflection of what is measurable.

6 · Evidence quality

Not all benchmark evidence is equal. Every benchmark record carries a quality tag so visitors can tell vendor self-reports apart from independent results.

independentPublic benchmark or third-party leaderboard.
vendor reportedNumbers self-published by the model maker.
replicatedReproduced by an independent group.
lab demoDemonstrated in controlled lab conditions.
deploymentObserved in production use.
anecdotalSingle report; not systematically measured.

Rule: a sub-capability score should not exceed 60 without at least one benchmark record attached. Above 60, the source URL must be verifiable.

7 · Scope rules

  • Technologies tracked: general-purpose AI models (e.g. Claude, GPT, Gemini, Llama) and general-purpose physical systems (e.g. Atlas, Optimus, Figure).
  • Excluded: industry-specific tools at v1 — no Harvey AI, no Bloomberg GPT, no IBM Watson Health.
  • Excluded: standalone hardware (GPUs, chips). Hardware that bottlenecks a specific capability is noted on the sub-cap page, not tracked as a technology.
  • Out of scope at v1: jobs, tasks, impact reports.

8 · Editorial workflow

At launch, scores are updated manually with a ~15–20 minute daily review cadence. Every change must carry a source_url. The data model is built from day one to accept agent-proposed updates as draft entries that require human approval before going live.

9 · Principles

  • Monochrome by default. Status and confidence are communicated by type density and pattern, never by color alone.
  • Downward pressure on scores. When evidence is ambiguous, the score stays low until clearer evidence arrives.
  • No single number. The overall AGI index is a convenience — the real information lives at the sub-capability level.
  • Transparent about uncertainty. Confidence levels and evidence quality tags exist so visitors can calibrate their trust.