Every scoring decision documented in public. This page is what separates agidex from opinion sites.
agidex tracks how close AI is to what humans can do. It does so at the capability level — not as a single countdown clock, and not as one person's opinion. Each capability is decomposed into sub-capabilities, and that is where the real work happens: each sub-capability has a score, a status, a confidence level, and a source.
The scoring scale is anchored to human baseline = 100. A score of 100 means AI matches what most adults can do on that sub-capability. A score above 100 means AI has surpassed typical human ability — which is already true for several sub-capabilities like multilingual translation and document OCR. A score of 50 means AI is roughly halfway from "no meaningful capability" to "matches a typical adult".
Each sub-capability also has an optional human frontier — what the best human specialists have proven. Frontier is reserved as a separate axis; it is not the same as baseline. Many sub-capabilities will have AI at or above baseline long before they reach frontier.
Evidence flows upward. Benchmarks set sub-capability scores; primary scores derive from sub-capabilities; track scores derive from primaries. We never hand-set anything above the sub-capability layer.
Every score on agidex maps to one of six interpretation bands. The rubric is the universal anchor — without it, percentages feel arbitrary. With it, you should be able to read any number on the site and know exactly what it means.
| Range | Interpretation |
|---|---|
| 0–20 | Early research |
| 20–40 | Narrow benchmark competence |
| 40–60 | Strong benchmark · reliability gaps |
| 60–80 | Human-competitive in common scenarios |
| 80–100 | Comparable to typical skilled adult |
| 100+ | Reliably exceeds typical human baseline |
Bands above 60 require benchmark evidence. Bands above 100 mean AI reliably exceeds the typical human baseline on that sub-capability — already true for several Vision and Hearing sub-capabilities, and approaching for multilingual generation.
A sub-capability can have score above 60 and still be marked Partial — the threshold is judgment about whether the gaps are significant, not arithmetic.
Confidence reflects how well-defined the human baseline is and how reliable the benchmark evidence is. It is shown publicly so visitors understand the strength of each score.
Physical capabilities generally score High confidence. Cognitive and social capabilities generally score Medium or Low. This is not a weakness in the methodology — it is an honest reflection of what is measurable.
Not all benchmark evidence is equal. Every benchmark record carries a quality tag so visitors can tell vendor self-reports apart from independent results.
Rule: a sub-capability score should not exceed 60 without at least one benchmark record attached. Above 60, the source URL must be verifiable.
At launch, scores are updated manually with a ~15–20 minute daily review cadence. Every change must carry a source_url. The data model is built from day one to accept agent-proposed updates as draft entries that require human approval before going live.