AGIdex
agidex/Capabilities/Vision/Visual Reasoning
SUB-CAPABILITY · VISION

Visual Reasoning

68%
vs. HUMAN BASELINE = 100
PartialConf · High

Scoring rubric

0–20
Early research
20–40
Narrow benchmark competence
40–60
Strong benchmark · reliability gaps
60–80
Human-competitive in common scenarios
80–100
Comparable to typical skilled adult
100+
Reliably exceeds typical human baseline

Score vs. baseline

Trend · Last 12 months
12mo ago   %
6mo ago   %
Today   %
Δ 12mo   +0

What this measures

Human baseline

Answer straightforward questions about image content — count, color, position.

Human frontier

Multi-step spatial and diagrammatic reasoning across disciplines without domain training.

Current state

What works

VQA on natural images; chart and figure understanding in academic papers.

Key gaps

3D spatial reasoning; novel diagram types not seen during training.

Evidence

2 sources
TechnologyQualitySourceScore vs. baselineScore
Claude Opus 4.7independentvellum.ai/blog/claude-opus-4-7-benchmarks-explained
84%
Gemini 3.1 Provendor reporteddeepmind.google/models/model-cards/gemini-3-1-pro
80%