AGIdex
agidex/Benchmarks
Evaluation landscape

Benchmarks

Standardized benchmarks used to measure AI system capabilities across reasoning, language, perception, and more.

25 benchmarks
BenchmarkDescriptionLeaderboard
ARC-AGI-2
arc-agi-2
Abstract reasoning benchmark testing fluid intelligence via novel grid puzzles.View →
AudioSet
audioset
Sound classification across 632 audio categories; mean average precision metric.View →
BEHAVIOR-1K
behavior-1k
Long-horizon robot task planning and execution in novel environments.View →
BIG-Bench Hard
big-bench-hard
Hard subset of BIG-Bench testing genuine few-shot generalization.View →
BigToM
bigtom
Theory of mind at scale; tests higher-order belief reasoning.View →
CLadder
cladder
Causal reasoning benchmark grounded in Pearl's causal hierarchy.View →
CULTURAL-BENCH
cultural-bench
Cultural knowledge and reasoning across 45 countries, 1,000+ questions.View →
Chatbot Arena (LMSYS)
chatbot-arena
Human preference ELO ranking across language generation and conversation tasks.View →
EmpatheticDialogues
empathetic-dialogues
Empathetic response generation in open-domain conversation.View →
FLORES-200
flores-200
Translation quality benchmark across 200 languages.View →
FOLIO
folio
First-order logic reasoning benchmark in natural language.View →
FrontierMath
frontiermath
Novel mathematical problem solving requiring genuine reasoning, not recall.View →
GAIA
gaia
Multi-step real-world task completion with verifiable answers; levels 1–3.View →
GPQA Diamond
gpqa-diamond
Graduate-level reasoning in physics, biology, and chemistry. Resists pattern matching.View →
ImageNet-1K
imagenet
Image classification benchmark; 1,000 object categories, top-1 accuracy metric.View →
LibriSpeech
librispeech
Speech recognition benchmark on clean and noisy audio; word error rate metric.View →
MMMU
mmmu
Massive multi-discipline multimodal understanding across text and images.View →
ManiSkill
maniskill
Robotic manipulation benchmark across dexterous and contact-rich tasks.View →
ManiSkill-ViTac
maniskill-vitac
Tactile and vision-tactile fusion manipulation challenge.View →
PARTNR
partnr
Human-robot collaboration across 100,000 household tasks in simulation.View →
Split-CIFAR-100
split-cifar-100
Continual learning benchmark measuring catastrophic forgetting across tasks.View →
SuperGLUE
superglue
Multi-task NLU benchmark testing reading comprehension, coreference, and entailment.View →
VTAB-1k
vtab-1k
Transfer learning across 19 diverse visual tasks from 1,000 examples.View →
VoxCeleb1
voxceleb
Speaker verification benchmark; equal error rate metric.View →
WinoGrande
winogrande
Commonsense reasoning via pronoun resolution; resists surface shortcuts.View →