Standardized benchmarks used to measure AI system capabilities across reasoning, language, perception, and more.
| Benchmark | Description | Leaderboard |
|---|---|---|
ARC-AGI-2 arc-agi-2 | Abstract reasoning benchmark testing fluid intelligence via novel grid puzzles. | View → |
AudioSet audioset | Sound classification across 632 audio categories; mean average precision metric. | View → |
BEHAVIOR-1K behavior-1k | Long-horizon robot task planning and execution in novel environments. | View → |
BIG-Bench Hard big-bench-hard | Hard subset of BIG-Bench testing genuine few-shot generalization. | View → |
BigToM bigtom | Theory of mind at scale; tests higher-order belief reasoning. | View → |
CLadder cladder | Causal reasoning benchmark grounded in Pearl's causal hierarchy. | View → |
CULTURAL-BENCH cultural-bench | Cultural knowledge and reasoning across 45 countries, 1,000+ questions. | View → |
Chatbot Arena (LMSYS) chatbot-arena | Human preference ELO ranking across language generation and conversation tasks. | View → |
EmpatheticDialogues empathetic-dialogues | Empathetic response generation in open-domain conversation. | View → |
FLORES-200 flores-200 | Translation quality benchmark across 200 languages. | View → |
FOLIO folio | First-order logic reasoning benchmark in natural language. | View → |
FrontierMath frontiermath | Novel mathematical problem solving requiring genuine reasoning, not recall. | View → |
GAIA gaia | Multi-step real-world task completion with verifiable answers; levels 1–3. | View → |
GPQA Diamond gpqa-diamond | Graduate-level reasoning in physics, biology, and chemistry. Resists pattern matching. | View → |
ImageNet-1K imagenet | Image classification benchmark; 1,000 object categories, top-1 accuracy metric. | View → |
LibriSpeech librispeech | Speech recognition benchmark on clean and noisy audio; word error rate metric. | View → |
MMMU mmmu | Massive multi-discipline multimodal understanding across text and images. | View → |
ManiSkill maniskill | Robotic manipulation benchmark across dexterous and contact-rich tasks. | View → |
ManiSkill-ViTac maniskill-vitac | Tactile and vision-tactile fusion manipulation challenge. | View → |
PARTNR partnr | Human-robot collaboration across 100,000 household tasks in simulation. | View → |
Split-CIFAR-100 split-cifar-100 | Continual learning benchmark measuring catastrophic forgetting across tasks. | View → |
SuperGLUE superglue | Multi-task NLU benchmark testing reading comprehension, coreference, and entailment. | View → |
VTAB-1k vtab-1k | Transfer learning across 19 diverse visual tasks from 1,000 examples. | View → |
VoxCeleb1 voxceleb | Speaker verification benchmark; equal error rate metric. | View → |
WinoGrande winogrande | Commonsense reasoning via pronoun resolution; resists surface shortcuts. | View → |