Apply deductive rules across short chains. Recognize valid vs. invalid arguments in plain language.
Construct novel multi-step proofs in formal logic and mathematics with self-verification.
Short formal proofs, propositional logic, well-scoped puzzles.
Long-horizon chains drift. Self-correction unreliable past ~10 steps.
| Technology | Quality | Source | Score vs. baseline | Score |
|---|---|---|---|---|
| Claude Opus 4.7 | independent | vellum.ai/blog/claude-opus-4-7-benchmarks-explained | 74% | |
| GPT-5.5 | vendor reported | openai.com/index/introducing-gpt-5-5 | 76% | |
| GPT-5.4 | vendor reported | openai.com/research/gpt-5-4 | 72% | |
| Gemini 3.1 Pro | vendor reported | deepmind.google/models/model-cards/gemini-3-1-pro | 71% | |
| DeepSeek V4 Pro | vendor reported | deepseek.com/v4 | 74% | |
| Mistral Large 3 | vendor reported | mistral.ai/news/mistral-large-3 | 68% |