AI benchmarks & best tools
Real, measured benchmark scores for the leading AI models, plus our pick of the best AI tool in each category: coding, image, video, audio, and agents. Refreshed three times a week.
Updated July 27, 2026
The state of play
Claude Opus 5 and GPT-5.6 Sol from Anthropic and OpenAI respectively lead in most AI categories. For small businesses, choosing the tool that aligns best with specific needs, such as coding or image generation, is crucial for maximizing efficiency and output quality.
Best AI tool by category
AI coding assistants
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Highest coding intelligence scores for complex tasks.
Also worth a look
- GPT-5.6 Sol (max) · OpenAI's offering with strong coding capabilities.
- Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) · Slightly lower score but still excellent performance.
AI image generation
Kimi K3
Leading intelligence for image generation with a focus on quality.
Also worth a look
- GPT-5.6 Sol (xhigh) · Balanced performance in coding and image generation.
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) · Good alternative for image generation with fallback support.
AI video generation
GPT-5.6 Sol (xhigh)
Best balance of intelligence and video generation capabilities.
Also worth a look
- Claude Opus 5 (Adaptive Reasoning, High Effort) · Solid performance in video generation with high effort setting.
- Kimi K3 · Good alternative with a focus on quality in image and video.
AI audio and voice
GPT-5.6 Sol (high)
High intelligence scores with a focus on voice generation.
Also worth a look
- Claude Opus 5 (Adaptive Reasoning, Medium Effort) · Offers voice generation with a medium effort setting.
- Kimi K3 · Competitive option for voice generation with quality focus.
Long-horizon and agentic tasks
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Highest intelligence scores for long-horizon and agentic tasks.
Also worth a look
- GPT-5.6 Sol (max) · From OpenAI with strong performance in long-term tasks.
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) · Fallback support adds reliability for complex tasks.
Everyday writing and chat
GPT-5.6 Sol (high)
Balanced intelligence and performance for everyday writing tasks.
Also worth a look
- Claude Opus 5 (Adaptive Reasoning, Medium Effort) · Good for everyday tasks with a medium effort setting.
- Kimi K3 · Another solid option with a focus on quality in writing.
LLM benchmark scores
Measured by Artificial Analysis. Higher is better, except price. Sorted by overall intelligence.
| Model | Intelligence | Coding | Math | MMLU-Pro | GPQA | Speed | $/1M |
|---|---|---|---|---|---|---|---|
| Claude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | 61 | 78 | 0 | 0% | 93% | 44/s | $10.00 |
| Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)Anthropic | 60 | 77 | 0 | 0% | 94% | 60/s | $10.00 |
| Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Anthropic | 60 | 77 | 0 | 0% | 93% | 58/s | $20.00 |
| GPT-5.6 Sol (max)OpenAI | 59 | 77 | 0 | 0% | 94% | 74/s | $11.25 |
| Claude Opus 5 (Adaptive Reasoning, High Effort)Anthropic | 59 | 77 | 0 | 0% | 94% | 50/s | $10.00 |
| GPT-5.6 Sol (xhigh)OpenAI | 58 | 78 | 0 | 0% | 93% | 64/s | $11.25 |
| Kimi K3Kimi | 57 | 76 | 0 | 0% | 94% | 33/s | $6.00 |
| Claude Opus 5 (Adaptive Reasoning, Medium Effort)Anthropic | 56 | 74 | 0 | 0% | 92% | 56/s | $10.00 |
| GPT-5.6 Sol (high)OpenAI | 56 | 77 | 0 | 0% | 93% | 66/s | $11.25 |
| Claude Opus 4.8 (Adaptive Reasoning, Max Effort)Anthropic | 56 | 74 | 0 | 0% | 92% | 63/s | $10.00 |
| GPT-5.6 Terra (max)OpenAI | 55 | 77 | 0 | 0% | 93% | 128/s | $5.63 |
| GPT-5.5 (xhigh)OpenAI | 55 | 75 | 0 | 0% | 94% | 0/s | $11.25 |
| Grok 4.5 (high)SpaceXAI | 54 | 72 | 0 | 0% | 93% | 56/s | $3.00 |
| GPT-5.6 Sol (medium)OpenAI | 54 | 76 | 0 | 0% | 93% | 60/s | $11.25 |
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort)Anthropic | 54 | 74 | 0 | 0% | 91% | 0/s | $10.00 |
| Claude Sonnet 5 (Adaptive Reasoning, Max Effort)Anthropic | 53 | 72 | 0 | 0% | 91% | 83/s | $4.00 |
Benchmark scores: Artificial Analysis · Search interest: DataForSEO · Category analysis: Moonshot AI, curated by Marin AI. Scores reflect public benchmarks and may not capture every real-world strength. Tool picks are our editorial view, refreshed three times a week.
Not sure which tools fit your business?
Knowing the best tool is the easy part. We help you pick the right ones, build them into your workflow, and train your team. Book a free fit call.
Book a free fit call