LABORATORY / 2026-09-08

Models, measured

Compare major models on the same task. Every number has a test version, configuration and source.

Models to compare9 / 9

Snapshot: 2026-09-08. Choose 2–4 models for a focused comparison. A dash means no value in this snapshot; some source values are rounded. Check each test's configuration and date.

Vendor comparison

DeepSWE v1.1 · release table

Long engineering tasks. OpenAI comparison table; maximum across reported reasoning efforts, not equal compute cost.

  1. GPT-6 Astrarelease-table maximum74.1%
  2. Gemini 3.8 Flashrelease-table maximum73.8%
  3. GPT-5.6 Solrelease-table maximum72.7%
  4. Claude Fable 5.1release-table maximum67.4%

GLM-5.3 · DeepSeek V4 Pro 0813 · Grok 4.6 · Muse Spark 1.3 · Kimi K3

Source and method
Vendor comparison

FrontierMath Tier 4 · v2

The hardest FrontierMath tier. Not comparable with Tier 1–3; see the source for settings.

  1. GPT-6 Astrarelease-table maximum97.6%
  2. Claude Fable 5.1release-table maximum87.8%
  3. GPT-5.6 Solrelease-table maximum83%

Gemini 3.8 Flash · GLM-5.3 · DeepSeek V4 Pro 0813 · Grok 4.6 · Muse Spark 1.3 · Kimi K3

Source and method
Same score, different strengths

Index v4.3 changed on September 7. It is not comparable with July's v4.1 numbers. Start with your task when choosing a model.

Find a model