Model rankings
Models ranked by quality, speed, and cost.
Compare the live MindsHub catalog by capability, response time, and measured cost per task. For models outside the catalog, see the blog.
What are you building?
Start with the kind of work you need to run.
-
Agentic coding
Many steps, many files, and the model has to keep its own plan straight for an hour. A one-shot index only half predicts that, so run your own repo through the top two.
- Pick
- GPT 5.6 Sol orClaude Opus 5
- On a budget
- DeepSeek V4-Pro-0813 $0.27/task at top effort
-
The genuinely hard problems
Research, analysis and anything where being wrong is expensive.
- Pick
- Claude Opus 5 orClaude Fable 5
- On a budget
- Kimi K3 $0.84/task at top effort
-
Anything a person waits on
Chat, voice, autocomplete. The first token has to land before attention does.
- Pick
- Gemini 3.7 Flash orGPT 5.6 Luna
-
Documents, images, audio, video
One request carrying a whole corpus of mixed media, reasoned over as a unit.
- Pick
- Gemini 3.7 Flash orMuse Spark 1.2
-
High volume, tight budget
Classification, extraction, tagging — millions of calls where the bill is the constraint.
- Pick
- GPT 5.6 Luna orGPT 5.4 Nano
-
Weights you can run yourself
Data that cannot leave your environment, or volume that beats renting.
- Pick
- Kimi K3 orDeepSeek V4-Pro-0813
Not sure which? Start with mindshub_air.
We keep this alias pointed at a balanced default, so its underlying
model can change without a code change. The first 5M tokens each month
are included. No web search and no reasoning dial. Its board row
is a snapshot and may briefly lag an alias change.
Compare the full catalog.
Sort and filter the table, then expand a row for context, modalities, reasoning settings, and billing notes.
| Frontier reasoning built for long, many-step agentic runs | 63 Frontier | 72.4 | 54 tok/sec | 3.2s first token | $2.34 | $5 / $25 our rate | 1M | Closed |
| ||||||||
| The previous Fable, for traffic already validated against it | 62 Frontier | 75 | 67 tok/sec | 3.6s first token | $3.14 | $10 / $50 our rate | 1M | Closed |
| ||||||||
| OpenAI's flagship all-rounder | 61 Frontier | 75.6 | 77 tok/sec | 2.0s first token | $0.95 | $5 / $30 our rate | 1M | Closed |
| ||||||||
| Long autonomous coding runs you can self-host | 60 Near-frontier | 60.2 | 38 tok/sec | 3.4s first token | $0.84 | $3 / $15 our rate | 1M | Open Kimi K3 licence (commercial use restricted) |
| ||||||||
| Reasoning with access to the live web | 60 Near-frontier | 60 | 55 tok/sec | 1.2s first token | $1.23 | $2 / $6 our rate | 500K | Closed |
| ||||||||
| The open checkpoint that matches the Qwen flagship | 58 Near-frontier | 56 | 41 tok/sec | 4.1s first token | $0.81 | $2 / $6 our rate | 984K | Open Qwen licence (commercial use restricted) |
| ||||||||
| Cheap closed model that powers the free Meta AI | 57 Strong | 54 | 117 tok/sec | 0.7s first token | $0.40 | $1.25 / $4.25 our rate | 1M | Closed |
| ||||||||
| Balanced mid-tier for everyday work | 57 Strong | 66 | 108 tok/sec | 1.3s first token | $0.53 | $2 / $12 our rate | 1M | Closed |
| ||||||||
| Fast multimodal work at a promotional price | 56 Strong | 72 | 285 tok/sec | 0.5s first token | $0.40 | $0.75 / $3.75 our rate | 1M | Closed |
| ||||||||
| The previous Grok, and now the better-value one | 56 Strong | 56 | 48 tok/sec | 1.4s first token | $0.43 | $2 / $6 our rate | 500K | Closed |
| ||||||||
| Balanced everyday workhorse | 55 Strong | 70 | 72 tok/sec | 2.4s first token | $1.72 | $2 / $10 our rate | 1M | Closed |
| ||||||||
| Every input type, at an open-model rate | 54 Solid | 52 | 150 tok/sec | 0.8s first token | $0.44 | $1.25 / $4.25 our rate | 1M | Closed |
| ||||||||
| Coding specialist for repo-scale changes | 54 Solid | 68 | 127 tok/sec | 1.7s first token | — | $1.75 / $14 our rate | 400K | Closed |
| ||||||||
| Open-weight reasoning and coding at a budget rate | 53 Solid | 58.6 | 54 tok/sec | 2.6s first token | $0.27 | $1.32 / $3.96 our rate | 1M | Open MIT |
| ||||||||
| The previous GLM release, with the same coding profile | 53 Solid | 62 | 68 tok/sec | 1.8s first token | $0.44 | $1.40 / $4.40 our rate | 1M | Open MIT |
| ||||||||
| Our alias, kept on whichever model balances best | 52 Solid | 48 | 128 tok/sec | 0.5s first token | $0.05 | $0.20 / $1.20 our rate | 1M | Closed |
| ||||||||
| Fast and cheap for high-volume work | 52 Solid | 48 | 128 tok/sec | 0.5s first token | $0.05 | $0.20 / $1.20 our rate | 1M | Closed |
| ||||||||
| Fast, high-volume everyday tasks | 52 Solid | 68.8 | 167 tok/sec | 0.6s first token | $0.34 | $0.75 / $3.75 our rate | 1M | Closed |
| ||||||||
| Small and cheap for routine automation | 48 Light | 50 | 150 tok/sec | 0.7s first token | $0.19 | $0.75 / $4.50 our rate | 400K | Closed |
| ||||||||
| Huge documents, images, audio and video | 48 Light | 68.8 | 103 tok/sec | 1.6s first token | $0.33 | $2 / $12 our rate | 1M | Closed |
| ||||||||
| Multilingual coverage at a fraction of the flagship rate | 39 Basic | 55 | 57 tok/sec | 1.5s first token | $0.17 | $0.40 / $1.60 our rate | 1M | Closed |
| ||||||||
| Built for throughput on bulk, simple work | 38 Basic | 30 | 220 tok/sec | 0.4s first token | $0.05 | $0.20 / $1.25 our rate | 400K | Closed |
| ||||||||
| Fastest Claude, for high-volume simple work | 30 Basic | 43.4 | 96 tok/sec | 1.1s first token | $0.22 | $1 / $5 our rate | 200K | Closed |
| ||||||||
| New in the catalog, not rated here yet | — | — | — | — | — | $1.50 / $9 our rate | — | Closed |
| ||||||||
| New in the catalog, not rated here yet | — | — | — | — | — | $0.50 / $3 our rate | — | Closed |
| ||||||||
| New in the catalog, not rated here yet | — | — | — | — | — | $0.25 / $1.50 our rate | — | Closed |
| ||||||||
| New in the catalog, not rated here yet | — | — | — | — | — | $1.74 / $3.48 our rate | — | Closed |
| ||||||||
- Quality & Cost/task
- Artificial Analysis's published figures, on Intelligence Index v4.1.1 — they measure independently of us — at each model's top reasoning effort, so both are a ceiling. Expand a row for the level we send.
- Coding, Speed & Latency
- Speed is Artificial Analysis's published median where they publish one; coding and latency are our own read. Close values are ties, not differences.
- Price
- Per 1M tokens, in / out. Everything else about the bill — prompt-cache rates, long-prompt tiers and search — is on the full rate card.
- Missing figures
- ≈ is our estimate. — means nobody has published one.
- Checked
- . Our rates are live from the catalog.
Small score differences are ties.
The top 5 scorers sit within 3 points of each other, which is close enough to be noise. When quality is that level, speed and cost decide.
A benchmark averages many tasks and can hide weaknesses that matter to you. Use the scores to shortlist two or three models, then test them on your own work. The method below shows how.
Each score is for a model at a specific reasoning setting, usually the maximum. Treat it as a ceiling. Across Claude Opus 5’s five settings, token use on the same suite spans eight-fold.
The model-by-model rundown What a benchmark hides Does a new release change your pick? Point a coding agent at any of them
- Claude Opus 5 63
- Claude Fable 5 62
- GPT 5.6 Sol 61
- Kimi K3 60
- Grok 4.6 60
- Qwen3.8-2.4T-A95B 58
- Muse Spark 1.2 57
- GPT 5.6 Terra 57
- Gemini 3.7 Flash 56
- Grok 4.5 56
- Claude Sonnet 5 55
- Muse Spark 1.1 ≈54
- GPT 5.3 Codex ≈54
- DeepSeek V4-Pro-0813 ≈53
- GLM 5.2 53
- MindsHub Air ≈52
- GPT 5.6 Luna 52
- Gemini 3.6 Flash 52
- GPT 5.4 Mini ≈48
- Gemini 3.1 Pro Preview 48
- Qwen3.7 Plus 39
- GPT 5.4 Nano ≈38
- Claude Haiku 4.5 30
Bars run from zero and are relative to the leader — a truncated axis would manufacture a gap that isn't there. Scores within 3 points of each other are a tie.
Measure the wait and the streaming speed.
Latency is the wait for the first token. Throughput is how quickly the rest arrives. The chart shows both because a model can be strong on one and weak on the other.
- Gemini 3.7 Flash 4.0s
- GPT 5.4 Nano 4.9s
- Gemini 3.6 Flash 6.6s
- GPT 5.4 Mini 7.4s
- Muse Spark 1.1 7.5s
- MindsHub Air 8.3s
- GPT 5.6 Luna 8.3s
- Muse Spark 1.2 9.2s
- GPT 5.3 Codex 9.6s
- GPT 5.6 Terra 10.6s
- Gemini 3.1 Pro Preview 11.3s
- Claude Haiku 4.5 11.5s
- GPT 5.6 Sol 15.0s
- Claude Sonnet 5 16.3s
- GLM 5.2 16.5s
- Claude Fable 5 18.5s
- Qwen3.7 Plus 19.0s
- Grok 4.6 19.4s
- DeepSeek V4-Pro-0813 21.1s
- Claude Opus 5 21.7s
- Grok 4.5 22.2s
- Qwen3.8-2.4T-A95B 28.5s
- Kimi K3 29.7s
Wait for the first token Streaming the answer
A 7-fold spread in completion time
Gemini 3.7 Flash finishes that answer in about 4 seconds and Kimi K3 takes about 30. Same answer length, same request. Nobody waits 30 seconds for a chat reply, and both of those models are on this board.
Short answers make latency more important
On a 200-token reply the wait is 24% to 55% of the wall clock, so latency is worth weighting. By a few thousand tokens it is a rounding error and throughput decides alone. The control above the chart switches between the two views.
Measure with your own traffic
Region, prompt length, cache hit rate and reasoning effort all move these figures. Use them to rank candidates, then measure the two you shortlist against your own traffic.
Price per token is not cost per task.
Token prices do not show how many tokens a model will use to finish the job. Reasoning models can vary widely on that second number.
You pay for reasoning tokens as well as the visible answer. A model with a lower token rate can still cost more per completed task.
It is somebody else's task, at each model's top effort, priced at the provider's own rate — so treat it as a way to rank candidates, not a forecast of your bill.
Why cheap models often aren't Two levers that cut the bill The full rate card
- Grok 4.6 3.3× cheaper on paper 1.3× dearer per task
- Claude Sonnet 5 2.0× cheaper on paper 1.8× dearer per task
- Qwen3.8-2.4T-A95B 3.3× cheaper on paper 1.2× cheaper per task
- Gemini 3.7 Flash 5.3× cheaper on paper 2.4× cheaper per task
- Qwen3.7 Plus 12.5× cheaper on paper 5.6× cheaper per task
Both figures come from the same rate card — each provider's own published price, which is what these tasks were costed at, and not always the rate in the Price column above. And the lesson is not that cheap models are a lie: Gemini 3.1 Pro Preview beats its own rate card — 1.7× cheaper on paper, 2.9× cheaper per task. It is that you cannot rank them by reading price pages.
Choose a model in an afternoon.
Use the board for a shortlist, then test it against real examples.
-
Shortlist two or three from the board
Use the task router above, or filter the board to your constraint — open weights, a modality, a latency ceiling. Stop at three.
-
Collect ten real tasks
With the messy context attached. Not clean samples: the actual ones you ran last week, including the awkward ones.
-
Run them at the effort you would really use
Not maximum. The reasoning dial moves cost further than switching models does, and every published cost-per-task figure is measured at the top setting.
-
Read the token counts off the response
Every provider returns a
usageobject. Multiply input, cached input and output separately — those three rates can differ by an order of magnitude. -
Divide by the results you would actually ship
If eight of ten came back usable, your real cost is the total divided by eight. That division is the one no leaderboard can do for you.
One API key reaches every model here, so a bake-off is a string change rather than three integrations. There's a runnable version, an afternoon costing recipe, a guide to running evals, and a free working session if you'd rather do it with an engineer.