Model rankings
Every model, ranked on quality, speed and cost.
Three things decide a model: how well it thinks, how fast it answers, and what one finished task costs. Every model we serve, plus what's coming, on all three.
What are you building?
The best model for one job is rarely the best for the next. Find the row that sounds like your week.
-
Agentic coding
Many steps, many files, and the model has to keep its own plan straight for an hour. A one-shot index only half predicts that, so run your own repo through the top two.
- Pick
- GPT 5.6 Sol orClaude Opus 5
- On a budget
- DeepSeek V4 Pro $0.27/task at top effort
-
The genuinely hard problems
Research, analysis and anything where being wrong is expensive.
- Pick
- Claude Opus 5 orClaude Fable 5
- On a budget
- Kimi K3 $0.84/task at top effort
-
Anything a person waits on
Chat, voice, autocomplete. The first token has to land before attention does.
- Pick
- Gemini 3.7 Flash soonorGPT 5.6 Luna
-
Documents, images, audio, video
One request carrying a whole corpus of mixed media, reasoned over as a unit.
- Pick
- Gemini 3.7 Flash soonorMuse Spark 1.2 soonorGemini 3.6 Flash
-
High volume, tight budget
Classification, extraction, tagging — millions of calls where the bill is the constraint.
- Pick
- GPT 5.6 Luna orDeepSeek V4 Flash soon
-
Weights you can run yourself
Data that cannot leave your environment, or volume that beats renting.
- Pick
- Kimi K3 orDeepSeek V4 Pro
- On a budget
- DeepSeek V4 Flash soon $0.11/task at top effort
Starting cheap? mindshub_air is our
low-cost tier — an alias pointed at the best-scoring model at the cheap
end of the rate card
(today GPT 5.6 Luna), with
5M tokens a month included. No web search, no reasoning dial, not a
frontier model — the name is the tier. It has no row here because it is
an alias, not a model.
Every model, every axis.
Sort by what you care about, filter to what you can use, and expand a row for the parts that don't fit in a cell.
| Tops the intelligence index; built for long agentic runs | 63 Frontier | 72.4 | 55 tok/sec | 3.2s first token | $2.34 | $5 / $25 our rate | 1M | Closed |
| ||||||||
| Anthropic's top tier for the hardest problems | 62 Frontier | 75 | 70 tok/sec | 3.6s first token | $3.14 | $10 / $50 our rate | 1M | Closed |
| ||||||||
| Flagship all-rounder, and first on the coding index | 61 Frontier | 75.6 | 73 tok/sec | 2.0s first token | $1.01 | $5 / $30 our rate | 1M | Closed |
| ||||||||
| Frontier-class coding at a low API price | 60 Near-frontier | 65 | 84 tok/sec | 1.6s first token | $0.68 | $1.40 / $4.40 list price | 1M | API only |
| ||||||||
| Long autonomous coding runs you can self-host | 60 Near-frontier | 60.2 | 39 tok/sec | 3.4s first token | $0.84 | $3 / $15 our rate | 1M | Open Moonshot licence |
| ||||||||
| Reasoning with access to the live web | 60 Near-frontier | 60 | 61 tok/sec | 1.2s first token | $1.23 | $2 / $6 our rate | 500K | Closed |
| ||||||||
| The open checkpoint that matches the Qwen flagship | 58 Near-frontier | 56 | 24 tok/sec | 4.1s first token | $0.81 | $2 / $6 list price | 984K | Open Qwen licence (commercial use restricted) |
| ||||||||
| Multimodal and multilingual work on a budget | 58 Near-frontier | 57 | 21 tok/sec | 3.9s first token | $0.91 | $2 / $6 list price | 1M | Closed |
| ||||||||
| A high score per dollar, on MIT-licensed weights | 57 Strong | 62 | 50 tok/sec | 1.5s first token | $0.09 | $0.15 / $0.50 list price | 1M | Open MIT |
| ||||||||
| Cheap closed model that powers the free Meta AI | 57 Strong | 54 | 158 tok/sec | 0.7s first token | $0.40 | $1.25 / $4.25 list price | 1M | Closed |
| ||||||||
| Balanced mid-tier for everyday work | 57 Strong | 66 | 109 tok/sec | 1.3s first token | $0.53 | $2 / $12 our rate | 1M | Closed |
| ||||||||
| Fast multimodal work at a promotional price | 56 Strong | 72 | 362 tok/sec | 0.5s first token | $0.40 | $0.75 / $3.75 list price | 1M | Closed |
| ||||||||
| The previous Grok, and now the better-value one | 56 Strong | 56 | 56 tok/sec | 1.4s first token | $0.43 | $2 / $6 our rate | 500K | Closed |
| ||||||||
| Balanced everyday workhorse | 55 Strong | 70 | 92 tok/sec | 2.4s first token | $1.72 | $2 / $10 our rate | 1M | Closed |
| ||||||||
| Every input type, at an open-model rate | 54 Solid | 52 | 150 tok/sec | 0.8s first token | $0.44 | $1.25 / $4.25 our rate | 1M | Closed |
| ||||||||
| Coding specialist for repo-scale changes | 54 Solid | 68 | 136 tok/sec | 1.7s first token | — | $1.75 / $14 our rate | 400K | Closed |
| ||||||||
| Open-weight reasoning and coding at a budget rate | 53 Solid | 58.6 | 68 tok/sec | 2.6s first token | $0.27 | $1.74 / $3.48 our rate | 1M | Open MIT |
| ||||||||
| Open weights with the GLM coding profile | 53 Solid | 62 | 69 tok/sec | 1.8s first token | $0.44 | $1.40 / $4.40 our rate | 1M | Open |
| ||||||||
| Fast and cheap for high-volume work | 52 Solid | 48 | 131 tok/sec | 0.5s first token | $0.05 | $0.20 / $1.20 our rate | 1M | Closed |
| ||||||||
| Low-cost, MIT-licensed, and you can host it yourself | 52 Solid | 52 | 122 tok/sec | 0.9s first token | $0.11 | $0.44 / $1.32 list price | 1M | Open MIT |
| ||||||||
| Fast, high-volume everyday tasks | 52 Solid | 68.8 | 203 tok/sec | 0.6s first token | $0.34 | $0.75 / $3.75 our rate | 1M | Closed |
| ||||||||
| Small and cheap for routine automation | 48 Light | 50 | 150 tok/sec | 0.7s first token | $0.19 | $0.75 / $4.50 our rate | 400K | Closed |
| ||||||||
| Huge documents, images, audio and video | 48 Light | 68.8 | 117 tok/sec | 1.6s first token | $0.33 | $2 / $12 our rate | 1M | Closed |
| ||||||||
| Korea's sovereign open model, free to download | 47 Light | — | — | — | — | — host-dependent | 262K | Open Motif research licence (non-commercial) |
| ||||||||
| A quietly strong multimodal all-rounder, very cheap | 45 Light | 44 | 114 tok/sec | 0.9s first token | $0.14 | $0.30 / $1.20 list price | 1M | Open MiniMax Community (commercial use restricted) |
| ||||||||
| Apache 2.0 weights that take audio in | 42 Light | — | 50 tok/sec | 2.9s first token | $0.34 | $1 / $4.05 list price | 1M | Open Apache 2.0 |
| ||||||||
| The same licence and inputs, two thirds smaller | 41 Light | — | 109 tok/sec | 1.7s first token | $0.07 | $0.30 / $1.20 list price | 1M | Open Apache 2.0 |
| ||||||||
| Multilingual coverage at a fraction of the flagship rate | 39 Basic | 55 | 56 tok/sec | 1.5s first token | $0.17 | $0.40 / $1.60 our rate | 1M | Closed |
| ||||||||
| Built for throughput on bulk, simple work | 38 Basic | 30 | 220 tok/sec | 0.4s first token | $0.05 | $0.20 / $1.25 our rate | 400K | Closed |
| ||||||||
| A small open agent model that runs on one consumer GPU | 35 Basic | — | 107 tok/sec | 0.8s first token | $0.06 | $0.32 / $1.35 list price | 131K | Open Apache 2.0 |
| ||||||||
| Fastest Claude, for high-volume simple work | 30 Basic | 43.4 | 119 tok/sec | 1.1s first token | $0.22 | $1 / $5 our rate | 200K | Closed |
| ||||||||
| Built for throughput on the routine steps of an agent run | 24 Basic | — | 300 tok/sec | 1.3s first token | $0.08 | $0.07 / $0.22 list price | 1M | Open OpenMDW-1.1 |
| ||||||||
- Quality & Cost/task
- Artificial Analysis's published figures, at each model's top reasoning effort — so both are a ceiling. Expand a row for the level we send.
- Coding, Speed & Latency
- Speed is Artificial Analysis's published median where they publish one; coding and latency are our own read. Close values are ties, not differences.
- Price
- Per 1M tokens, in / out. Pure token prices — the 5% fee is a separate line on the bill. Full rate card.
- Missing figures
- ≈ is our estimate. — means nobody has published one.
- Checked
- . Our rates are live from the catalog.
The top of the board is flat.
The top 6 scorers sit within 3 points of each other, which is close enough to be noise. When quality is that level, speed and cost decide.
A benchmark is a generic picture, not a verdict on your work. It averages somebody else's tasks, and the average is exactly where a model's weak spot hides — one that aces graduate physics can still write clunky support replies. Use these scores to draw a shortlist. Then run your own tasks through the top two or three and keep the one whose answers you would actually send; there's a way to do that in an afternoon below.
One thing the number hides: every entry is a model at a reasoning-effort setting, and labs submit at maximum. So these are ceilings. Turning the dial down moves cost further than switching models does — across Claude Opus 5's five settings, token use on the same suite spans eight-fold. A frontier model turned down usually beats a mid-tier model turned up, on both.
The model-by-model rundown What a benchmark hides Does a new release change your pick? Point a coding agent at any of them
- Claude Opus 5 63
- Claude Fable 5 62
- GPT 5.6 Sol 61
- GLM-5.3 soon 60
- Kimi K3 60
- Grok 4.6 60
- Qwen3.8 2.4T A95B soon 58
- Qwen3.8-Max soon 58
- GLM-5.3-Flash soon 57
- Muse Spark 1.2 soon 57
- GPT 5.6 Terra 57
- Gemini 3.7 Flash soon 56
- Grok 4.5 56
- Claude Sonnet 5 55
- Muse Spark 1.1 ≈54
- GPT 5.3 Codex ≈54
- DeepSeek V4 Pro 53
- GLM 5.2 53
- GPT 5.6 Luna 52
- DeepSeek V4 Flash soon 52
- Gemini 3.6 Flash 52
- GPT 5.4 Mini ≈48
- Gemini 3.1 Pro Preview 48
- Motif 3 soon 47
- MiniMax-M3 soon 45
- Inkling soon 42
- Inkling Small soon 41
- Qwen3.7 Plus 39
- GPT 5.4 Nano ≈38
- Muse Glimmer soon 35
- Claude Haiku 4.5 30
- Nemotron 3.5 Lightning soon 24
Bars run from zero and are relative to the leader — a truncated axis would manufacture a gap that isn't there. Scores within 3 points of each other are a tie.
Speed is two numbers, not one.
Latency is the wait before an answer starts. Throughput is how fast the rest arrives. Plenty of models are good at one and poor at the other, which is why the chart below shows both.
- Gemini 3.7 Flash soon 3.3s
- Nemotron 3.5 Lightning soon 4.6s
- GPT 5.4 Nano 4.9s
- Gemini 3.6 Flash 5.5s
- Muse Spark 1.2 soon 7.0s
- GPT 5.4 Mini 7.4s
- Muse Spark 1.1 7.5s
- GPT 5.6 Luna 8.1s
- GPT 5.3 Codex 9.1s
- DeepSeek V4 Flash soon 9.1s
- Claude Haiku 4.5 9.5s
- MiniMax-M3 soon 9.7s
- Muse Glimmer soon 10.1s
- Gemini 3.1 Pro Preview 10.1s
- GPT 5.6 Terra 10.5s
- Inkling Small soon 10.9s
- Claude Sonnet 5 13.3s
- GLM-5.3 soon 13.5s
- GPT 5.6 Sol 15.7s
- GLM 5.2 16.3s
- DeepSeek V4 Pro 17.3s
- Grok 4.6 17.6s
- Claude Fable 5 17.9s
- Grok 4.5 19.3s
- Qwen3.7 Plus 19.4s
- Claude Opus 5 21.4s
- GLM-5.3-Flash soon 21.5s
- Inkling soon 22.9s
- Kimi K3 29.0s
- Qwen3.8 2.4T A95B soon 45.8s
- Qwen3.8-Max soon 51.5s
Wait for the first token Streaming the answer
The spread is 6-fold
GPT 5.4 Nano finishes that answer in about 5 seconds and Kimi K3 takes about 29. Same answer length, same request. Nobody waits 29 seconds for a chat reply, and both of those models are on this board.
Short answers change what to rank on
On a 200-token reply the wait is 25% to 66% of the wall clock, so latency is worth weighting. By a few thousand tokens it is a rounding error and throughput decides alone. The control above the chart switches between the two views.
Reasoning is charged in seconds too
The models that think longest before answering are the ones with the worst latency, because the thinking happens before the first token appears. Turning reasoning effort down shortens the wait and the bill at the same time.
Throughput is a batch-job number
Nobody watches a nightly summarisation run. If no human is waiting, optimise tokens per second and ignore latency entirely — and check whether the provider has a batch endpoint, which often halves the bill for work that is not time-sensitive.
Your numbers will differ
Region, prompt length, cache hit rate and reasoning effort all move these figures. Use them to rank candidates, then measure the two you shortlist against your own traffic.
The price page stopped predicting the bill.
A price per million tokens is a unit price. It says what one unit costs and nothing about how many units the job will take — and that second number now varies between models far more than the first one does.
Models now decide for themselves how long to think, and you pay for that hidden reasoning exactly as for the visible answer. A model five times cheaper per token can land within pennies of one that isn't. Rank on cost per task instead.
It is somebody else's task, at each model's top effort, priced at the provider's own rate — so treat it as a way to rank candidates, not a forecast of your bill.
Why cheap models often aren't Two levers that cut the bill The full rate card
- Nemotron 3.5 Lightning 90.9× cheaper on paper 12.6× cheaper per task
- Grok 4.6 3.3× cheaper on paper 1.2× dearer per task
- GLM-5.3-Flash 40.0× cheaper on paper 11.2× cheaper per task
- Claude Sonnet 5 2.0× cheaper on paper 1.7× dearer per task
- GLM-5.3 4.5× cheaper on paper 1.5× cheaper per task
Both figures come from the same rate card — each provider's own published price, which is what these tasks were costed at, and not always the rate in the Price column above. And the lesson is not that cheap models are a lie: Gemini 3.1 Pro Preview beats its own rate card — 1.7× cheaper on paper, 3.1× cheaper per task. It is that you cannot rank them by reading price pages.
How to actually decide, in an afternoon.
No leaderboard can rank models on your work. This is the shortest thing that can.
-
Shortlist two or three from the board
Use the task router above, or filter the board to your constraint — open weights, a modality, a latency ceiling. Stop at three.
-
Collect ten real tasks
With the messy context attached. Not clean samples: the actual ones you ran last week, including the awkward ones.
-
Run them at the effort you would really use
Not maximum. The reasoning dial moves cost further than switching models does, and every published cost-per-task figure is measured at the top setting.
-
Read the token counts off the response
Every provider returns a
usageobject. Multiply input, cached input and output separately — those three rates can differ by an order of magnitude. -
Divide by the results you would actually ship
If eight of ten came back usable, your real cost is the total divided by eight. That division is the one no leaderboard can do for you.
One API key reaches every model here, so a bake-off is a string change rather than three integrations. There's a runnable version, an afternoon costing recipe, a guide to running evals, and a free working session if you'd rather do it with an engineer.