Model rankings

Every model, ranked on quality, speed and cost.

Three things decide a model: how well it thinks, how fast it answers, and what one finished task costs. Every model we serve, plus what's coming, on all three.

Start here

What are you building?

The best model for one job is rarely the best for the next. Find the row that sounds like your week.

  • Agentic coding

    Many steps, many files, and the model has to keep its own plan straight for an hour. A one-shot index only half predicts that, so run your own repo through the top two.

    On a budget
    DeepSeek V4 Pro $0.27/task at top effort
  • The genuinely hard problems

    Research, analysis and anything where being wrong is expensive.

    On a budget
    Kimi K3 $0.84/task at top effort
  • Anything a person waits on

    Chat, voice, autocomplete. The first token has to land before attention does.

  • Documents, images, audio, video

    One request carrying a whole corpus of mixed media, reasoned over as a unit.

  • High volume, tight budget

    Classification, extraction, tagging — millions of calls where the bill is the constraint.

  • Weights you can run yourself

    Data that cannot leave your environment, or volume that beats renting.

    On a budget
    DeepSeek V4 Flash soon $0.11/task at top effort

Starting cheap? mindshub_air is our low-cost tier — an alias pointed at the best-scoring model at the cheap end of the rate card (today GPT 5.6 Luna), with 5M tokens a month included. No web search, no reasoning dial, not a frontier model — the name is the tier. It has no row here because it is an alias, not a model.

Leaderboard

Every model, every axis.

Sort by what you care about, filter to what you can use, and expand a row for the parts that don't fit in a cell.

Large language models ranked by quality index, with coding index, output speed, latency, measured cost per task at top reasoning effort, price per million tokens, context window and weight availability.
Anthropic opus Tops the intelligence index; built for long agentic runs 63 Frontier 72.4 55 tok/sec 3.2s first token $2.34 $5 / $25 our rate 1M Closed
Pick it when
The task is hard, runs for many steps, or the output carries risk that someone has to sign off on.
Skip it when
You're paying by the token for routine work — it keeps chewing long after a cheaper model would have stopped.
Inputs
TextImages
API alias
opus The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
The top of its dial is a plateau, not a peak. This row quotes max; one notch down at xhigh Artificial Analysis measures the same 63 for $1.80 a task instead of $2.34. Read the score as a ceiling and the cost as a worst case.
Anthropic fable Anthropic's top tier for the hardest problems 62 Frontier 75 70 tok/sec 3.6s first token $3.14 $10 / $50 our rate 1M Closed
Pick it when
Nothing else has cleared the bar and the answer is worth a premium rate.
Skip it when
Opus 5 sits one index point away at half the price — start there and only move up if it falls short.
Inputs
TextImages
API alias
fable The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
OpenAI gpt Flagship all-rounder, and first on the coding index 61 Frontier 75.6 73 tok/sec 2.0s first token $1.01 $5 / $30 our rate 1M Closed
Pick it when
You want one provider that is competent at everything, with the deepest tooling ecosystem around it.
Skip it when
Price is the binding constraint — GPT 5.6 Terra scores four points lower for about half the cost per finished task.
Inputs
TextImages
API alias
gpt The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send medium unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Zhipu coming soon Frontier-class coding at a low API price 60 Near-frontier 65 84 tok/sec 1.6s first token $0.68 $1.40 / $4.40 list price 1M API only
Pick it when
You want strong coding and security review at a fraction of frontier prices, and renting is fine.
Skip it when
You need multimodal input, or weights you can host today.
Inputs
Text
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Worth knowing
Zhipu has not released the weights, so this is an API you rent rather than a model you can host.
Price shown
Zhipu's own published list price, not ours — we don't serve this model yet.
Moonshot AI kimi Long autonomous coding runs you can self-host 60 Near-frontier 60.2 39 tok/sec 3.4s first token $0.84 $3 / $15 our rate 1M Open Moonshot licence
Pick it when
You want the strongest model you can genuinely download and run yourself.
Skip it when
The API bill is the point — GLM is cheaper for similar work.
Inputs
TextImages
API alias
kimi The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
Licence
Moonshot licence
Served by
Fireworks AI Moonshot AI makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Worth knowing
The licence is Moonshot's own, not MIT or Apache. Self-hosting and fine-tuning are fine; reselling access at scale needs a read.
SpaceXAI grok Reasoning with access to the live web 60 Near-frontier 60 61 tok/sec 1.2s first token $1.23 $2 / $6 our rate 500K Closed
Pick it when
You need information from this morning in the same call as the reasoning.
Skip it when
Your agents live in a terminal, or your prompts are enormous — the window is half what most of this board offers.
Inputs
TextImages
API alias
grok The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhigh We send high unless you ask for another. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
Its dial stops paying off at the top. This row quotes xhigh, the highest setting Artificial Analysis measures, and one notch down at high the same model scores a point higher for a quarter less per task. What the top setting buys is four tokens a second.
Alibaba coming soon The open checkpoint that matches the Qwen flagship 58 Near-frontier 56 24 tok/sec 4.1s first token $0.81 $2 / $6 list price 984K Open Qwen licence (commercial use restricted)
Pick it when
You need the weights, and text-only input plus a restrictive licence are acceptable.
Skip it when
You are reselling access, or your inputs are not text.
Inputs
Text
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
Qwen licence (commercial use restricted)
Worth knowing
Artificial Analysis costs it at Alibaba's own $2 / $6 — the same rate card as the proprietary Max — but the weights are downloadable, so your actual bill is whatever your host or your hardware charges.
Price shown
Alibaba's own published list price, not ours — we don't serve this model yet.
Alibaba coming soon Multimodal and multilingual work on a budget 58 Near-frontier 57 21 tok/sec 3.9s first token $0.91 $2 / $6 list price 1M Closed
Pick it when
You want a capable multimodal API with a million-token window and strong non-English coverage.
Skip it when
Speed or cost per task is the priority — it is slow and verbose, and keeps almost none of its token discount.
Inputs
TextImagesVideo
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Worth knowing
The Price column puts it at a fifth of GPT 5.6 Sol per output token, and it still comes out only 1.1x cheaper per finished task — it is slow and it thinks at length, and both show up on the invoice.
Price shown
Alibaba's own published list price, not ours — we don't serve this model yet.
Zhipu coming soon A high score per dollar, on MIT-licensed weights 57 Strong 62 50 tok/sec 1.5s first token $0.09 $0.15 / $0.50 list price 1M Open MIT
Pick it when
You want a high score per dollar and MIT weights, and you can absorb a slow first token.
Skip it when
You are waiting on a person — it is one of the slower models here.
Inputs
TextImages
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
MIT
Worth knowing
The score comes at a price in patience: Artificial Analysis measures it as slow and very verbose, so a finished answer takes longer than its cost per task suggests. First natively multimodal model in the GLM-5 line — 320B total parameters, 18B active.
Price shown
Zhipu's own published list price, not ours — we don't serve this model yet.
Meta coming soon Cheap closed model that powers the free Meta AI 57 Strong 54 158 tok/sec 0.7s first token $0.40 $1.25 / $4.25 list price 1M Closed
Pick it when
You want a fast, capable, every-input model and Meta is an acceptable processor of your data.
Skip it when
You need weights — the Muse family is closed — or your data cannot go to Meta.
Inputs
TextImagesAudioVideo
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Worth knowing
Meta publishes a "contributor" tier at roughly a fifteenth of the standard rate in exchange for training on your prompts and replies. Fine for hobby projects, an easy no for anything confidential.
Price shown
Meta's own published list price, not ours — we don't serve this model yet.
OpenAI gpt-terra Balanced mid-tier for everyday work 57 Strong 66 109 tok/sec 1.3s first token $0.53 $2 / $12 our rate 1M Closed
Pick it when
You want the sensible default: strong enough for most work, and it keeps almost all of its discount on the invoice.
Skip it when
The last few index points decide the outcome.
Inputs
TextImages
API alias
gpt-terra The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send medium unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Google coming soon Fast multimodal work at a promotional price 56 Strong 72 362 tok/sec 0.5s first token $0.40 $0.75 / $3.75 list price 1M Closed
Pick it when
Inputs are large or mixed-media — PDFs, images, audio, video — and you want them read quickly.
Skip it when
You need frontier-level reasoning, or weights you can host yourself.
Inputs
TextImagesAudioVideo
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Price shown
Google's own published list price, not ours — we don't serve this model yet. Promotional through December 31, 2026.
SpaceXAI grok-4-5 The previous Grok, and now the better-value one 56 Strong 56 56 tok/sec 1.4s first token $0.43 $2 / $6 our rate 500K Closed
Pick it when
You want most of 4.6 for much less per task, or you have production traffic validated against this version.
Skip it when
You need the four index points 4.6 adds.
Inputs
TextImages
API alias
grok-4-5 The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhigh We send high unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Anthropic sonnet Balanced everyday workhorse 55 Strong 70 92 tok/sec 2.4s first token $1.72 $2 / $10 our rate 1M Closed
Pick it when
You want one model that drafts, summarizes, analyzes and codes without hitting a wall.
Skip it when
Volume is the point. It costs more per finished task than its rate card suggests.
Inputs
TextImages
API alias
sonnet The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
Three times cheaper than GPT 5.6 Sol per output token and about 70% more per finished task — the plainest case on this board that a rate card does not predict a bill. Its streaming rate is the one on this board that moves most with effort: 92 tok/s at the top setting these figures come from, against 68 at the high setting we send by default.
Meta muse-spark Every input type, at an open-model rate 54 Solid 52 150 tok/sec 0.8s first token $0.44 $1.25 / $4.25 our rate 1M Closed
Pick it when
You want every input type at a low rate and a fast first token.
Skip it when
1.2 is three index points better on the same rate card.
Inputs
TextImagesAudioVideo
API alias
muse-spark The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
OpenAI gpt-codex Coding specialist for repo-scale changes 54 Solid 68 136 tok/sec 1.7s first token $1.75 / $14 our rate 400K Closed
Pick it when
The job is software, across many files, and you want a model tuned for that rather than for general reasoning.
Skip it when
The work is not code. Its general index score sits below its coding score for a reason.
Inputs
TextImages
API alias
gpt-codex The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhigh We send medium unless you ask for another. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
Artificial Analysis scores this version only provisionally — on an older component set that omits all three of the benchmarks they run on every fully scored row (GDPval-AA v2, Terminal-Bench v2.1 and t3-Banking) — and publishes no cost per task for it, as they do for none of their provisional scores. Their figure is not on the same scale as the rest of this column, so both published columns here are ours.
DeepSeek deepseek Open-weight reasoning and coding at a budget rate 53 Solid 58.6 68 tok/sec 2.6s first token $0.27 $1.74 / $3.48 our rate 1M Open MIT
Pick it when
Volume is high, the budget is tight, and you want permissive weights as an exit route.
Skip it when
You need multimodal input or the last few points of quality.
Inputs
Text
API alias
deepseek The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Licence
MIT
Served by
Fireworks AI DeepSeek makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Zhipu glm Open weights with the GLM coding profile 53 Solid 62 69 tok/sec 1.8s first token $0.44 $1.40 / $4.40 our rate 1M Open
Pick it when
You want the GLM coding profile and weights you can host, and can spare the points 5.3 adds.
Skip it when
5.3 is seven index points better on the same rate card, if renting is fine.
Inputs
Text
API alias
glm The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
highmax We send max unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Served by
Fireworks AI Zhipu makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Worth knowing
Artificial Analysis now lists these weights as open, while 5.3 is still API-only — so the older release is the one you can host. The licence name is not recorded here; read Zhipu's terms before you build on it.
OpenAI gpt-luna Fast and cheap for high-volume work 52 Solid 48 131 tok/sec 0.5s first token $0.05 $0.20 / $1.20 our rate 1M Closed
Pick it when
Volume is high, tasks are routine, and the bill matters more than the last few points of quality.
Skip it when
The work is genuinely hard. It scores alongside good open models, not alongside the frontier.
Inputs
TextImages
API alias
gpt-luna The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send medium unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
DeepSeek coming soon Low-cost, MIT-licensed, and you can host it yourself 52 Solid 52 122 tok/sec 0.9s first token $0.11 $0.44 / $1.32 list price 1M Open MIT
Pick it when
High-volume routine work where you want the option to stop renting and self-host.
Skip it when
Quality is the constraint rather than cost.
Inputs
Text
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
MIT
Price shown
DeepSeek's own published list price, not ours — we don't serve this model yet.
Google gemini-flash Fast, high-volume everyday tasks 52 Solid 68.8 203 tok/sec 0.6s first token $0.34 $0.75 / $3.75 our rate 1M Closed
Pick it when
You want multimodal input and a million-token window at a throughput price.
Skip it when
3.7 Flash scores four index points better and is stronger at code.
Inputs
TextImagesAudioVideo
API alias
gemini-flash The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
mediumhigh We send medium unless you ask for another — the lowest setting it has. Quality and Cost/task above were measured at high, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
OpenAI gpt-mini Small and cheap for routine automation 48 Light 50 150 tok/sec 0.7s first token $0.19 $0.75 / $4.50 our rate 400K Closed
Pick it when
Classification, extraction, tagging and other high-frequency jobs with a narrow definition of correct.
Skip it when
Anything open-ended. Luna is both cheaper and better at this point in the line-up.
Inputs
TextImages
API alias
gpt-mini The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhigh We send none unless you ask for another — the lowest setting it has, and no reasoning tokens at all. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Google gemini Huge documents, images, audio and video 48 Light 68.8 117 tok/sec 1.6s first token $0.33 $2 / $12 our rate 1M Closed
Pick it when
A single request has to carry a whole corpus of mixed media, and it has to be reasoned over rather than skimmed.
Skip it when
You are paying by the token — the Flash line reads the same inputs at a third of the rate card, and 3.7 Flash outscores this by eight points.
Inputs
TextImagesAudioVideo
API alias
gemini The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
minimallowmediumhigh We send high unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Moreh coming soon Korea's sovereign open model, free to download 47 Light host-dependent 262K Open Motif research licence (non-commercial)
Pick it when
You want a strong open checkpoint for research or internal evaluation and text-only input is fine.
Skip it when
It is going anywhere near a product — the licence rules out commercial use.
Inputs
Text
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
Motif research licence (non-commercial)
Worth knowing
Four columns are blank because nobody has published them. The weights are free and there is no priced first-party endpoint, so Artificial Analysis reports its cost per task as unknown — the bill is whatever your own hardware or host costs. It is also very verbose: 260M output tokens across the Intelligence Index against a 100M median, which is the shape of a model that thinks in long answers.
Price shown
Moreh's own published list price, not ours — we don't serve this model yet.
MiniMax coming soon A quietly strong multimodal all-rounder, very cheap 45 Light 44 114 tok/sec 0.9s first token $0.14 $0.30 / $1.20 list price 1M Open MiniMax Community (commercial use restricted)
Pick it when
You want a million-token window and image input at the bottom of the price range.
Skip it when
You are building a commercial product on the weights — read the licence first.
Inputs
TextImages
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
MiniMax Community (commercial use restricted)
Price shown
MiniMax's own published list price, not ours — we don't serve this model yet.
Thinking Machines Lab coming soon Apache 2.0 weights that take audio in 42 Light 50 tok/sec 2.9s first token $0.34 $1 / $4.05 list price 1M Open Apache 2.0
Pick it when
You need a permissive licence with no commercial catch, and audio input in the same model.
Skip it when
The score is what you are buying — several open models here rank higher for less.
Inputs
TextImagesAudio
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
Apache 2.0
Worth knowing
A genuinely unrestricted licence, at a price: several times the cost per task of its own Small sibling for one index point more, and slower than most of this board. 975B total parameters, 41B active.
Price shown
Thinking Machines Lab's own published list price, not ours — we don't serve this model yet.
Thinking Machines Lab coming soon The same licence and inputs, two thirds smaller 41 Light 109 tok/sec 1.7s first token $0.07 $0.30 / $1.20 list price 1M Open Apache 2.0
Pick it when
You want the Apache 2.0 terms and the multimodal inputs on hardware you can actually afford.
Skip it when
You need the one index point Inkling adds.
Inputs
TextImagesAudio
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
Apache 2.0
Worth knowing
One index point behind Inkling on well under a third of the parameters, twice as fast and a fifth of the cost per task. On this board it is the better of the pair on every axis except the score.
Price shown
Thinking Machines Lab's own published list price, not ours — we don't serve this model yet.
Alibaba qwen Multilingual coverage at a fraction of the flagship rate 39 Basic 55 56 tok/sec 1.5s first token $0.17 $0.40 / $1.60 our rate 1M Closed
Pick it when
You want Qwen-family behaviour and broad multilingual coverage without paying a flagship rate.
Skip it when
You need the current flagship — 3.8-Max is nineteen index points ahead.
Inputs
TextImages
API alias
qwen The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
Served by
Fireworks AI Alibaba makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
OpenAI gpt-nano Built for throughput on bulk, simple work 38 Basic 30 220 tok/sec 0.4s first token $0.05 $0.20 / $1.25 our rate 400K Closed
Pick it when
Millions of tiny calls where latency is the product and the task fits in one sentence.
Skip it when
Any judgement is required.
Inputs
Text
API alias
gpt-nano The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhigh We send none unless you ask for another — the lowest setting it has, and no reasoning tokens at all. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Meta coming soon A small open agent model that runs on one consumer GPU 35 Basic 107 tok/sec 0.8s first token $0.06 $0.32 / $1.35 list price 131K Open Apache 2.0
Pick it when
You want something genuinely local: Apache 2.0 weights, 30B parameters, and a score that still holds up on well-scoped agent steps.
Skip it when
The work is open-ended. This is the smallest capable thing here, not a substitute for the rows above it.
Inputs
TextImages
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
Apache 2.0
Worth knowing
Meta going closed with Muse Spark did not stop it shipping small open models. The 131K window is the shortest on this board, and inputs beyond text and images are not recorded here.
Price shown
Meta's own published list price, not ours — we don't serve this model yet.
Anthropic haiku Fastest Claude, for high-volume simple work 30 Basic 43.4 119 tok/sec 1.1s first token $0.22 $1 / $5 our rate 200K Closed
Pick it when
You want Claude behaviour and Claude tooling on work that does not need Claude-sized thinking.
Skip it when
The job needs reasoning headroom — this is the previous generation, and it shows on hard tasks.
Inputs
TextImages
API alias
haiku The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
NVIDIA coming soon Built for throughput on the routine steps of an agent run 24 Basic 300 tok/sec 1.3s first token $0.08 $0.07 / $0.22 list price 1M Open OpenMDW-1.1
Pick it when
An agent is grinding through thousands of small, well-defined steps and the wall clock is the constraint.
Skip it when
Any step needs judgement — this is the lowest score on the board.
Inputs
Text
Reasoning effort
Not known yet. Which levels a model accepts comes from our catalog, and this one isn't in it yet — so we can't tell you, rather than there being nothing to tell.
Licence
OpenMDW-1.1
Worth knowing
It wins the accuracy-versus-speed frontier in its size class rather than the accuracy one: 31.6B total parameters, 3.6B active, and among the fastest output on this board. Weights, and a permissive licence, but a score that only suits work you have already decomposed.
Price shown
NVIDIA's own published list price, not ours — we don't serve this model yet.
Quality & Cost/task
Artificial Analysis's published figures, at each model's top reasoning effort — so both are a ceiling. Expand a row for the level we send.
Coding, Speed & Latency
Speed is Artificial Analysis's published median where they publish one; coding and latency are our own read. Close values are ties, not differences.
Price
Per 1M tokens, in / out. Pure token prices — the 5% fee is a separate line on the bill. Full rate card.
Missing figures
is our estimate. means nobody has published one.
Checked
. Our rates are live from the catalog.
01 · Quality

The top of the board is flat.

The top 6 scorers sit within 3 points of each other, which is close enough to be noise. When quality is that level, speed and cost decide.

A benchmark is a generic picture, not a verdict on your work. It averages somebody else's tasks, and the average is exactly where a model's weak spot hides — one that aces graduate physics can still write clunky support replies. Use these scores to draw a shortlist. Then run your own tasks through the top two or three and keep the one whose answers you would actually send; there's a way to do that in an afternoon below.

One thing the number hides: every entry is a model at a reasoning-effort setting, and labs submit at maximum. So these are ceilings. Turning the dial down moves cost further than switching models does — across Claude Opus 5's five settings, token use on the same suite spans eight-fold. A frontier model turned down usually beats a mid-tier model turned up, on both.

Quality index, whole board. The flat stretch at the top is the point; the drop-off below it is where the index starts deciding things.
  1. Claude Opus 5 63
  2. Claude Fable 5 62
  3. GPT 5.6 Sol 61
  4. GLM-5.3 soon 60
  5. Kimi K3 60
  6. Grok 4.6 60
  7. Qwen3.8 2.4T A95B soon 58
  8. Qwen3.8-Max soon 58
  9. GLM-5.3-Flash soon 57
  10. Muse Spark 1.2 soon 57
  11. GPT 5.6 Terra 57
  12. Gemini 3.7 Flash soon 56
  13. Grok 4.5 56
  14. Claude Sonnet 5 55
  15. Muse Spark 1.1 54
  16. GPT 5.3 Codex 54
  17. DeepSeek V4 Pro 53
  18. GLM 5.2 53
  19. GPT 5.6 Luna 52
  20. DeepSeek V4 Flash soon 52
  21. Gemini 3.6 Flash 52
  22. GPT 5.4 Mini 48
  23. Gemini 3.1 Pro Preview 48
  24. Motif 3 soon 47
  25. MiniMax-M3 soon 45
  26. Inkling soon 42
  27. Inkling Small soon 41
  28. Qwen3.7 Plus 39
  29. GPT 5.4 Nano 38
  30. Muse Glimmer soon 35
  31. Claude Haiku 4.5 30
  32. Nemotron 3.5 Lightning soon 24

Bars run from zero and are relative to the leader — a truncated axis would manufacture a gap that isn't there. Scores within 3 points of each other are a tie.

02 · Performance

Speed is two numbers, not one.

Latency is the wait before an answer starts. Throughput is how fast the rest arrives. Plenty of models are good at one and poor at the other, which is why the chart below shows both.

From the request to a finished answer. Pale is the wait before the first token, solid is the time to stream the rest. A ranking, not a stopwatch: the streaming rate is Artificial Analysis's published median wherever they publish one, the first-token wait is our own read on every row. Answer length changes both the stakes and the order. A 200-token reply lands between 1.1s and 13.4s; at 4,000 tokens the same board runs 12s to 194s, with 24 of 31 rows changing places.
  1. Gemini 3.7 Flash soon 3.3s
  2. Nemotron 3.5 Lightning soon 4.6s
  3. GPT 5.4 Nano 4.9s
  4. Gemini 3.6 Flash 5.5s
  5. Muse Spark 1.2 soon 7.0s
  6. GPT 5.4 Mini 7.4s
  7. Muse Spark 1.1 7.5s
  8. GPT 5.6 Luna 8.1s
  9. GPT 5.3 Codex 9.1s
  10. DeepSeek V4 Flash soon 9.1s
  11. Claude Haiku 4.5 9.5s
  12. MiniMax-M3 soon 9.7s
  13. Muse Glimmer soon 10.1s
  14. Gemini 3.1 Pro Preview 10.1s
  15. GPT 5.6 Terra 10.5s
  16. Inkling Small soon 10.9s
  17. Claude Sonnet 5 13.3s
  18. GLM-5.3 soon 13.5s
  19. GPT 5.6 Sol 15.7s
  20. GLM 5.2 16.3s
  21. DeepSeek V4 Pro 17.3s
  22. Grok 4.6 17.6s
  23. Claude Fable 5 17.9s
  24. Grok 4.5 19.3s
  25. Qwen3.7 Plus 19.4s
  26. Claude Opus 5 21.4s
  27. GLM-5.3-Flash soon 21.5s
  28. Inkling soon 22.9s
  29. Kimi K3 29.0s
  30. Qwen3.8 2.4T A95B soon 45.8s
  31. Qwen3.8-Max soon 51.5s

Wait for the first token Streaming the answer

The spread is 6-fold

GPT 5.4 Nano finishes that answer in about 5 seconds and Kimi K3 takes about 29. Same answer length, same request. Nobody waits 29 seconds for a chat reply, and both of those models are on this board.

Short answers change what to rank on

On a 200-token reply the wait is 25% to 66% of the wall clock, so latency is worth weighting. By a few thousand tokens it is a rounding error and throughput decides alone. The control above the chart switches between the two views.

Reasoning is charged in seconds too

The models that think longest before answering are the ones with the worst latency, because the thinking happens before the first token appears. Turning reasoning effort down shortens the wait and the bill at the same time.

Throughput is a batch-job number

Nobody watches a nightly summarisation run. If no human is waiting, optimise tokens per second and ignore latency entirely — and check whether the provider has a batch endpoint, which often halves the bill for work that is not time-sensitive.

Your numbers will differ

Region, prompt length, cache hit rate and reasoning effort all move these figures. Use them to rank candidates, then measure the two you shortlist against your own traffic.

03 · Cost

The price page stopped predicting the bill.

A price per million tokens is a unit price. It says what one unit costs and nothing about how many units the job will take — and that second number now varies between models far more than the first one does.

Models now decide for themselves how long to think, and you pay for that hidden reasoning exactly as for the visible answer. A model five times cheaper per token can land within pennies of one that isn't. Rank on cost per task instead.

It is somebody else's task, at each model's top effort, priced at the provider's own rate — so treat it as a way to rank candidates, not a forecast of your bill.

How much of the discount survives the invoice. Each model against GPT 5.6 Sol: what its published output rate promises, then what its measured cost per task delivers — both at top effort. These 5 keep the least.

Both figures come from the same rate card — each provider's own published price, which is what these tasks were costed at, and not always the rate in the Price column above. And the lesson is not that cheap models are a lie: Gemini 3.1 Pro Preview beats its own rate card — 1.7× cheaper on paper, 3.1× cheaper per task. It is that you cannot rank them by reading price pages.

04 · Method

How to actually decide, in an afternoon.

No leaderboard can rank models on your work. This is the shortest thing that can.

  1. Shortlist two or three from the board

    Use the task router above, or filter the board to your constraint — open weights, a modality, a latency ceiling. Stop at three.

  2. Collect ten real tasks

    With the messy context attached. Not clean samples: the actual ones you ran last week, including the awkward ones.

  3. Run them at the effort you would really use

    Not maximum. The reasoning dial moves cost further than switching models does, and every published cost-per-task figure is measured at the top setting.

  4. Read the token counts off the response

    Every provider returns a usage object. Multiply input, cached input and output separately — those three rates can differ by an order of magnitude.

  5. Divide by the results you would actually ship

    If eight of ten came back usable, your real cost is the total divided by eight. That division is the one no leaderboard can do for you.

One API key reaches every model here, so a bake-off is a string change rather than three integrations. There's a runnable version, an afternoon costing recipe, a guide to running evals, and a free working session if you'd rather do it with an engineer.

FAQ

The questions that decide it.

Which model should I use?
It depends on the job, and the board names a pick for each. Agentic coding: GPT 5.6 Sol or Claude Opus 5 — Sol tops the coding index and Opus tops the intelligence index, and long agentic runs want both. The hardest reasoning: Claude Opus 5, with Kimi K3 as the value option — three index points behind for about a third of the cost per task. GLM-5.3 is cheaper again but is not in the catalog yet. Anything a person waits on: the Gemini Flash line, or GPT 5.6 Luna. Documents, images, audio and video: the Gemini Flash line again. High volume on a budget: GPT 5.6 Luna, at the board's lowest cost per task, $0.05 — matched only by an estimate. Or GLM-5.3-Flash, five index points higher and still under a tenth of a dollar a task, with MIT-licensed weights — though it is not in the MindsHub catalog yet, and not for anything a person is waiting on, because it is one of the slower models on the board. Weights you can host: Kimi K3. Starting cheap rather than choosing on capability: the mindshub_air alias is the low-cost entry point, kept pointed at the best-scoring model at the cheap end of the rate card, with 5M tokens a month at no charge — not a frontier model, and not meant to be.
What does "coming soon" mean on this page?
That the model is benchmarked on the board but is not in the MindsHub catalog yet, so you cannot call it on the API today. It is one small marker on an otherwise identical row, and it clears itself: the board reads the live catalog on every request and binds a row to a model by name, so the moment we start serving one the marker goes, the API alias appears and the price switches to ours — with no change to the page. Until then those rows show the provider's own published list price rather than ours; every model we already serve takes its rates live from our catalog, the same source the pricing page reads.
Why rank on cost per task instead of price per token?
Because a unit price says nothing about how many units the job will take, and that second number now varies between models far more than the first: models decide for themselves how long to think and how many turns to take, and hidden reasoning tokens are billed like visible ones. Grok 4.6 is five times cheaper than GPT 5.6 Sol per output token and costs about 20% more per finished task. The task being measured is one piece of work from Artificial Analysis's suite — a mix of agentic runs, terminal coding, graduate science questions and tool-calling exercises — so treat it as a way to rank candidates rather than as a forecast of your bill. Both price columns on the board are pure token prices with nothing folded in; the platform fee is a separate line on the bill, and the pricing page states it in full.
Where do these benchmark numbers come from?
The page says, column by column, where each figure comes from. The quality index and the cost per task are Artificial Analysis's published measurements, both taken at each model's top reasoning-effort setting, checked on the date shown at the top of the page; a "≈" marks a row they have not covered at that exact model version, where the figure is our own estimate. The coding index and the first-token latency are our own read on every row; the speed column is Artificial Analysis's published median tokens per second wherever they publish one for that exact version, and our read on the four rows they publish no median for. Close values in those three columns are ties rather than measurements. Prices for models we serve are read live from our own catalog.
At what reasoning effort were these numbers measured?
At each model's top setting. Artificial Analysis publishes a separate score and a separate cost for every effort level a model exposes, and labs submit at maximum because that maximises the score — so both published columns on this board are ceilings, and the cost is a worst case. Expand a row to see the levels a model accepts and the one MindsHub sends by default; on most models the default sits below the measured setting, so real bills come in lower. How much lower depends on the model: across Claude Opus 5's five settings, token consumption on the same suite spans roughly eight-fold at one unchanged price per token. We do not publish a per-setting score of our own, because we have not measured one.
Does this page replace the pricing page?
No. The board carries two price dimensions — input and output per million tokens — because that is what a ranking needs, and both are pure token prices with nothing folded into them. Cached input, cache writes, long-prompt surcharge tiers, web search and failover order are all on the pricing page, which is the only place those numbers are authoritative — as is the 5% platform fee, which is a separate line on the bill rather than part of any rate.