Model rankings

Models ranked by quality, speed, and cost.

Compare the live MindsHub catalog by capability, response time, and measured cost per task. For models outside the catalog, see the blog.

Use cases

What are you building?

Start with the kind of work you need to run.

  • Agentic coding

    Many steps, many files, and the model has to keep its own plan straight for an hour. A one-shot index only half predicts that, so run your own repo through the top two.

    On a budget
    DeepSeek V4-Pro-0813 $0.27/task at top effort
  • The genuinely hard problems

    Research, analysis and anything where being wrong is expensive.

    On a budget
    Kimi K3 $0.84/task at top effort
  • Anything a person waits on

    Chat, voice, autocomplete. The first token has to land before attention does.

  • Documents, images, audio, video

    One request carrying a whole corpus of mixed media, reasoned over as a unit.

  • High volume, tight budget

    Classification, extraction, tagging — millions of calls where the bill is the constraint.

  • Weights you can run yourself

    Data that cannot leave your environment, or volume that beats renting.

Not sure which? Start with mindshub_air. We keep this alias pointed at a balanced default, so its underlying model can change without a code change. The first 5M tokens each month are included. No web search and no reasoning dial. Its board row is a snapshot and may briefly lag an alias change.

Leaderboard

Compare the full catalog.

Sort and filter the table, then expand a row for context, modalities, reasoning settings, and billing notes.

The models MindsHub serves, ranked by quality index, with a coding read, output speed, latency, measured cost per task at top reasoning effort, price per million tokens, context window and weight availability.
Anthropic opus Frontier reasoning built for long, many-step agentic runs 63 Frontier 72.4 54 tok/sec 3.2s first token $2.34 $5 / $25 our rate 1M Closed
Pick it when
The task is hard, runs for many steps, or the output carries risk that someone has to sign off on.
Skip it when
You're paying by the token for routine work — it keeps chewing long after a cheaper model would have stopped.
Inputs
TextImages
API alias
opus The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
The top of its dial is a plateau, not a peak. This row quotes max; one notch down at xhigh Artificial Analysis measures the same 63 for $1.80 a task instead of $2.34. Read the score as a ceiling and the cost as a worst case.
Anthropic fable The previous Fable, for traffic already validated against it 62 Frontier 75 67 tok/sec 3.6s first token $3.14 $10 / $50 our rate 1M Closed
Pick it when
You have production traffic validated against this version and re-testing it on 5.1 is not yet worth the interruption.
Skip it when
You are starting fresh — Fable 5.1 scores four points higher on the same list price, and at high it matches this row's 62 for $1.43 a task against $3.14.
Inputs
TextImages
API alias
fable The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
OpenAI gpt OpenAI's flagship all-rounder 61 Frontier 75.6 77 tok/sec 2.0s first token $0.95 $5 / $30 our rate 1M Closed
Pick it when
You want one provider that is competent at everything, with the deepest tooling ecosystem around it.
Skip it when
Price is the binding constraint — GPT 5.6 Terra scores four points lower for about half the cost per finished task.
Inputs
TextImages
API alias
gpt The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send medium unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Moonshot AI kimi Long autonomous coding runs you can self-host 60 Near-frontier 60.2 38 tok/sec 3.4s first token $0.84 $3 / $15 our rate 1M Open Kimi K3 licence (commercial use restricted)
Pick it when
You want the top open-weights score — shared with GLM-5.3 — with image input as well as text.
Skip it when
The API bill is the point — GLM is cheaper for similar work.
Inputs
TextImages
API alias
kimi The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
Licence
Kimi K3 licence (commercial use restricted)
Served by
Fireworks AI Moonshot AI makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Worth knowing
The licence is Moonshot's own, not MIT or Apache. Self-hosting and fine-tuning are fine; reselling access at scale needs a read.
SpaceXAI grok Reasoning with access to the live web 60 Near-frontier 60 55 tok/sec 1.2s first token $1.23 $2 / $6 our rate 500K Closed
Pick it when
You need information from this morning in the same call as the reasoning.
Skip it when
Your agents live in a terminal, or your prompts are enormous — the window is half what most of this board offers.
Inputs
TextImages
API alias
grok The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhigh We send high unless you ask for another. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
Its dial stops paying off at the top. This row quotes xhigh, the highest setting Artificial Analysis measures, and one notch down at high the same model scores a point higher for a quarter less per task. What the top setting buys is one token a second.
Alibaba qwen The open checkpoint that matches the Qwen flagship 58 Near-frontier 56 41 tok/sec 4.1s first token $0.81 $2 / $6 our rate 984K Open Qwen licence (commercial use restricted)
Pick it when
You need the weights, and text-only input plus a restrictive licence are acceptable.
Skip it when
You are reselling access, or your inputs are not text.
Inputs
Text
API alias
qwen The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumxhigh We send xhigh unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Licence
Qwen licence (commercial use restricted)
Served by
Fireworks AI Alibaba makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Worth knowing
Artificial Analysis costs it at Alibaba's own $2 / $6 — the same rate card as the proprietary Max — but the weights are downloadable, so your actual bill is whatever your host or your hardware charges.
Meta muse-spark Cheap closed model that powers the free Meta AI 57 Strong 54 117 tok/sec 0.7s first token $0.40 $1.25 / $4.25 our rate 1M Closed
Pick it when
You want a fast, capable, every-input model and Meta is an acceptable processor of your data.
Skip it when
You need weights — the Muse family is closed — or your data cannot go to Meta.
Inputs
TextImagesAudioVideo
API alias
muse-spark The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
Worth knowing
Meta publishes a "contributor" tier at roughly a fifteenth of the standard rate in exchange for training on your prompts and replies. Fine for hobby projects, an easy no for anything confidential.
OpenAI gpt-terra Balanced mid-tier for everyday work 57 Strong 66 108 tok/sec 1.3s first token $0.53 $2 / $12 our rate 1M Closed
Pick it when
You want the sensible default: strong enough for most work, and it keeps almost all of its discount on the invoice.
Skip it when
The last few index points decide the outcome.
Inputs
TextImages
API alias
gpt-terra The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send medium unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Google gemini-flash Fast multimodal work at a promotional price 56 Strong 72 285 tok/sec 0.5s first token $0.40 $0.75 / $3.75 our rate 1M Closed
Pick it when
Inputs are large or mixed-media — PDFs, images, audio, video — and you want them read quickly.
Skip it when
You need frontier-level reasoning, or weights you can host yourself.
Inputs
TextImagesAudioVideo
API alias
gemini-flash The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhigh We send medium unless you ask for another. Quality and Cost/task above were measured at high, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
SpaceXAI grok-4-5 The previous Grok, and now the better-value one 56 Strong 56 48 tok/sec 1.4s first token $0.43 $2 / $6 our rate 500K Closed
Pick it when
You want most of 4.6 for much less per task, or you have production traffic validated against this version.
Skip it when
You need the four index points 4.6 adds.
Inputs
TextImages
API alias
grok-4-5 The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhigh We send high unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Anthropic sonnet Balanced everyday workhorse 55 Strong 70 72 tok/sec 2.4s first token $1.72 $2 / $10 our rate 1M Closed
Pick it when
You want one model that drafts, summarizes, analyzes and codes without hitting a wall.
Skip it when
Volume is the point. It costs more per finished task than its rate card suggests.
Inputs
TextImages
API alias
sonnet The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
Twice as cheap as GPT 5.6 Sol per output token and about 80% more per finished task — a plain case of a rate card not predicting the bill. Its streaming rate moves with effort too: 72 tok/s at the top setting these figures come from, against 61 at the high setting we send by default.
Meta muse-spark-1-1 Every input type, at an open-model rate 54 Solid 52 150 tok/sec 0.8s first token $0.44 $1.25 / $4.25 our rate 1M Closed
Pick it when
You want every input type at a low rate and a fast first token.
Skip it when
1.2 is three index points better on the same rate card.
Inputs
TextImagesAudioVideo
API alias
muse-spark-1-1 The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
OpenAI gpt-codex Coding specialist for repo-scale changes 54 Solid 68 127 tok/sec 1.7s first token $1.75 / $14 our rate 400K Closed
Pick it when
The job is software, across many files, and you want a model tuned for that rather than for general reasoning.
Skip it when
The work is not code. Its general index score sits below its coding score for a reason.
Inputs
TextImages
API alias
gpt-codex The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhighxhigh We send medium unless you ask for another. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Worth knowing
Artificial Analysis scores this version only provisionally — on an older component set that omits all three of the benchmarks they run on every fully scored row (GDPval-AA v2, Terminal-Bench v2.1 and t3-Banking) — and publishes no cost per task for it, as they do for none of their provisional scores. Their figure is not on the same scale as the rest of this column, so both published columns here are ours.
DeepSeek deepseek Open-weight reasoning and coding at a budget rate 53 Solid 58.6 54 tok/sec 2.6s first token $0.27 $1.32 / $3.96 our rate 1M Open MIT
Pick it when
Volume is high, the budget is tight, and you want permissive weights as an exit route.
Skip it when
You need multimodal input or the last few points of quality.
Inputs
Text
API alias
deepseek The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Licence
MIT
Served by
Fireworks AI DeepSeek makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Version
The alias above now serves DeepSeek V4-Pro-0813, and the scores on this row were measured on DeepSeek V4 Pro — which is why they carry a . Treat them as the closest published read until the newer version is benchmarked.
Zhipu glm The previous GLM release, with the same coding profile 53 Solid 62 68 tok/sec 1.8s first token $0.44 $1.40 / $4.40 our rate 1M Open MIT
Pick it when
You want the GLM coding profile and can spare the points 5.3 adds.
Skip it when
5.3 is seven index points better on the same rate card, and its weights are listed as open too.
Inputs
Text
API alias
glm The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
highmax We send max unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Licence
MIT
Served by
Fireworks AI Zhipu makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
Worth knowing
MIT weights, where 5.3 ships under Zhipu's own licence with commercial restrictions — so if the terms decide it, this is the GLM you can build on without a read.
MindsHub mindshub_air Free allowance Our alias, kept on whichever model balances best 52 Solid 48 128 tok/sec 0.5s first token $0.05 $0.20 / $1.20 our rate 1M Closed
Pick it when
You want a good default without re-picking a model every time the field moves. Point at the alias and we keep it on the model that balances cost, speed and intelligence best.
Skip it when
You are benchmarking, pinning a version, or otherwise need the numbers in this row to hold still. Name a model directly instead.
Inputs
TextImages
API alias
mindshub_air The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. They are mirrored figures, though — see the caveat below before you budget on them.
Worth knowing
Every figure on this row is mirrored from whatever model the alias resolves to today, so it can match another row on this board exactly — and it moves when we move the alias, without notice and with no version to pin. Name a model directly if you need the numbers to stay put.
OpenAI gpt-luna Fast and cheap for high-volume work 52 Solid 48 128 tok/sec 0.5s first token $0.05 $0.20 / $1.20 our rate 1M Closed
Pick it when
Volume is high, tasks are routine, and the bill matters more than the last few points of quality.
Skip it when
The work is genuinely hard. It scores alongside good open models, not alongside the frontier.
Inputs
TextImages
API alias
gpt-luna The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhighmax We send medium unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Google gemini-flash-3-6 Fast, high-volume everyday tasks 52 Solid 68.8 167 tok/sec 0.6s first token $0.34 $0.75 / $3.75 our rate 1M Closed
Pick it when
You want multimodal input and a million-token window at a throughput price.
Skip it when
3.7 Flash scores four index points better and is stronger at code.
Inputs
TextImagesAudioVideo
API alias
gemini-flash-3-6 The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
minimallowmediumhigh We send medium unless you ask for another. Quality and Cost/task above were measured at high, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
OpenAI gpt-mini Small and cheap for routine automation 48 Light 50 150 tok/sec 0.7s first token $0.19 $0.75 / $4.50 our rate 400K Closed
Pick it when
Classification, extraction, tagging and other high-frequency jobs with a narrow definition of correct.
Skip it when
Anything open-ended. Luna is both cheaper and better at this point in the line-up.
Inputs
TextImages
API alias
gpt-mini The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhigh We send none unless you ask for another — the lowest setting it has, and no reasoning tokens at all. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Google gemini Huge documents, images, audio and video 48 Light 68.8 103 tok/sec 1.6s first token $0.33 $2 / $12 our rate 1M Closed
Pick it when
A single request has to carry a whole corpus of mixed media, and it has to be reasoned over rather than skimmed.
Skip it when
You are paying by the token — the Flash line reads the same inputs at a third of the rate card, and 3.7 Flash outscores this by eight points.
Inputs
TextImagesAudioVideo
API alias
gemini The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowmediumhigh We send high unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Alibaba qwen-3-7-plus Multilingual coverage at a fraction of the flagship rate 39 Basic 55 57 tok/sec 1.5s first token $0.17 $0.40 / $1.60 our rate 1M Closed
Pick it when
You want Qwen-family behaviour and broad multilingual coverage without paying a flagship rate.
Skip it when
You need the current flagship — 3.8-Max is nineteen index points ahead.
Inputs
TextImages
API alias
qwen-3-7-plus The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
Served by
Fireworks AI Alibaba makes it; we buy the inference from Fireworks AI, which is the name the rate card shows.
OpenAI gpt-nano Built for throughput on bulk, simple work 38 Basic 30 220 tok/sec 0.4s first token $0.05 $0.20 / $1.25 our rate 400K Closed
Pick it when
Millions of tiny calls where latency is the product and the task fits in one sentence.
Skip it when
Any judgement is required.
Inputs
Text
API alias
gpt-nano The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
nonelowmediumhighxhigh We send none unless you ask for another — the lowest setting it has, and no reasoning tokens at all. Quality and Cost/task above were measured at xhigh, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Anthropic haiku Fastest Claude, for high-volume simple work 30 Basic 43.4 96 tok/sec 1.1s first token $0.22 $1 / $5 our rate 200K Closed
Pick it when
You want Claude behaviour and Claude tooling on work that does not need Claude-sized thinking.
Skip it when
The job needs reasoning headroom — this is the previous generation, and it shows on hard tasks.
Inputs
TextImages
API alias
haiku The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
Single setting — no dial. Nothing to turn down, so the columns above are what you get. That also makes it predictable to budget.
Google gemini-flash-3-5 New in the catalog, not rated here yet $1.50 / $9 our rate Closed
Pick it when
Not rated on this board yet.
Skip it when
Not rated on this board yet.
Inputs
Text
API alias
gemini-flash-3-5 The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
minimallowmediumhigh We send medium unless you ask for another. Quality and Cost/task above were measured at high, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Google gemini-flash-3 New in the catalog, not rated here yet $0.50 / $3 our rate Closed
Pick it when
Not rated on this board yet.
Skip it when
Not rated on this board yet.
Inputs
Text
API alias
gemini-flash-3 The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
minimallowmediumhigh We send high unless you ask for another, and that is this model's top setting — the one Quality and Cost/task above were measured at. Turning it down is the saving on offer here, not a surprise on the invoice.
Google gemini-flash-lite New in the catalog, not rated here yet $0.25 / $1.50 our rate Closed
Pick it when
Not rated on this board yet.
Skip it when
Not rated on this board yet.
Inputs
Text
API alias
gemini-flash-lite The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
minimallowmediumhigh We send minimal unless you ask for another — the lowest setting it has, and no reasoning tokens at all. Quality and Cost/task above were measured at high, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Fireworks AI deepseek-v4-pro New in the catalog, not rated here yet $1.74 / $3.48 our rate Closed
Pick it when
Not rated on this board yet.
Skip it when
Not rated on this board yet.
Inputs
Text
API alias
deepseek-v4-pro The whole integration is this string. It follows model upgrades, so pinning to it survives a version bump.
Reasoning effort
lowhighmax We send high unless you ask for another. Quality and Cost/task above were measured at max, this model's top setting, so this row's board position is a ceiling rather than its out-of-the-box behaviour: at our default you pay less, and on hard work you may score lower.
Quality & Cost/task
Artificial Analysis's published figures, on Intelligence Index v4.1.1 — they measure independently of us — at each model's top reasoning effort, so both are a ceiling. Expand a row for the level we send.
Coding, Speed & Latency
Speed is Artificial Analysis's published median where they publish one; coding and latency are our own read. Close values are ties, not differences.
Price
Per 1M tokens, in / out. Everything else about the bill — prompt-cache rates, long-prompt tiers and search — is on the full rate card.
Missing figures
is our estimate. means nobody has published one.
Checked
. Our rates are live from the catalog.
01 · Quality

Small score differences are ties.

The top 5 scorers sit within 3 points of each other, which is close enough to be noise. When quality is that level, speed and cost decide.

A benchmark averages many tasks and can hide weaknesses that matter to you. Use the scores to shortlist two or three models, then test them on your own work. The method below shows how.

Each score is for a model at a specific reasoning setting, usually the maximum. Treat it as a ceiling. Across Claude Opus 5’s five settings, token use on the same suite spans eight-fold.

Quality index from Artificial Analysis for the full board.
  1. Claude Opus 5 63
  2. Claude Fable 5 62
  3. GPT 5.6 Sol 61
  4. Kimi K3 60
  5. Grok 4.6 60
  6. Qwen3.8-2.4T-A95B 58
  7. Muse Spark 1.2 57
  8. GPT 5.6 Terra 57
  9. Gemini 3.7 Flash 56
  10. Grok 4.5 56
  11. Claude Sonnet 5 55
  12. Muse Spark 1.1 54
  13. GPT 5.3 Codex 54
  14. DeepSeek V4-Pro-0813 53
  15. GLM 5.2 53
  16. MindsHub Air 52
  17. GPT 5.6 Luna 52
  18. Gemini 3.6 Flash 52
  19. GPT 5.4 Mini 48
  20. Gemini 3.1 Pro Preview 48
  21. Qwen3.7 Plus 39
  22. GPT 5.4 Nano 38
  23. Claude Haiku 4.5 30

Bars run from zero and are relative to the leader — a truncated axis would manufacture a gap that isn't there. Scores within 3 points of each other are a tie.

02 · Performance

Measure the wait and the streaming speed.

Latency is the wait for the first token. Throughput is how quickly the rest arrives. The chart shows both because a model can be strong on one and weak on the other.

From the request to a finished answer. Pale is the wait before the first token, solid is the time to stream the rest. A ranking, not a stopwatch: the streaming rate is Artificial Analysis's published median wherever they publish one, the first-token wait is our own read on every row. Answer length changes both the stakes and the order. A 200-token reply lands between 1.2s and 9.0s; at 4,000 tokens the same board runs 15s to 109s, with 17 of 23 rows changing places.
  1. Gemini 3.7 Flash 4.0s
  2. GPT 5.4 Nano 4.9s
  3. Gemini 3.6 Flash 6.6s
  4. GPT 5.4 Mini 7.4s
  5. Muse Spark 1.1 7.5s
  6. MindsHub Air 8.3s
  7. GPT 5.6 Luna 8.3s
  8. Muse Spark 1.2 9.2s
  9. GPT 5.3 Codex 9.6s
  10. GPT 5.6 Terra 10.6s
  11. Gemini 3.1 Pro Preview 11.3s
  12. Claude Haiku 4.5 11.5s
  13. GPT 5.6 Sol 15.0s
  14. Claude Sonnet 5 16.3s
  15. GLM 5.2 16.5s
  16. Claude Fable 5 18.5s
  17. Qwen3.7 Plus 19.0s
  18. Grok 4.6 19.4s
  19. DeepSeek V4-Pro-0813 21.1s
  20. Claude Opus 5 21.7s
  21. Grok 4.5 22.2s
  22. Qwen3.8-2.4T-A95B 28.5s
  23. Kimi K3 29.7s

Wait for the first token Streaming the answer

A 7-fold spread in completion time

Gemini 3.7 Flash finishes that answer in about 4 seconds and Kimi K3 takes about 30. Same answer length, same request. Nobody waits 30 seconds for a chat reply, and both of those models are on this board.

Short answers make latency more important

On a 200-token reply the wait is 24% to 55% of the wall clock, so latency is worth weighting. By a few thousand tokens it is a rounding error and throughput decides alone. The control above the chart switches between the two views.

Measure with your own traffic

Region, prompt length, cache hit rate and reasoning effort all move these figures. Use them to rank candidates, then measure the two you shortlist against your own traffic.

03 · Cost

Price per token is not cost per task.

Token prices do not show how many tokens a model will use to finish the job. Reasoning models can vary widely on that second number.

You pay for reasoning tokens as well as the visible answer. A model with a lower token rate can still cost more per completed task.

It is somebody else's task, at each model's top effort, priced at the provider's own rate — so treat it as a way to rank candidates, not a forecast of your bill.

How much of the discount survives the invoice. Each model against GPT 5.6 Sol: what its published output rate promises, then what its measured cost per task delivers — both at top effort. These 5 keep the least.

Both figures come from the same rate card — each provider's own published price, which is what these tasks were costed at, and not always the rate in the Price column above. And the lesson is not that cheap models are a lie: Gemini 3.1 Pro Preview beats its own rate card — 1.7× cheaper on paper, 2.9× cheaper per task. It is that you cannot rank them by reading price pages.

04 · Method

Choose a model in an afternoon.

Use the board for a shortlist, then test it against real examples.

  1. Shortlist two or three from the board

    Use the task router above, or filter the board to your constraint — open weights, a modality, a latency ceiling. Stop at three.

  2. Collect ten real tasks

    With the messy context attached. Not clean samples: the actual ones you ran last week, including the awkward ones.

  3. Run them at the effort you would really use

    Not maximum. The reasoning dial moves cost further than switching models does, and every published cost-per-task figure is measured at the top setting.

  4. Read the token counts off the response

    Every provider returns a usage object. Multiply input, cached input and output separately — those three rates can differ by an order of magnitude.

  5. Divide by the results you would actually ship

    If eight of ten came back usable, your real cost is the total divided by eight. That division is the one no leaderboard can do for you.

One API key reaches every model here, so a bake-off is a string change rather than three integrations. There's a runnable version, an afternoon costing recipe, a guide to running evals, and a free working session if you'd rather do it with an engineer.

FAQ

Questions about the rankings.

Which model should I use?
Use the task recommendations on the board. Current picks include GPT 5.6 Sol or Claude Opus 5 for agentic coding, Claude Opus 5 for hard reasoning, Gemini Flash or GPT 5.6 Luna for interactive work, and Kimi K3 or DeepSeek V4 Pro when you need open weights. Start with MindsHub Air if you do not yet have a strong preference.
Which models does this page cover?
Every model currently in the MindsHub catalog, and only those models. Each row includes the alias used to call it.
Why rank on cost per task instead of price per token?
Models use different numbers of tokens and reasoning steps to finish the same task. Cost per task captures that difference. Use it to compare candidates, not as a forecast of your own bill.
Where do these benchmark numbers come from?
Quality and task cost come from Artificial Analysis at each model's highest reasoning setting. A “≈” marks a MindsHub estimate. Coding scores and first-token latency are MindsHub assessments. Throughput uses Artificial Analysis medians where available, and prices come from the live MindsHub catalog.
At what reasoning effort were these numbers measured?
Artificial Analysis figures use each model's highest setting. MindsHub often sends a lower default, shown when you expand a row, so actual cost may be lower. MindsHub does not publish its own per-setting scores.
Does this page replace the pricing page?
No. The board shows input and output token prices for comparison. Use the pricing page for cache rates, long-prompt tiers, web search, failover order, and billing details.