LLM Cost per Task: Are Cheap Models Really Cheap?

Here is a question that is easy to get wrong.

Two AI models can do the same piece of work. One of them charges five times more for the text it produces. At the end of the month, which one has cost you more?

Often enough, neither. They land within a few cents of each other.

The two models are real. Alibaba’s Qwen3.8-Max charges $6 per million output tokens and OpenAI’s GPT-5.6 Sol charges $30, both published on their own price pages, and five-to-one is not a subtle difference. Run the same set of hard tasks through both and total what lands on the invoice: Qwen3.8-Max comes to about $1.13 a task, GPT-5.6 Sol to about $1.23.

A five-fold discount, worth ten cents.

Ten cents is not worth writing about until you scale it. An agent grinding through two thousand tasks a month on Sol runs about $2,460. Move that workload to Qwen because the price page promises five-to-one, and you have budgeted $492 for it. The invoice arrives at $2,260.

Dollars per million tokens is the unit the labs publish, so it’s the unit every comparison quotes, ours included. Our rundown of the leading models lines them up on quality, context, licensing and list price, and that’s the right way to draw up a shortlist: it tells you who is in the running and roughly what tier of spend you’re entering.

This is the layer underneath it. A per-token price predicts a bill only while the amount of work per job stays put, and over the past year, as models started thinking for longer and running for dozens of turns at a time, it stopped staying put.

Qwen3.8-Max costs five times less per output token than GPT-5.6 Sol, but only eight percent less per task.

Why the sticker price stopped predicting the bill

A price per million tokens is a unit price. It tells you what one unit costs and nothing at all about how many units the job will take, and that second number now varies between models far more than the first one does.

Start with how long a model thinks. Modern models generate hidden reasoning tokens before they answer and you are billed for those exactly as for the visible ones, so two models can hand back identically short answers and differ tenfold in what it took to get there. Artificial Analysis rates Qwen3.8-Max maximally verbose, describing it as “notably slow and very verbose.” That is not an insult and it isn’t a defect: Qwen is a strong model and the thoroughness is the point. You just pay for the thoroughness by the word. Then multiply that by turns. Most serious AI work now runs as an agent, reading and acting and checking and retrying, and Qwen3.8-Max averages around 64 turns on Artificial Analysis’s agentic tasks where its predecessor averaged 14. Same job, four and a half times the round trips. And then the part that quietly costs the most: on turn 40, the model re-reads turns 1 through 39 as input. Turn count doesn’t just multiply generation, it multiplies your input bill, which is why caching is now priced as its own line item. For the leanest models on the board, generation isn’t even the main cost any more. Grok 4.5’s $0.36 per task is mostly input: the same conversation, re-sent every turn.

A cheap model that thinks twice as long and takes four times the turns is not cheap.

The comparison table

The Cost per task column below is Artificial Analysis’s measurement, taken on August 11, 2026. They run every model through the same suite of hard tasks and total the real bill, counting input, cached, reasoning and answer tokens alike. The last column is our arithmetic on top of theirs.

Rows are sorted by cost per task, most expensive first. Compare that order against the price column.

ModelPriceper 1M · in / outIntelligenceindex scoreCost per taskSticker vs realityvs GPT-5.6 Sol
Claude Fable 5$10 / $5062$3.141.7x dearer per token2.6x dearer per task
Claude Opus 5$5 / $2563$2.341.2x cheaper per token1.9x dearer per task
Claude Sonnet 5$3 / $1555$2.292.0x cheaper per token1.9x dearer per task
Claude Opus 4.8$5 / $2556$1.801.2x cheaper per token1.5x dearer per task
GPT-5.6 Sol$5 / $3061$1.23the yardstick1.0x by definition
Qwen3.8-Max$2 / $658$1.135.0x cheaper per token1.1x cheaper per task
Kimi K3$3 / $1560$0.842.0x cheaper per token1.5x cheaper per task
Gemini 3.6 Flash$1.50 / $7.5052$0.564.0x cheaper per token2.2x cheaper per task
GPT-5.6 Terra$2 / $1257$0.512.5x cheaper per token2.4x cheaper per task
Muse Spark 1.2$1.25 / $4.2554$0.407.1x cheaper per token3.1x cheaper per task
Grok 4.5$2 / $656$0.365.0x cheaper per token3.4x cheaper per task
GLM-5.2$1.40 / $4.4053$0.316.8x cheaper per token4.0x cheaper per task
MiniMax-M3$0.30 / $1.2045$0.1425x cheaper per token8.8x cheaper per task
GPT-5.6 Luna$0.20 / $1.2052$0.0525x cheaper per token25x cheaper per task
DeepSeek V4 Pro$0.44 / $0.8745$0.0434x cheaper per token31x cheaper per task
DeepSeek V4 Flash$0.14 / $0.2852$0.027107x cheaper per token46x cheaper per task

Prices are API list rates as of August 11, 2026. The Intelligence column is Artificial Analysis’s composite quality score on a 0–100 scale, where the current leader sits at 63. The “sticker vs reality” column compares each model’s output-token price and its cost per task against GPT-5.6 Sol, chosen as the yardstick because it’s a widely deployed flagship at the top of the price sheet. Claude Sonnet 5’s figures use its standard $3 / $15 rate rather than the promotional $2 / $10, which lowers it. Always check the provider for the current rate.

We left Gemini 3.1 Pro out. Its published score and its cost figures currently disagree across sources, and we would rather drop a row than print a number we can’t stand behind.

The table disagrees with the price pages

Every model gives up some of its advantage between the price page and the invoice, which is what you would expect: cost per task counts input and cached tokens too, so a model with cheap output relative to input has further to fall. The surprise is the spread. GPT-5.6 Terra keeps nearly all of its discount, 2.5x on the sticker and 2.4x in practice. Qwen3.8-Max keeps almost none of its own, 5.0x becoming 1.1x.

A shrinking discount can still be an enormous discount. DeepSeek V4 Flash sheds more than half its multiple, 107x on paper down to 46x in practice, and remains by a distance the cheapest way to get work done on this board. What you cannot do is rank these models by reading their price pages.

Then there is Anthropic, which inverts the rule outright.

Claude Sonnet 5 costs 40% less than Claude Opus 4.8 per token. It costs more per task: $2.29 against $1.80. That is Artificial Analysis’s finding, not our reading of it.

Three-fifths the price. Higher bill.

Every Anthropic model in the table lands above what its price page predicts, which reads less like an accident than a house style: these are models built to keep chewing on a problem, and on genuinely hard work that is usually what you want. It does mean the most common cost-saving move in the Claude ecosystem, stepping down from Opus to Sonnet, can quietly go the wrong way. If you made that switch this year to save money, go and check whether it did.

OpenAI is the exception. Across a 25-fold price range, from Luna to Sol, its lineup tracks its own price ladder almost exactly.

The gap, drawn

Each line runs from the discount a model’s price page promises (hollow marker) to the discount its measured cost per task delivers (solid marker). The longer the line, the more of the advertised saving evaporates.

Promised by the price page Delivered per task

DeepSeek V4 Flash107x → 46x
DeepSeek V4 Pro34x → 31x
GPT-5.6 Luna25x → 25x
MiniMax-M325x → 8.8x
GLM-5.26.8x → 4.0x
Grok 4.55.0x → 3.4x
Qwen3.8-Max5.0x → 1.1x
Kimi K32.0x → 1.5x
1x3x10x30x100x

Eight of the eleven models that beat GPT-5.6 Sol on both measures, picked to span the range; the table above has all of them. The axis is logarithmic so a 100-fold range fits on one line, so read the numbers rather than the distances.

Five things this number can’t see

Cost per task beats price per token, and it still comes off a benchmark. Five limits are worth carrying around with it.

It prices attempts, not successes. The metric counts what a model spent, never whether the answer was usable. A model that fails fast and cheap scores beautifully. In production you re-run the failures, and often a person reads the output before it goes anywhere, so a model that’s 10% cheaper per attempt and wrong 30% more often is more expensive in every way that matters. No public leaderboard measures cost per finished task, because nobody but you can judge finished.

A “task” here is Artificial Analysis’s task, not yours. The average blends long agentic runs, terminal coding, graduate science questions and tool-calling exercises. Roughly half the index is agentic work, which we think is the right call for a general-purpose number in 2026, because that’s where the money is going. It’s also why your own ranking will come out much flatter than this one if your job is summarizing support tickets.

Caching is modelled, not observed. The calculation applies each model’s typical cache hit rate rather than the rate that actually occurred during the run. Prompts that change on every call cost more than published. A stable prefix costs less.

It’s list-price, text-only, English-only, and blind to tiered rates. Gemini and Grok both roughly double their rates above ~200K-token prompts, and no blended figure can show you how often that line was crossed, so treat the published number as a floor for long-document work. Batch endpoints, meanwhile, can halve the bill for anything that isn’t time-sensitive.

It moves for reasons that have nothing to do with the model. GPT-5.6 Luna’s cost per task fell from $0.21 to $0.05 in July because OpenAI cut its price by 80%. GPT-5.6 Sol’s published figure has read anywhere from $1.04 to $1.23 this summer at an unchanged $5 / $30, because the benchmark suite and the costing method were both revised underneath it. Every figure in the table above is one day’s snapshot and will drift the same way. Any cost-per-task number you quote needs a date attached, or it means nothing. That includes ours.

The dial matters more than the model

Modern models expose a reasoning effort setting, which controls how long the model may think before answering, and it moves cost further than swapping models does. Claude Opus 5 is published at five effort settings, and across them its token consumption on the same benchmark suite spans roughly eight-fold, from 12 million tokens at the lowest setting to 100 million at the highest. Same model, same rate card, same tasks.

This has a quiet consequence for every leaderboard you read. Labs submit at maximum effort, because that is what maximizes the score, so the published cost per task usually reflects the most expensive setting a model has, which is rarely the one you’d run all day. Turning the dial down is the fastest saving available to most teams, and a frontier model turned down will often beat a mid-tier model turned up on quality and cost together.

Work out your own cost per task

Every API already reports what you need for this, and it takes an afternoon.

Collect ten real tasks, with the messy context attached, and run each one through two or three candidates at the reasoning effort you would actually use. Read the token counts off the response instead of estimating them: every provider returns a usage object with input, output and cached counts, and reasoning tokens sit inside the output number. Multiply by the rate card separately for input, cached input and output, because those three prices can differ by an order of magnitude. Then divide by the number of results you would actually ship, not the number you ran. If eight of ten came back usable, your cost per finished task is the total divided by eight, and that division is the one no leaderboard can do for you.

Ten tasks will not survive a statistician, and a ten-task run can absolutely crown a different winner than fifty tasks would. Run ten anyway. A rough number you have beats a precise one you never get around to. If you want to do this properly, with a real evaluation set behind it, our guide to choosing and testing a model covers that side in more depth.

Where this leaves us

None of this makes the price pages dishonest. They answer a question that used to be the same as the one you were asking, and quietly stopped being.

The unit you are billed in has come apart from the unit you care about. You buy tokens, you consume finished jobs, and for years those two moved together closely enough that the price of one implied the price of the other. That link only holds while the amount of work per job is roughly fixed. Once a model decides for itself how long to think and how many turns to take, the amount stops being a property of your job and becomes a property of the model you hired. Qwen3.8-Max and GPT-5.6 Sol sit five-to-one apart on tokens and neck and neck on work.

Which means cheap and efficient are now two different properties, and only one of them is published. What you actually save is the advertised discount multiplied by how sparing the model is, and no vendor prints the second number. GPT-5.6 Luna and MiniMax-M3 advertise the identical 25x discount against Sol; one delivers 25x and the other 8.8x.

The practical damage is that the moves which used to save money have turned into questions. Dropping a tier, moving to an open model, taking the cheaper of two similar scores: all of these were reliable, and none of them are reliable now. Claude Sonnet 5 costs 40% less than Claude Opus 4.8 per token and more per finished task. If you made a switch like that this year on the strength of a price page, the honest thing is to go and look at what happened to the bill.

And the number that would settle any of this doesn’t exist yet. Every public figure, ours on this page included, prices attempts. None of them price successes. A model that fails cheaply looks superb on a leaderboard and costs you a re-run plus somebody’s afternoon, and until a benchmark can judge “good enough” for your particular work, that measurement has to come from you.

So the shift is a small one, and it’s mostly a habit. Compare candidates on a task instead of a price page. Set the effort dial before you go shopping for a cheaper model, because it moves more money than the model choice does. And before any switch made in the name of cost, put ten real jobs through both options and count what you would have shipped.

The honest answer to “which model is cheapest” is: for which job, at which effort setting, measured on which date. It’s an irritating answer, and it’s the only one that survives contact with an invoice.

It’s also why we built MindsHub to treat the model as a setting. MindsHub Cowork is a workspace where you hand a whole task to an open-source AI agent and collect the finished work, and the Model Router underneath it spans the frontier providers and the leading open models, so running the same task through three of them and comparing what each one actually cost is a dropdown. Try it free.

Frequently asked questions

What is cost per task for an LLM? The average total bill to finish one real task, counting every token the model consumed: input, cached input, hidden reasoning, and the visible answer. Artificial Analysis publishes it for its Intelligence Index suite, and it regularly ranks models in a different order than their price pages do.

Does a cheaper model per token ever cost more per task? Yes, and it isn’t rare. Claude Sonnet 5 costs 40% less per token than Claude Opus 4.8 and more per task, $2.29 against $1.80, which Artificial Analysis flagged at launch. Every Anthropic model behaves this way relative to its sticker price, because they’re built to think longer and persist through more steps.

Does reasoning effort change cost more than switching models? Often, yes. Claude Opus 5’s token consumption spans roughly eight-fold across its five published effort settings at one unchanged price per token. Leaderboards typically report the highest setting, so the published figure is usually a worst case, and turning the dial down is the fastest saving most teams have available.

Why do published cost-per-task numbers keep changing? Three reasons that aren’t the model: providers cut prices, benchmark suites get revised, and the costing methodology itself gets updated. GPT-5.6 Luna went from $0.21 to $0.05 per task on a price cut, and GPT-5.6 Sol has been quoted between $1.04 and $1.23 this summer on an unchanged rate card. Always check the date on the number.


MindsHub by MindsDB is the unified workspace where open-source models get things done for you. Delegate entire projects through MindsHub Cowork and collect finished, shareable results - work runs on interchangeable open-source agent harnesses, Anton and Hermes. The Model Router spans commercial and open models, so you can match every job to the right engine and switch anytime. Founded 2018 in Berkeley. Backed by Benchmark, Mayfield, Y Combinator, and NVIDIA.