LLM Cost per Task: Are Cheap Models Really Cheap?
A lower token price looks like an easy saving. It stays a saving only if the model can finish the work without spending the difference on extra reasoning and retries.
The bill depends on the whole attempt: how much context the model reads, how long it reasons, which tools it calls, and whether you have to run the job again. A rate card gives you the price of each token. It cannot tell you how many the job will need.
Take Qwen3.8-Max (0902) and the previous-generation GPT-5.6 Sol. Their output rates are $6 and $20 per million tokens. On Artificial Analysis’s current benchmark, however, Qwen costs $5.41 per task and Sol at max effort costs $1.99. Qwen’s output rate is 70% lower, while its measured task cost is 172% higher.

The point is to measure how a model works through the job. Some low-price models retain a substantial cost advantage. Others consume enough tokens to lose it.
What cost per task measures
Artificial Analysis calculates a weighted average cost across its Intelligence Index evaluations. For each evaluation, it accounts for input, cache reads and writes, reasoning, and answer tokens, divides by the task count, and applies the evaluation’s index weight.
That is more useful than token price alone, but it still measures benchmark attempts. It does not divide spending by the number of results your team would accept, or include your staff’s review time.
For a model shortlist and input capabilities, use our LLM comparison. This article focuses on the bill.
The September 2026 cost snapshot
The table uses Intelligence Index v4.3.2, checked September 22, 2026, and the same models and settings as the comparison article. Rows are sorted by rounded score, then cost. Token prices are US dollars per million uncached input and output tokens.
We show the highest published effort, or the single published reasoning row where no effort is named. GPT-6 Sol and Luna’s task costs were added on September 23. These are reference settings; the largest effort label does not guarantee the best result or establish a maximum bill.
| Model | Priceper 1M · in / out | Intelligenceindex v4.3.2 | Cost per taskbenchmark average |
|---|---|---|---|
| Claude Opus 5.5max, with fallback | $4 / $20 | 58 | $5.98 |
| GPT-6 Astramax | $10 / $50 | 53 | $3.26 |
| Claude Fable 5.1max, with fallback | $10 / $50 | 53 | $7.63 |
| Claude Opus 5max | $5 / $25 | 51 | $5.86 |
| Claude Fable 5max, with fallback | $10 / $50 | 50 | $8.75 |
| GPT-6 Solmax | $2 / $10 | 48 | $1.06 |
| Muse Spark 1.3max, Standard pricing | $1.25 / $4.25 | 48 | $1.60 |
| GPT-5.6 Solmax | $4 / $20 | 47 | $1.99 |
| MiMo-V2.6-Propublished reasoning row | $0.435 / $0.87 | 46 | $0.13 |
| Grok 4.7xhigh | $2 / $6 | 46 | $3.74 |
| GLM-5.3max | $1.4 / $4.4 | 45 | $2.01 |
| Qwen3.8-Max (0902)published reasoning row | $2 / $6 | 45 | $5.41 |
| Kimi K3max | $3 / $15 | 44 | $2.00 |
| Grok 4.6xhigh | $2 / $6 | 44 | $2.32 |
| GLM-5.3-Flashpublished reasoning row | $0.15 / $0.5 | 42 | $0.25 |
| GPT-5.6 Terramax | $2 / $12 | 42 | $1.40 |
| Gemini 3.8 Flashhigh | $0.75 / $3.75 | 41 | $1.24 |
| Qwen3.8-Flash-Nextpublished reasoning row | $0.15 / $0.47 | 40 | $0.37 |
| DeepSeek V4.1 Flashmax, peak token rates | $0.3 / $1.2 | 39 | $0.27 |
| Claude Sonnet 5max | $2 / $10 | 38 | $5.09 |
| GPT-6 Lunamax | $0.10 / $0.50 | 37 | $0.07 |
| GPT-5.6 Lunamax | $0.2 / $1.2 | 37 | $0.18 |
| DeepSeek V4 Pro 0813max, peak token rates | $1.32 / $3.96 | 36 | $0.67 |
The table uses provider rates recorded by Artificial Analysis, including Standard pricing for Muse and peak token rates for DeepSeek. Cached input, time-of-day discounts, batch processing, long prompts, and tool fees can change your bill. Gemini’s rate is introductory through December 31, 2026. Check the provider terms for the requests you will send; MindsHub’s rate card is separate.
GPT-6 Sol and Luna: lower rates and measured task costs
OpenAI’s new Sol and Luna releases cut the price of each token. Sol costs $2 input / $10 output per million tokens, half GPT-5.6 Sol’s rates. Luna costs $0.10 / $0.50, versus $0.20 / $1.20 for GPT-5.6 Luna.
Artificial Analysis now measures Sol at 48 and $1.06 per task, against 47 and $1.99 for GPT-5.6 Sol. Luna scores 37 at $0.07, against the same rounded score and $0.18 for GPT-5.6 Luna. All four use max effort. These are new measurements, not the old task costs multiplied by a price cut. The opening example remains a comparison with the previous-generation Sol.
Where the newer low-cost models fit
MiMo-V2.6-Pro reaches 46 at $0.13 per benchmark task, using the rates recorded by Artificial Analysis. DeepSeek V4.1 Flash reaches 39 at $0.27. GLM-5.3-Flash reaches 42 at $0.25, while Qwen3.8-Flash-Next reaches 40 at $0.37. These are useful candidates for a budget evaluation, especially when their supported input types match the work.
Two details keep the comparison honest. First, a model name can hide a release change: the current Qwen3.8-Max (0902) row scores 45 at $5.41, while the older 0803 row had different measurements. Flash-Next is a separate model, not a lower effort setting. Second, this is a new DeepSeek checkpoint. The older V4 Flash task cost does not describe V4.1 Flash, even when an API alias moves to the new version. V4 Pro remains available separately. See the DeepSeek section of the comparison for the current routing details.
A new version can change the bill without changing token prices
Grok 4.7 and Grok 4.6 both list $2 input / $6 output per million tokens. At xhigh, their task costs are $3.74 and $2.32, respectively. The newer model scores 46 against 44, but costs about 61% more per benchmark task. Its launch announcement describes changes aimed at longer coding and knowledge-work tasks; test whether those changes improve the results you accept.
Opus 5.5 offers a different comparison. Its $4 / $20 token rates are 20% below Opus 5’s $5 / $25, while its max-effort task cost is about 2% higher: $5.98 against $5.86. Its score rises from 51 to 58. The 5.5 result includes the default fallback, so record that configuration with the comparison.
Try a lower effort setting
Grok 4.7 xhigh and high both round to 46, at $3.74 and $2.73 per task. Opus 5.5 max scores 58 at $5.98; high scores 54 at $1.82. GPT-6 Astra max scores 53 at $3.26, against 52 and $2.31 at xhigh.
Those tradeoffs give you a useful experiment: run your tasks at both settings. Keep the extra reasoning where it improves accepted results, and lower it where it only adds time and cost.
Read discounted tiers as contracts
Meta’s Muse Contributor option charges $0.10 per million input tokens and $0.20 per million output tokens in exchange for permission to use prompts and completions in training. That can change which data you are allowed to send. The $1.60 Muse result above uses Standard pricing; it is not a measured Contributor result.
DeepSeek’s peak, off-peak, and cache rates are another reason to record billing conditions with an evaluation. A discount matters only when your actual requests qualify for it.
MindsHub free tier
The MindsHub free tier includes selected models at no charge. Currently included: the MindsHub Air model and Jev. No credit card is required; fair use and daily rate limits apply.
Measure the cost of results you can use
Start with ten representative tasks as a pilot. Include the documents, instructions, and tool access the model will have in practice, along with a few cases you know are difficult.
- Set acceptance criteria before running the models. For extraction, that might mean every required field is correct. For a code change, it may mean the tests pass and a reviewer accepts the patch.
- Run two or three candidates. Record the model version, reasoning setting, prompts, and billing conditions. Repeat tasks where inconsistent results would matter.
- Count the whole workflow. Include failed attempts, retries, tool charges, and any follow-up model calls. Record human review time separately so a cheap but demanding model does not appear free to operate.
- Divide total spend by accepted results. If you spend $12 across all attempts and accept eight results, the cost is $1.50 per accepted result. If none pass, the model has not established a usable cost for that workflow.
Ten tasks can reveal a large mismatch; they cannot establish a dependable success rate. Expand the evaluation before a consequential rollout, using the process in our model selection guide.
Re-run it when the model, alias, prompt, or pricing changes. Keep the evaluation set and scoring rules outside any one agent workspace; the workspace migration checklist covers the context and workflow details to preserve.
MindsHub Agents provides a workspace with a choice of models and Anton, an open-source agent harness. MindsHub Inference gives developers access to the MindsHub model catalog through one API. Check model availability and MindsHub pricing for the service you plan to use.