The Best LLMs in 2026: A Plain-English Comparison

The model at the top of the leaderboard may be too expensive for everyday work. The cheapest API may need more attempts to get the answer right.

That leaves a practical question: which models deserve a test on your work? A team reviewing contracts needs different evidence from one building a coding agent or processing thousands of images.

This guide compares the main model families and the costs behind their headline prices. The September 2026 snapshot includes the newly released GPT-6 Sol and Luna, alongside Claude Opus 5.5, Grok 4.7, and MiMo-V2.6-Pro. It also explains where the free MindsHub Air model and Jev fit. Scores use Artificial Analysis’s Intelligence Index v4.3.2; they should not be compared with scores from earlier index versions.

Where to start testing

  • Difficult reasoning and coding: compare Claude Opus 5.5, GPT-6 Astra, and Claude Fable 5.1. They lead this snapshot, but their task costs differ substantially.
  • Everyday work: test GPT-6 Sol for coding and multi-step workflows, and GPT-6 Luna for focused, high-volume tasks. Compare them at the effort level you would use in production.
  • Mixed-media input: Gemini 3.8 Flash accepts text, images, audio, and video. Muse Spark 1.3 and Qwen3.8-Flash-Next are also candidates when you need image or video input.
  • Lower-cost text and image work: put MiMo-V2.6-Pro, DeepSeek V4.1 Flash, and GLM-5.3-Flash on the shortlist. All three have downloadable, MIT-licensed weights.
  • More control over deployment: compare GLM-5.3 and Kimi K3 as well. Check their licences and hardware requirements before choosing to host them yourself.

These are starting points. Run the same tasks through two or three candidates and inspect the results before committing.

Models at a glance

Prices are US dollars per million uncached input and output tokens. The benchmark setting appears beneath each model name. We use the highest published reasoning effort, including Grok’s xhigh row. When a source publishes one reasoning row without an effort label, we say so. GPT-6 Sol and Luna now also have published task costs, checked September 23.

Cost per task is the benchmark’s weighted average token cost. It is neither a spending limit nor the cost of a result your team has accepted. The suggestions in the third column are our assessment of what to test, based on the capabilities and measurements discussed below.

ModelTypeTest forQualityintelligence index v4.3.2Priceper 1M · in / outPer taskbenchmark average
GPT-6 AstraOpenAI · maxClosedComplex work across code, browsers, and documents53$10 / $50$3.26
GPT-6 SolOpenAI · maxClosedCoding and multi-step agent work48$2 / $10$1.06
GPT-6 LunaOpenAI · maxClosedFocused text and image tasks at volume37$0.10 / $0.50$0.07
GPT-5.6 SolOpenAI · maxClosedGeneral coding, analysis, and professional work47$4 / $20$1.99
GPT-5.6 TerraOpenAI · maxClosedAn upgrade reference for existing GPT-5.6 workflows42$2 / $12$1.40
GPT-5.6 LunaOpenAI · maxClosedRoutine tasks at high volume37$0.2 / $1.2$0.18
Claude Opus 5.5Anthropic · max, with fallbackClosedDifficult reasoning, coding, and document work58$4 / $20$5.98
Claude Fable 5.1Anthropic · max, with fallbackClosedDifficult reasoning, coding, and review53$10 / $50$7.63
Claude Opus 5Anthropic · maxClosedComplex coding and analysis below Fable token rates51$5 / $25$5.86
Claude Fable 5Anthropic · max, with fallbackClosedAn older reference for existing Fable workloads50$10 / $50$8.75
Claude Sonnet 5Anthropic · maxClosedShorter Claude workflows, with effort tuned to the task38$2 / $10$5.09
Gemini 3.8 FlashGoogle · highClosedText, image, audio, and video input41$0.75 / $3.75$1.24
Muse Spark 1.3Meta · max, Standard pricingClosedLong agent workflows with text, images, and video48$1.25 / $4.25$1.60
Grok 4.7SpaceXAI · xhighClosedLong coding and research workflows with web and X tools46$2 / $6$3.74
Grok 4.6SpaceXAI · xhighClosedWork that uses web and X search tools44$2 / $6$2.32
GLM-5.3Z AI · maxOpenText and coding work with downloadable weights45$1.4 / $4.4$2.01
GLM-5.3-FlashZ AI · published reasoning rowOpenLow-cost text and image tasks42$0.15 / $0.5$0.25
Kimi K3Moonshot AI · maxOpenLong coding workflows and image input44$3 / $15$2.00
MiMo-V2.6-ProXiaomi · published reasoning rowOpenLow-cost text, image, audio, and video tasks46$0.435 / $0.87$0.13
Qwen3.8-Max (0902)Alibaba · published reasoning rowClosedMultimodal work through a hosted API45$2 / $6$5.41
Qwen3.8-Flash-NextAlibaba · published reasoning rowOpenLower-cost text, image, and video tasks40$0.15 / $0.47$0.37
DeepSeek V4.1 FlashDeepSeek · max, peak token ratesOpenLow-cost text and image work with MIT-licensed weights39$0.3 / $1.2$0.27
DeepSeek V4 Pro 0813DeepSeek · max, peak token ratesOpenExisting V4 Pro text and reasoning workflows36$1.32 / $3.96$0.67

Pricing and access details

The token rates shown are those recorded by Artificial Analysis, checked September 22, 2026. They are developer API rates; consumer subscriptions and MindsHub’s own rate card are separate.

  • OpenAI and Grok: long prompts can cost more. Check the provider’s current context thresholds, caching terms, and tool fees.
  • Gemini: Google calls the 3.8 Flash rate introductory pricing through December 31, 2026.
  • Muse: the table uses Standard pricing. Contributor has different data-use terms and is discussed below.
  • DeepSeek: the V4.1 Flash row uses peak rates of $0.30 input / $1.20 output. Its off-peak rates are $0.15 / $0.60, and cached input has a separate rate. V4 Pro also has peak and off-peak pricing. Check the current rate card for the applicable hours and rates.
  • Open weights: download access does not promise a particular licence, hosting price, or fit on your hardware. Review the exact checkpoint and host.

This article covers a wider set of models than MindsHub serves. Use the MindsHub model board to check availability and pricing for MindsHub billing.

What each model family offers

OpenAI: Astra, Sol, and Luna

GPT-6 Astra is OpenAI’s model for its hardest work across code, browsers, and professional files. At max effort it scores 53 in this snapshot, alongside Fable 5.1, at $3.26 per benchmark task. That makes it a useful candidate for complicated workflows with several tools and review steps.

GPT-6 Sol and Luna extend that family at lower token prices. Sol lists $2 input / $10 output per million tokens, half GPT-5.6 Sol’s rates. Luna lists $0.10 / $0.50, down from $0.20 / $1.20. Both accept text and image input. Sol is a candidate for coding and multi-step agent work; Luna is worth testing for focused tasks at volume.

At max effort, Sol scores 48 at $1.06 per benchmark task, while Luna scores 37 at $0.07. GPT-5.6 Sol scores 47 at $1.99, and GPT-5.6 Luna scores 37 at $0.18. Those older rows remain as upgrade references. The newer models cost less on this benchmark; test whether that advantage holds for results your team accepts.

Effort is part of the choice. Astra’s xhigh row scores 52 and costs $2.31 per task, against 53 and $3.26 at max. Compare the settings on your own evaluation set before paying for the extra reasoning.

Claude: strong results with substantial reasoning costs

Claude Opus 5.5 leads this snapshot at 58 and $5.98 per benchmark task at max effort with the default fallback. Its high-effort row scores 54 at $1.82. Those measurements make effort a practical part of the choice, especially for recurring code and document work.

Fable 5.1 remains at 53 and $7.63, while Opus 5 scores 51 at $5.86. Sonnet 5 has lower token rates, but its max-effort task cost is still $5.09. We retain the older rows for existing users comparing an upgrade. Opus 5.5 and Fable measurements include a default fallback; Artificial Analysis explains the Fable setup. Treat each result as a measurement of the named configuration.

Gemini: one model for several input types

Gemini 3.8 Flash accepts text, images, audio, and video. That range is useful when a workflow combines recordings, screenshots, and documents. Its high-effort row scores 41 at $1.24 per benchmark task.

The broad score is below the leaders, so check the tasks that matter to you. Google also notes that difficult work can use more tokens and tool calls than with 3.7 Flash. Include those calls in the evaluation, along with the introductory pricing deadline.

Muse: long agent workflows, with a separate Contributor option

Muse Spark 1.3 supports text, image, and video input. At max effort, its score of 48 and Standard task cost of $1.60 make it worth comparing for long coding and agent workflows. It is a closed model; Meta’s history with Llama does not mean Spark weights are downloadable.

Meta’s Contributor option charges $0.10 per million input tokens and $0.20 per million output tokens in exchange for permission to use submitted prompts and completions in training. Treat the data terms as part of the price. The benchmark figure here uses Standard and does not establish a Contributor task cost.

Grok: search access and an effort setting to test

Grok 4.7 adds a newer candidate for long coding and knowledge-work tasks. It accepts text and images, with a 500k context window and access to the provider’s web and X search tools. Search can supply current information; the model still needs to check and use it correctly.

At xhigh, Grok 4.7 scores 46 and costs $3.74 per benchmark task. High also rounds to 46 at $2.73. Grok 4.6 remains at 44 and $2.32 at xhigh. Both generations list $2 input / $6 output per million tokens, yet their task costs differ. The table uses the standard 4.7 API, not its separately priced fast variant.

GLM and Kimi: downloadable models with different costs

GLM-5.3 is a text model with a score of 45 and a task cost of $2.01. GLM-5.3-Flash adds image input, uses an MIT licence, and reaches 42 at $0.25 per task. The larger model’s custom licence and higher cost deserve separate consideration.

Kimi K3 supports image input and long coding workflows. Its max-effort score is 44 at $2.00 per task. Downloadable weights give you a hosting option, but its custom licence and hardware needs still determine whether that option works for your team.

MiMo: a new low-cost multimodal candidate

MiMo-V2.6-Pro scores 46 at $0.13 per benchmark task, using the API rates recorded by Artificial Analysis: $0.435 input / $0.87 output per million tokens. It is worth testing when cost matters, though one broad score cannot show whether it handles your tasks reliably.

Xiaomi’s Pro-RL model card lists text, image, audio, and video input, a one-million-token context window, and MIT-licensed weights. With 1.02 trillion total parameters, it needs substantial hardware to self-host.

Qwen: distinguish Max from Flash-Next

Qwen3.8-Max (0902) is a proprietary multimodal API. Qwen3.8-Flash-Next has open weights, text, image, and video input, and a 256k context window.

The current 0902 Max release scores 45 at $5.41 per task; Flash-Next scores 40 at $0.37. The older 0803 Max release has different measurements, so record the release as well as the model name. Context windows, supported input types, and task behavior also differ. The downloadable Qwen3.8 2.4T A95B checkpoint is another distinct model; it is not the Max API.

DeepSeek: V4.1 Flash replaces the older Flash API

DeepSeek V4.1 Flash launched on September 10 with native text and image input, a one-million-token context window, and MIT-licensed weights. Artificial Analysis measures it at 39 and $0.27 per task at max effort. That puts it close in task cost to GLM-5.3-Flash, with a different capability profile to test.

DeepSeek’s current API documentation names the new model deepseek-flash. The older V4 Flash and Flash Vision API names now route to V4.1 Flash. V4 Pro remains available: the current documentation withdraws the earlier plan to replace it after September 14 and says its billing will stay unchanged. Check the model behind an alias when comparing older results.

The name Flash does not mean it will run on a laptop. Its model card lists 552 billion backbone parameters plus a 196-billion-parameter Engram memory component; self-hosting needs a suitable serving setup and hardware budget.

MindsHub free tier

The MindsHub free tier includes selected models at no charge. Currently included: the MindsHub Air model and Jev. No credit card is required; fair use and daily rate limits apply.

Make the final choice on your own work

Choose representative tasks and agree on what a usable result looks like. Check correctness, required input types, tool use, response time, and data-handling terms. Then divide the total spend, including failed attempts, by the number of results you would accept.

Our cost-per-task guide explains that calculation. The model selection guide covers the evaluation and hosting decisions. If you are changing the workspace as well as the model, use the workspace migration checklist to preserve instructions, connections, and reusable work.


MindsHub Agents provides a workspace with a choice of models and Anton, an open-source agent harness. MindsHub Inference gives developers access to the MindsHub model catalog through one API. Check model availability and MindsHub pricing for the service you plan to use.