Field reference · August 2026
Twenty-five models from thirteen vendors, priced and labeled. No single model wins everything, and the right pick is usually decided by budget and constraints before it's decided by benchmark. Read down for the shape of the market, or jump to the end to work out your own setup.
Why cost comes first
What each model charges to produce a million words' worth of output, on a log scale so the whole market fits on one line. The spread is roughly 380x from cheapest to most expensive, which is why "just use the best one" is advice for people with a budget line rather than a rule.
Dollars per million output tokens. Colors follow the five tiers below. Scroll sideways on a narrow window. This plots a representative eleven spanning the full range rather than all twenty-five; the complete list with prices is in the reference below.
Who leads
Scores are the Artificial Analysis composite index, which measures intelligence rather than value. Gemini 3 Pro (Google) leads multimodal work rather than the composite. Grok 4.5 (xAI) is the cheapest model in the top 10 at $2 per million output tokens. And the open-weight field sits close behind the podium: GLM-5.2 beats GPT-5.5 on SWE-bench Pro. The podium is real, but most working decisions happen below it.
The shape of the market
Nearly every model on the market falls into one of these five groups. Knowing the group tells you the tradeoff you are making.
Claude Fable 5 / Opus 5 · GPT-5.6 Sol · Gemini 3 Pro · Grok 4.5
Deep reasoning, long-running agents, the hardest problems. Correctness matters more than cost here, which makes it a tier for funded work rather than routine volume. Grok 4.5 is the budget door in.
Claude Sonnet 5 · GPT-5.6 standard · Gemini 3.5 Flash · Grok 4.5 · Perplexity Sonar Pro · Nemotron 3 Ultra
Roughly 85 to 95 percent of frontier quality at a third the price or less. The sane default for production work across every vendor. Perplexity's Sonar models sit here with a twist: live web search and citations are bundled into the token price, so they answer from the current web rather than from memory.
Qwen 3.7 Flash · DeepSeek V4-Flash · Nemotron 3 Nano · Nemotron 3 Super · Claude Haiku 4.5 · Gemini Flash-Lite
Sorting, tagging, extracting, routing. Anything where volume is the whole problem and a large model is wasted money.
GLM-5.2 · DeepSeek V4 Pro · Qwen 3 235B · MiniMax M3 · Kimi K2.6 · Mistral Large 3 · Llama 4 · Nemotron 3 (NVIDIA)
Models you can download and run yourself. The fastest-moving part of the market, and where most real volume goes. Each of these has a genuine specialty rather than being a cheap copy of a frontier model.
Image and video generation · embeddings · computer-use models
A different axis entirely. Ranked per task rather than on intelligence indexes, and out of scope for this map.
Every model, expandable
Grouped by whether you can download the weights, because that decides more than price does: it sets whether you can self-host, fine-tune, run disconnected from the internet, or get cut off by a vendor policy change. Click any row for full specs.
The same job, twelve ways
One heavy month on each model: roughly 50 million words in and 5 million out, with no caching discounts. This is the honest version of why almost nobody runs a frontier model for everything.
Quick answers
If you already know the job, this is the short version. Two columns because the answer genuinely differs depending on whether cost is a constraint.
| Job | If budget allows | Bootstrapped pick |
|---|---|---|
| Hardest reasoning, research, overnight agents | Claude Fable 5 / Opus 5, GPT-5.6 Sol | GLM-5.2 or DeepSeek V4 Pro (~6-10x cheaper) |
| Questions about current events, with sources | Sonar Reasoning Pro or Sonar Deep Research | Sonar ($1/$1, search included) |
| Writing and fixing code | Claude Opus 5 or GPT-5.6 Sol | GLM-5.2 (62.1% SWE-bench Pro), MiniMax M3 |
| Math-heavy or scientific reasoning | Claude Fable 5, GPT-5.6 Sol | Qwen 3 235B (77.2% GPQA), DeepSeek V4 Pro |
| Everyday app or product work | Claude Sonnet 5, GPT-5.6 standard | Nemotron 3 Super, Gemini 3.5 Flash, Grok 4.5 |
| Images, video, dense documents | Gemini 3 Pro | MiniMax M3, Kimi K2.6, Gemini 3.5 Flash |
| Writing and editing prose | Claude Opus tier | Claude Sonnet 5, Kimi K2.6 |
| Many languages, or EU data residency | Mistral Large 3 | Qwen 3 235B |
| Running privately on your own machine | Qwen 3.6 (27B dense) or Nemotron 3 Nano | Llama 4 Scout, DeepSeek distills |
| Sorting or tagging huge amounts of text | Claude Haiku 4.5, Gemini Flash-Lite | Qwen 3.7 Flash ($0.03/$0.13, cheapest) |
| Bulk coding volume through OpenRouter | GLM-5.2, MiniMax M3 | Tencent Hy3 (free tier) |
The pattern that works
Almost nobody should run one model for everything. The cost curve rewards mixing tiers, and setting that up is cheap.
Get the task working where capability is not the bottleneck, so you learn what the job actually requires before you optimize.
Re-run the same prompts on a small or open model and compare the results. Most production work holds up, and the bill drops by an order of magnitude.
Send the requests that measurably fail on the cheap model back to the expensive one. That subset is usually far smaller than expected.
Caching repeated context and using batch pricing cuts large-model costs substantially. It changes the absolute numbers, never the ordering.
If your company forbids data retention
Some frontier models require the vendor to keep your data for a period and are unavailable to companies that forbid it. Claude Fable 5 is one. Open-weight models sidestep the question entirely when you host them yourself. The tool below has a filter for this.
One caveat on that filter: Fable 5 is the only model here confirmed to require retention, so the others pass by default. Confirm the current terms with any vendor before you rely on it contractually.
Now make it yours
Answer as few or as many of these as you like. The list underneath narrows to the models that actually qualify, cheapest first, with a real monthly estimate at your own usage.
Pick the one that fits best. Leave it blank to see everything.
Only tick these if they are genuine requirements. Each one removes models from the list.
A ceiling on what a model charges to produce a million words of output.
Used only to estimate your monthly bill. Models charge separately for text you send in and text they write back, measured in tokens: roughly 750 words per thousand tokens.