Field reference · August 2026

The Model Map

Twenty-five models from thirteen vendors, priced and labeled. No single model wins everything, and the right pick is usually decided by budget and constraints before it's decided by benchmark. Read down for the shape of the market, or jump to the end to work out your own setup.

Why cost comes first

The price range

What each model charges to produce a million words' worth of output, on a log scale so the whole market fits on one line. The spread is roughly 380x from cheapest to most expensive, which is why "just use the best one" is advice for people with a budget line rather than a rule.

$1
$10
Qwen Flash$0.13
DeepSeek Flash$0.28
DeepSeek Pro$0.87
MiniMax M3$1.20
Mistral L3$1.50
Grok 4.5$2
Gemini Flash$4.50
Haiku 4.5$5
Sonnet 5$15
Opus 5$25
Fable 5$50

Dollars per million output tokens. Colors follow the five tiers below. Scroll sideways on a narrow window. This plots a representative eleven spanning the full range rather than all twenty-five; the complete list with prices is in the reference below.

Who leads

The top of the board

1Claude Opus 5Anthropic60.7
2Claude Fable 5Anthropic59.9
3GPT-5.6 SolOpenAI58.9

Scores are the Artificial Analysis composite index, which measures intelligence rather than value. Gemini 3 Pro (Google) leads multimodal work rather than the composite. Grok 4.5 (xAI) is the cheapest model in the top 10 at $2 per million output tokens. And the open-weight field sits close behind the podium: GLM-5.2 beats GPT-5.5 on SWE-bench Pro. The podium is real, but most working decisions happen below it.

The shape of the market

The five tiers

Nearly every model on the market falls into one of these five groups. Knowing the group tells you the tradeoff you are making.

Frontier$2 to $50 / M out

Claude Fable 5 / Opus 5 · GPT-5.6 Sol · Gemini 3 Pro · Grok 4.5

Deep reasoning, long-running agents, the hardest problems. Correctness matters more than cost here, which makes it a tier for funded work rather than routine volume. Grok 4.5 is the budget door in.

Workhorse$0.75 to $15 / M out

Claude Sonnet 5 · GPT-5.6 standard · Gemini 3.5 Flash · Grok 4.5 · Perplexity Sonar Pro · Nemotron 3 Ultra

Roughly 85 to 95 percent of frontier quality at a third the price or less. The sane default for production work across every vendor. Perplexity's Sonar models sit here with a twist: live web search and citations are bundled into the token price, so they answer from the current web rather than from memory.

Small / fast$0.03 to $5 / M out

Qwen 3.7 Flash · DeepSeek V4-Flash · Nemotron 3 Nano · Nemotron 3 Super · Claude Haiku 4.5 · Gemini Flash-Lite

Sorting, tagging, extracting, routing. Anything where volume is the whole problem and a large model is wasted money.

Open-weightself-host or ~6-10x under proprietary

GLM-5.2 · DeepSeek V4 Pro · Qwen 3 235B · MiniMax M3 · Kimi K2.6 · Mistral Large 3 · Llama 4 · Nemotron 3 (NVIDIA)

Models you can download and run yourself. The fastest-moving part of the market, and where most real volume goes. Each of these has a genuine specialty rather than being a cheap copy of a frontier model.

Specialistpriced per task

Image and video generation · embeddings · computer-use models

A different axis entirely. Ranked per task rather than on intelligence indexes, and out of scope for this map.

Every model, expandable

Full reference

Grouped by whether you can download the weights, because that decides more than price does: it sets whether you can self-host, fine-tune, run disconnected from the internet, or get cut off by a vendor policy change. Click any row for full specs.

Closed weights Rented through an API. The vendor sets the terms.
Open weights Downloadable. Self-host, fine-tune, or rent from any provider.

The same job, twelve ways

What it actually costs

One heavy month on each model: roughly 50 million words in and 5 million out, with no caching discounts. This is the honest version of why almost nobody runs a frontier model for everything.

Claude Opus 5Anthropic · $5 / $25
~$375
Claude Sonnet 5Anthropic · $3 / $15
~$225
Gemini 3.5 FlashGoogle · $0.75 / $4.50
~$60
SonarPerplexity · $1 / $1
~$55
Mistral Large 3Mistral · $0.50 / $1.50
~$32
DeepSeek V4 Proopen weight · $0.44 / $0.87
~$26
MiniMax M3open weight · $0.30 / $1.20
~$21
Llama 4 Maverickopen weight · $0.27 / $0.85
~$18
DeepSeek V4-Flashopen weight · $0.14 / $0.28
~$8
Nemotron 3 SuperNVIDIA · $0.085 / $0.40
~$6
Nemotron 3 NanoNVIDIA · $0.06 / $0.24
~$4
Qwen 3.7 FlashAlibaba · $0.03 / $0.13
~$2

Caching and batch discounts cut the large-model numbers substantially, but they never change the ordering. Want these figures at your own volume instead? The tool at the bottom does that.

Quick answers

What to use when

If you already know the job, this is the short version. Two columns because the answer genuinely differs depending on whether cost is a constraint.

JobIf budget allowsBootstrapped pick
Hardest reasoning, research, overnight agentsClaude Fable 5 / Opus 5, GPT-5.6 SolGLM-5.2 or DeepSeek V4 Pro (~6-10x cheaper)
Questions about current events, with sourcesSonar Reasoning Pro or Sonar Deep ResearchSonar ($1/$1, search included)
Writing and fixing codeClaude Opus 5 or GPT-5.6 SolGLM-5.2 (62.1% SWE-bench Pro), MiniMax M3
Math-heavy or scientific reasoningClaude Fable 5, GPT-5.6 SolQwen 3 235B (77.2% GPQA), DeepSeek V4 Pro
Everyday app or product workClaude Sonnet 5, GPT-5.6 standardNemotron 3 Super, Gemini 3.5 Flash, Grok 4.5
Images, video, dense documentsGemini 3 ProMiniMax M3, Kimi K2.6, Gemini 3.5 Flash
Writing and editing proseClaude Opus tierClaude Sonnet 5, Kimi K2.6
Many languages, or EU data residencyMistral Large 3Qwen 3 235B
Running privately on your own machineQwen 3.6 (27B dense) or Nemotron 3 NanoLlama 4 Scout, DeepSeek distills
Sorting or tagging huge amounts of textClaude Haiku 4.5, Gemini Flash-LiteQwen 3.7 Flash ($0.03/$0.13, cheapest)
Bulk coding volume through OpenRouterGLM-5.2, MiniMax M3Tencent Hy3 (free tier)

The pattern that works

How to route

Almost nobody should run one model for everything. The cost curve rewards mixing tiers, and setting that up is cheap.

First

Prototype on a frontier model

Get the task working where capability is not the bottleneck, so you learn what the job actually requires before you optimize.

Then

Move volume to a cheap tier

Re-run the same prompts on a small or open model and compare the results. Most production work holds up, and the bill drops by an order of magnitude.

Keep

Escalate only what earns it

Send the requests that measurably fail on the cheap model back to the expensive one. That subset is usually far smaller than expected.

Always

Cache aggressively

Caching repeated context and using batch pricing cuts large-model costs substantially. It changes the absolute numbers, never the ordering.

If your company forbids data retention

Some frontier models require the vendor to keep your data for a period and are unavailable to companies that forbid it. Claude Fable 5 is one. Open-weight models sidestep the question entirely when you host them yourself. The tool below has a filter for this.

One caveat on that filter: Fable 5 is the only model here confirmed to require retention, so the others pass by default. Confirm the current terms with any vendor before you rely on it contractually.

Now make it yours

Work out your own setup

Answer as few or as many of these as you like. The list underneath narrows to the models that actually qualify, cheapest first, with a real monthly estimate at your own usage.

1. What do you want it to do?

Pick the one that fits best. Leave it blank to see everything.

2. Any rules you cannot break?

Only tick these if they are genuine requirements. Each one removes models from the list.

3. How much do you want to spend?

A ceiling on what a model charges to produce a million words of output.

4. How much will you use it?

Used only to estimate your monthly bill. Models charge separately for text you send in and text they write back, measured in tokens: roughly 750 words per thousand tokens.