Measure

Model price comparison

Maintained per-million token rates for major providers. Use this table for quick comparisons; verify against provider pages before budgeting. Each model also has its own page with worked costs and price history.

For engineers choosing a model for a workload or a team. Cost math is the same everywhere: (tokens / 1,000,000) × the listed rate, with input and output billed separately — so your workload's output ratio matters as much as the sticker price.

Model prices verified 2026-07-02 against provider pricing pages.

qwen-flashAlibaba$0.05$0.40—1,000,000YesBulk / cheapest Qwen
nova-liteAmazon (Bedrock)$0.06$0.24—300,000YesAWS-native high-volume
gemini-2-5-flash-liteGoogle$0.10$0.40$0.01—YesCheapest Google tier
llama-4-scoutMeta (via OpenRouter)$0.10$0.30—10,000,000NoOpen-weights, 10M context
deepseek-v4-flashDeepSeek$0.14$0.28$0.001,000,000YesExtreme price-performance
command-r-08-2024Cohere$0.15$0.60—128,000NoLight RAG workloads
llama-4-maverickMeta (via OpenRouter)$0.15$0.60—1,048,576NoOpen-weights, hosted
mistral-small-4Mistral$0.15$0.60——NoCheap general-purpose
gpt-5-4-nanoOpenAI$0.20$1.25$0.02—YesBulk / classification
gemini-3-1-flash-liteGoogle$0.25$1.50$0.03—YesCheap high-volume
gemini-2-5-flashGoogle$0.30$2.50$0.03—YesFast general-purpose
codestralMistral$0.30$0.90——NoCode completion
qwen3-7-plusAlibaba$0.40$1.60—1,000,000YesBalanced Qwen tier
devstral-2Mistral$0.40$2.00——NoCoding agents (Mistral)
deepseek-v4-proDeepSeek$0.43$0.87$0.001,000,000YesFrontier value, 1M context
mistral-large-3Mistral$0.50$1.50—262,144NoEuropean flagship value
kimi-k2-5Moonshot AI$0.60$3.00$0.10262,144YesValue agentic tier
glm-4-7Zhipu (Z.ai)$0.60$2.20$0.11—YesZhipu value tier
gpt-5-4-miniOpenAI$0.75$4.50$0.07—YesHigh-volume subtasks
nova-proAmazon (Bedrock)$0.80$3.20—300,000YesAWS-native balanced tier
kimi-k2-7-codeMoonshot AI$0.95$4.00$0.19262,144YesAgentic coding (K2 line)
claude-haiku-4-5Anthropic$1.00$5.00$0.10200,000YesFast, cheap subtasks
qwen3-coder-plusAlibaba$1.00$5.00—1,000,000YesCoding (Qwen line)
grok-build-0-1xAI$1.00$2.00$0.20256,000YesCoding (xAI)
llama-3-3-70bMeta (via Together AI)$1.04$1.04——NoOpen-weights workhorse
gemini-2-5-proGoogle$1.25$10.00$0.13—YesLong-context analysis
grok-4-3xAI$1.25$2.50$0.201,000,000YesxAI flagship
glm-5-2Zhipu (Z.ai)$1.40$4.40$0.261,048,576YesZhipu flagship
gemini-3-5-flashGoogle$1.50$9.00$0.15—YesGoogle flagship-speed tier
mistral-medium-3-5Mistral$1.50$7.50—262,144NoMistral premium tier
gpt-5-3-codexOpenAI$1.75$14.00$0.17—YesCoding agents
claude-sonnet-5Anthropic$2.00$10.00$0.201,000,000YesCoding default — intro pricing
gemini-3-1-pro-previewGoogle$2.00$12.00$0.20—YesLong-context analysis (preview)
gpt-5-4OpenAI$2.50$15.00$0.25—YesBalanced OpenAI tier
qwen3-7-maxAlibaba$2.50$7.50—1,000,000YesAlibaba flagship
nova-premierAmazon (Bedrock)$2.50$12.50$0.631,000,000YesAWS-native flagship
command-aCohere$2.50$10.00—256,000NoEnterprise RAG / tools
claude-sonnet-4-6Anthropic$3.00$15.00$0.301,000,000YesCoding, 1M context
claude-sonnet-4-5Anthropic$3.00$15.00$0.30200,000YesCoding, analysis, balanced cost
claude-opus-4-8Anthropic$5.00$25.00$0.501,000,000YesDeep coding sessions, agentic work
claude-opus-4-7Anthropic$5.00$25.00$0.501,000,000YesLegacy Opus tier
claude-opus-4-5Anthropic$5.00$25.00$0.50200,000YesLegacy Opus tier
gpt-5-5OpenAI$5.00$30.00$0.50—YesOpenAI flagship
claude-fable-5Anthropic$10.00$50.00$1.001,000,000YesFrontier reasoning, hardest problems
gpt-5-5-proOpenAI$30.00$180.00——NoMax reasoning, price-insensitive

Provider logos via Lobe Icons (MIT). Logos are trademarks of their respective owners, shown to identify the products compared. tokencheat is not affiliated with, endorsed by, or sponsored by any provider listed.

Last updated: 2026-07-02 · Rates are indicative; verify on provider sites before budgeting.

How to read this table

In/Out are dollars per million tokens; Ctx is the context window. Worked example: at $3/M input, 1M input tokens = $3.00. Output rates are typically several times input rates, so a chatty model on a verbose workload can cost more than a pricier model that answers tersely. Rates in this table are maintained and revalidated, but providers change pricing — verify before committing a budget.

FAQ

Why are input and output priced separately?
Providers bill the two lanes at different rates, and workloads skew differently: agentic coding reads far more than it writes, while generation workloads are output-heavy. Comparing on input price alone routinely picks the wrong model.
What about cache pricing?
Not shown here. Providers that discount cache reads (Anthropic prices them at roughly 10% of the input rate per their pricing page) can change the effective cost dramatically for stable-prefix workloads — use the token calculator, which includes cache lanes where we have rates.
Which model is cheapest for me?
Depends on your token mix. Take your real input and output counts (from a usage dashboard or local logs) and run them through the token calculator against each candidate — the ranking often flips between read-heavy and write-heavy workloads.
What are the limitations?
The table shows list rates only: no batch-API discounts, volume tiers, long-context surcharges, or subscription plans. Treat it as a comparison aid, and verify against your provider invoice.
When should I use this vs the full audit?
Use this to pick a model. The full stack audit addresses the other half of the bill — how many tokens your setup sends per request — which usually moves cost more than switching models does.

For your full setup, run the free stack audit — or see the 100-configs report.

Pricing sources

Every rate is verified against the provider's own pricing page (or a pass-through aggregator where noted in the catalog). Verified 2026-07-02.

Next step

Picking a cheaper model is the smaller lever.

Rates differ by a few dollars per million tokens. How much context you send on every turn differs by far more, and that is what the audit measures.

Audit my instructions