Model price comparison

Maintained per-million token rates for major providers. Use this table for quick comparisons; verify against provider pages before budgeting.

For engineers choosing a model for a workload or a team. Cost math is the same everywhere: (tokens / 1,000,000) × the listed rate, with input and output billed separately — so your workload's output ratio matters as much as the sticker price.

Model prices verified 2026-07-02 against provider pricing pages.

ProviderModelIn $/MtokOut $/MtokCtx
Alibabaqwen-flash0.050.41000000
Alibabaqwen3-7-max2.57.51000000
Alibabaqwen3-7-plus0.41.61000000
Alibabaqwen3-coder-plus151000000
Amazon (Bedrock)nova-lite0.060.24300000
Amazon (Bedrock)nova-premier2.512.51000000
Amazon (Bedrock)nova-pro0.83.2300000
Anthropicclaude-fable-510501000000
Anthropicclaude-haiku-4-515200000
Anthropicclaude-opus-4-5525200000
Anthropicclaude-opus-4-75251000000
Anthropicclaude-opus-4-85251000000
Anthropicclaude-sonnet-4-5315200000
Anthropicclaude-sonnet-4-63151000000
Anthropicclaude-sonnet-52101000000
Coherecommand-a2.510256000
Coherecommand-r-08-20240.150.6128000
DeepSeekdeepseek-v4-flash0.140.281000000
DeepSeekdeepseek-v4-pro0.4350.871000000
Googlegemini-2-5-flash0.32.5
Googlegemini-2-5-flash-lite0.10.4
Googlegemini-2-5-pro1.2510
Googlegemini-3-1-flash-lite0.251.5
Googlegemini-3-1-pro-preview212
Googlegemini-3-5-flash1.59
Meta (via OpenRouter)llama-4-maverick0.150.61048576
Meta (via OpenRouter)llama-4-scout0.10.310000000
Meta (via Together AI)llama-3-3-70b1.041.04
Mistralcodestral0.30.9
Mistraldevstral-20.42
Mistralmistral-large-30.51.5262144
Mistralmistral-medium-3-51.57.5262144
Mistralmistral-small-40.150.6
Moonshot AIkimi-k2-50.63262144
Moonshot AIkimi-k2-7-code0.954262144
OpenAIgpt-5-3-codex1.7514
OpenAIgpt-5-42.515
OpenAIgpt-5-4-mini0.754.5
OpenAIgpt-5-4-nano0.21.25
OpenAIgpt-5-5530
OpenAIgpt-5-5-pro30180
Zhipu (Z.ai)glm-4-70.62.2
Zhipu (Z.ai)glm-5-21.44.41048576
xAIgrok-4-31.252.51000000
xAIgrok-build-0-112256000

How to read this table

In/Out are dollars per million tokens; Ctx is the context window. Worked example: at $3/M input, 1M input tokens = $3.00. Output rates are typically several times input rates, so a chatty model on a verbose workload can cost more than a pricier model that answers tersely. Rates in this table are maintained and revalidated, but providers change pricing — verify before committing a budget.

FAQ

Why are input and output priced separately?
Providers bill the two lanes at different rates, and workloads skew differently: agentic coding reads far more than it writes, while generation workloads are output-heavy. Comparing on input price alone routinely picks the wrong model.
What about cache pricing?
Not shown here. Providers that discount cache reads (Anthropic prices them at roughly 10% of the input rate per their pricing page) can change the effective cost dramatically for stable-prefix workloads — use the token calculator, which includes cache lanes where we have rates.
Which model is cheapest for me?
Depends on your token mix. Take your real input and output counts (from a usage dashboard or local logs) and run them through the token calculator against each candidate — the ranking often flips between read-heavy and write-heavy workloads.
What are the limitations?
The table shows list rates only: no batch-API discounts, volume tiers, long-context surcharges, or subscription plans. Treat it as a comparison aid, and verify against your provider invoice.
When should I use this vs the full audit?
Use this to pick a model. The full stack audit addresses the other half of the bill — how many tokens your setup sends per request — which usually moves cost more than switching models does.

For your full setup, run the free stack audit — or see the 100-configs report.

Pricing sources

Every rate is verified against the provider's own pricing page (or a pass-through aggregator where noted in the catalog). Verified 2026-07-02.