Measure
Model price comparison
Maintained per-million token rates for major providers. Use this table for quick comparisons; verify against provider pages before budgeting. Each model also has its own page with worked costs and price history.
For engineers choosing a model for a workload or a team. Cost math is the same everywhere: (tokens / 1,000,000) × the listed rate, with input and output billed separately — so your workload's output ratio matters as much as the sticker price.
Model prices verified 2026-07-02 against provider pricing pages.
| qwen-flash | Alibaba | $0.05 | $0.40 | — | 1,000,000 | Yes | Bulk / cheapest Qwen |
| nova-lite | Amazon (Bedrock) | $0.06 | $0.24 | — | 300,000 | Yes | AWS-native high-volume |
| gemini-2-5-flash-lite | $0.10 | $0.40 | $0.01 | — | Yes | Cheapest Google tier | |
| llama-4-scout | Meta (via OpenRouter) | $0.10 | $0.30 | — | 10,000,000 | No | Open-weights, 10M context |
| deepseek-v4-flash | DeepSeek | $0.14 | $0.28 | $0.00 | 1,000,000 | Yes | Extreme price-performance |
| command-r-08-2024 | Cohere | $0.15 | $0.60 | — | 128,000 | No | Light RAG workloads |
| llama-4-maverick | Meta (via OpenRouter) | $0.15 | $0.60 | — | 1,048,576 | No | Open-weights, hosted |
| mistral-small-4 | Mistral | $0.15 | $0.60 | — | — | No | Cheap general-purpose |
| gpt-5-4-nano | OpenAI | $0.20 | $1.25 | $0.02 | — | Yes | Bulk / classification |
| gemini-3-1-flash-lite | $0.25 | $1.50 | $0.03 | — | Yes | Cheap high-volume | |
| gemini-2-5-flash | $0.30 | $2.50 | $0.03 | — | Yes | Fast general-purpose | |
| codestral | Mistral | $0.30 | $0.90 | — | — | No | Code completion |
| qwen3-7-plus | Alibaba | $0.40 | $1.60 | — | 1,000,000 | Yes | Balanced Qwen tier |
| devstral-2 | Mistral | $0.40 | $2.00 | — | — | No | Coding agents (Mistral) |
| deepseek-v4-pro | DeepSeek | $0.43 | $0.87 | $0.00 | 1,000,000 | Yes | Frontier value, 1M context |
| mistral-large-3 | Mistral | $0.50 | $1.50 | — | 262,144 | No | European flagship value |
| kimi-k2-5 | Moonshot AI | $0.60 | $3.00 | $0.10 | 262,144 | Yes | Value agentic tier |
| glm-4-7 | Zhipu (Z.ai) | $0.60 | $2.20 | $0.11 | — | Yes | Zhipu value tier |
| gpt-5-4-mini | OpenAI | $0.75 | $4.50 | $0.07 | — | Yes | High-volume subtasks |
| nova-pro | Amazon (Bedrock) | $0.80 | $3.20 | — | 300,000 | Yes | AWS-native balanced tier |
| kimi-k2-7-code | Moonshot AI | $0.95 | $4.00 | $0.19 | 262,144 | Yes | Agentic coding (K2 line) |
| claude-haiku-4-5 | Anthropic | $1.00 | $5.00 | $0.10 | 200,000 | Yes | Fast, cheap subtasks |
| qwen3-coder-plus | Alibaba | $1.00 | $5.00 | — | 1,000,000 | Yes | Coding (Qwen line) |
| grok-build-0-1 | xAI | $1.00 | $2.00 | $0.20 | 256,000 | Yes | Coding (xAI) |
| llama-3-3-70b | Meta (via Together AI) | $1.04 | $1.04 | — | — | No | Open-weights workhorse |
| gemini-2-5-pro | $1.25 | $10.00 | $0.13 | — | Yes | Long-context analysis | |
| grok-4-3 | xAI | $1.25 | $2.50 | $0.20 | 1,000,000 | Yes | xAI flagship |
| glm-5-2 | Zhipu (Z.ai) | $1.40 | $4.40 | $0.26 | 1,048,576 | Yes | Zhipu flagship |
| gemini-3-5-flash | $1.50 | $9.00 | $0.15 | — | Yes | Google flagship-speed tier | |
| mistral-medium-3-5 | Mistral | $1.50 | $7.50 | — | 262,144 | No | Mistral premium tier |
| gpt-5-3-codex | OpenAI | $1.75 | $14.00 | $0.17 | — | Yes | Coding agents |
| claude-sonnet-5 | Anthropic | $2.00 | $10.00 | $0.20 | 1,000,000 | Yes | Coding default — intro pricing |
| gemini-3-1-pro-preview | $2.00 | $12.00 | $0.20 | — | Yes | Long-context analysis (preview) | |
| gpt-5-4 | OpenAI | $2.50 | $15.00 | $0.25 | — | Yes | Balanced OpenAI tier |
| qwen3-7-max | Alibaba | $2.50 | $7.50 | — | 1,000,000 | Yes | Alibaba flagship |
| nova-premier | Amazon (Bedrock) | $2.50 | $12.50 | $0.63 | 1,000,000 | Yes | AWS-native flagship |
| command-a | Cohere | $2.50 | $10.00 | — | 256,000 | No | Enterprise RAG / tools |
| claude-sonnet-4-6 | Anthropic | $3.00 | $15.00 | $0.30 | 1,000,000 | Yes | Coding, 1M context |
| claude-sonnet-4-5 | Anthropic | $3.00 | $15.00 | $0.30 | 200,000 | Yes | Coding, analysis, balanced cost |
| claude-opus-4-8 | Anthropic | $5.00 | $25.00 | $0.50 | 1,000,000 | Yes | Deep coding sessions, agentic work |
| claude-opus-4-7 | Anthropic | $5.00 | $25.00 | $0.50 | 1,000,000 | Yes | Legacy Opus tier |
| claude-opus-4-5 | Anthropic | $5.00 | $25.00 | $0.50 | 200,000 | Yes | Legacy Opus tier |
| gpt-5-5 | OpenAI | $5.00 | $30.00 | $0.50 | — | Yes | OpenAI flagship |
| claude-fable-5 | Anthropic | $10.00 | $50.00 | $1.00 | 1,000,000 | Yes | Frontier reasoning, hardest problems |
| gpt-5-5-pro | OpenAI | $30.00 | $180.00 | — | — | No | Max reasoning, price-insensitive |
Provider logos via Lobe Icons (MIT). Logos are trademarks of their respective owners, shown to identify the products compared. tokencheat is not affiliated with, endorsed by, or sponsored by any provider listed.
Last updated: 2026-07-02 · Rates are indicative; verify on provider sites before budgeting.
How to read this table
In/Out are dollars per million tokens; Ctx is the context window. Worked example: at $3/M input, 1M input tokens = $3.00. Output rates are typically several times input rates, so a chatty model on a verbose workload can cost more than a pricier model that answers tersely. Rates in this table are maintained and revalidated, but providers change pricing — verify before committing a budget.
FAQ
- Why are input and output priced separately?
- Providers bill the two lanes at different rates, and workloads skew differently: agentic coding reads far more than it writes, while generation workloads are output-heavy. Comparing on input price alone routinely picks the wrong model.
- What about cache pricing?
- Not shown here. Providers that discount cache reads (Anthropic prices them at roughly 10% of the input rate per their pricing page) can change the effective cost dramatically for stable-prefix workloads — use the token calculator, which includes cache lanes where we have rates.
- Which model is cheapest for me?
- Depends on your token mix. Take your real input and output counts (from a usage dashboard or local logs) and run them through the token calculator against each candidate — the ranking often flips between read-heavy and write-heavy workloads.
- What are the limitations?
- The table shows list rates only: no batch-API discounts, volume tiers, long-context surcharges, or subscription plans. Treat it as a comparison aid, and verify against your provider invoice.
- When should I use this vs the full audit?
- Use this to pick a model. The full stack audit addresses the other half of the bill — how many tokens your setup sends per request — which usually moves cost more than switching models does.
For your full setup, run the free stack audit — or see the 100-configs report.
Pricing sources
Every rate is verified against the provider's own pricing page (or a pass-through aggregator where noted in the catalog). Verified 2026-07-02.
- https://platform.claude.com/docs/en/about-claude/pricing.md
- https://developers.openai.com/api/docs/pricing
- https://ai.google.dev/gemini-api/docs/pricing
- https://docs.x.ai/docs/models
- https://api-docs.deepseek.com/quick_start/pricing
- https://mistral.ai/pricing/api
- https://www.together.ai/pricing
- https://platform.kimi.ai/docs/pricing/chat-k27-code.md
- https://docs.z.ai/guides/overview/pricing
- https://www.alibabacloud.com/help/en/model-studio/model-pricing
- https://cohere.com/pricing
- https://openrouter.ai/api/v1/models
Next step
Picking a cheaper model is the smaller lever.
Rates differ by a few dollars per million tokens. How much context you send on every turn differs by far more, and that is what the audit measures.