Model pricing table

Filter by provider, caching, and context window. Data comes from packages/ai/lib/model-pricing-catalog.ts.

Last updated:

qwen-flashAlibaba$0.05$0.401,000,000YesBulk / cheapest Qwen
nova-liteAmazon (Bedrock)$0.06$0.24300,000YesAWS-native high-volume
gemini-2-5-flash-liteGoogle$0.10$0.40$0.01YesCheapest Google tier
llama-4-scoutMeta (via OpenRouter)$0.10$0.3010,000,000NoOpen-weights, 10M context
deepseek-v4-flashDeepSeek$0.14$0.28$0.001,000,000YesExtreme price-performance
command-r-08-2024Cohere$0.15$0.60128,000NoLight RAG workloads
llama-4-maverickMeta (via OpenRouter)$0.15$0.601,048,576NoOpen-weights, hosted
mistral-small-4Mistral$0.15$0.60NoCheap general-purpose
gpt-5-4-nanoOpenAI$0.20$1.25$0.02YesBulk / classification
gemini-3-1-flash-liteGoogle$0.25$1.50$0.03YesCheap high-volume
gemini-2-5-flashGoogle$0.30$2.50$0.03YesFast general-purpose
codestralMistral$0.30$0.90NoCode completion
qwen3-7-plusAlibaba$0.40$1.601,000,000YesBalanced Qwen tier
devstral-2Mistral$0.40$2.00NoCoding agents (Mistral)
deepseek-v4-proDeepSeek$0.43$0.87$0.001,000,000YesFrontier value, 1M context
mistral-large-3Mistral$0.50$1.50262,144NoEuropean flagship value
kimi-k2-5Moonshot AI$0.60$3.00$0.10262,144YesValue agentic tier
glm-4-7Zhipu (Z.ai)$0.60$2.20$0.11YesZhipu value tier
gpt-5-4-miniOpenAI$0.75$4.50$0.07YesHigh-volume subtasks
nova-proAmazon (Bedrock)$0.80$3.20300,000YesAWS-native balanced tier
kimi-k2-7-codeMoonshot AI$0.95$4.00$0.19262,144YesAgentic coding (K2 line)
claude-haiku-4-5Anthropic$1.00$5.00$0.10200,000YesFast, cheap subtasks
qwen3-coder-plusAlibaba$1.00$5.001,000,000YesCoding (Qwen line)
grok-build-0-1xAI$1.00$2.00$0.20256,000YesCoding (xAI)
llama-3-3-70bMeta (via Together AI)$1.04$1.04NoOpen-weights workhorse
gemini-2-5-proGoogle$1.25$10.00$0.13YesLong-context analysis
grok-4-3xAI$1.25$2.50$0.201,000,000YesxAI flagship
glm-5-2Zhipu (Z.ai)$1.40$4.40$0.261,048,576YesZhipu flagship
gemini-3-5-flashGoogle$1.50$9.00$0.15YesGoogle flagship-speed tier
mistral-medium-3-5Mistral$1.50$7.50262,144NoMistral premium tier
gpt-5-3-codexOpenAI$1.75$14.00$0.17YesCoding agents
claude-sonnet-5Anthropic$2.00$10.00$0.201,000,000YesCoding default — intro pricing
gemini-3-1-pro-previewGoogle$2.00$12.00$0.20YesLong-context analysis (preview)
gpt-5-4OpenAI$2.50$15.00$0.25YesBalanced OpenAI tier
qwen3-7-maxAlibaba$2.50$7.501,000,000YesAlibaba flagship
nova-premierAmazon (Bedrock)$2.50$12.50$0.631,000,000YesAWS-native flagship
command-aCohere$2.50$10.00256,000NoEnterprise RAG / tools
claude-sonnet-4-6Anthropic$3.00$15.00$0.301,000,000YesCoding, 1M context
claude-sonnet-4-5Anthropic$3.00$15.00$0.30200,000YesCoding, analysis, balanced cost
claude-opus-4-8Anthropic$5.00$25.00$0.501,000,000YesDeep coding sessions, agentic work
claude-opus-4-7Anthropic$5.00$25.00$0.501,000,000YesLegacy Opus tier
claude-opus-4-5Anthropic$5.00$25.00$0.50200,000YesLegacy Opus tier
gpt-5-5OpenAI$5.00$30.00$0.50YesOpenAI flagship
claude-fable-5Anthropic$10.00$50.00$1.001,000,000YesFrontier reasoning, hardest problems
gpt-5-5-proOpenAI$30.00$180.00NoMax reasoning, price-insensitive

Last updated: April 2026 · Rates are indicative; verify on provider sites before budgeting.

FAQ

Where do these prices come from?
Each rate is read from the provider's published pricing page; entries that could only be verified through a pass-through aggregator are marked as secondary sources in the catalog. Every entry records its source URL.
How current is the pricing?
The catalog carries a last-verified date — currently 2026-07-02 — shown above the table. Models withdrawn from the catalog are flagged as deprecated rather than deleted.
What do the cache read and cache write rates mean?
Providers with prompt caching bill cached input at a discounted read rate and charge a write rate to create the cache entry. Models without caching support show no cache pricing.
Why are prices listed per million tokens?
Per-million-token (MTok) rates are the industry convention and keep small per-request costs comparable across providers. Multiply the rates by your actual input and output volumes to estimate spend.