OpenAI vs Anthropic vs Google: 2026 AI Coding Cost Breakdown

TokenCheat Team

4/19/2026

#cost-analysis#pricing#2026
OpenAI vs Anthropic vs Google: 2026 AI Coding Cost Breakdown

Model pricing is a moving target — this post has been re-verified against provider pricing pages as of 2026-07-02. Prices change; check the live price comparison before you budget. All rates are USD per 1M tokens, standard pay-as-you-go list price.

Horizontal bar chart of input price per 1M tokens across current coding models, from Claude Fable 5 at $10 down to DeepSeek V4 Flash at $0.14, verified 2026-07-02

How to compare apples

  • Input vs output — coding agents skew input-heavy (file reads, tool schemas, history), but output rates are 4–6× input on most tiers, so both matter.
  • Cache reads — all three providers now price cache reads at 10% of the input rate. The differences are in write charges and TTL control, not the discount.
  • Context window — bigger windows enable fewer round trips, but tempt larger prompts. Google bills a higher rate above 200K prompt tokens on Pro tiers.

Anthropic

ModelTierInputOutputCache read
Claude Fable 5Frontier$10.00$50.00$1.00
Claude Opus 4.8Deep coding / agentic$5.00$25.00$0.50
Claude Sonnet 5Coding default$2.00*$10.00*$0.20
Claude Sonnet 4.6Coding, 1M context$3.00$15.00$0.30
Claude Haiku 4.5Fast subtasks$1.00$5.00$0.10

* Sonnet 5 is intro pricing through 2026-08-31 ($3/$15 after), and its new tokenizer produces roughly 30% more tokens per text than Sonnet 4.6 — compare per-task cost, not just per-token rate.

Anthropic charges for cache writes: 1.25× input for the 5-minute TTL (Opus 4.8: $6.25/M), 2× input for the 1-hour TTL ($10/M on Opus 4.8).

Note the retirement story: the old $15/$75 Opus tier is gone. Opus 4.8 at $5/$25 is a 3× input price cut versus what "Opus pricing" meant in early 2025 — any cost model still carrying the old Opus input rate is stale.

OpenAI

ModelTierInputOutputCache read
GPT-5.5 ProMax reasoning$30.00$180.00— (no caching)
GPT-5.5Flagship$5.00$30.00$0.50
GPT-5.4Balanced$2.50$15.00$0.25
GPT-5.3 CodexCoding agents$1.75$14.00$0.175
GPT-5.4 miniHigh-volume subtasks$0.75$4.50$0.075
GPT-5.4 nanoBulk / classification$0.20$1.25$0.02

OpenAI does not charge for cache writes — caching is free to populate, which makes their effective caching economics slightly better than the identical read discount suggests. OpenAI's previous generation — the GPT-4-class chat models and the o-series reasoning models — is no longer sold; the GPT-5.5/5.4 family replaced it.

Google

ModelTierInputOutputCache read
Gemini 3.1 Pro (Preview)Flagship$2.00*$12.00*$0.20
Gemini 3.5 FlashSpeed flagship$1.50$9.00$0.15
Gemini 2.5 ProLong-context analysis$1.25*$10.00*$0.125
Gemini 2.5 FlashFast general-purpose$0.30$2.50$0.03
Gemini 2.5 Flash-LiteCheapest tier$0.10$0.40$0.01

* Base tier for prompts ≤200K tokens; above that, 3.1 Pro bills $4/$18 and 2.5 Pro bills $2.50/$15. Google's explicit context caching also bills hourly storage (e.g., $1.00/hr on Gemini 3.5 Flash per Google's pricing page).

Beyond the big three

The value tier is where the price war actually lives:

ModelProviderInputOutput
Grok 4.3xAI$1.25$2.50
Kimi K2.7 CodeMoonshot AI$0.95$4.00
Qwen3-Coder-PlusAlibaba$1.00*$5.00*
GLM-4.7Zhipu (Z.ai)$0.60$2.20
Mistral Large 3Mistral$0.50$1.50
DeepSeek V4 ProDeepSeek$0.435$0.87
DeepSeek V4 FlashDeepSeek$0.14$0.28

* Qwen3-Coder-Plus is heavily tiered by context length (input $1–$6, output $5–$60); base tier shown.

A worked month of agent traffic

Assume a team pushes 100M input tokens (70% served from cache) and 5M output tokens through one model in a month. Arithmetic: 30M uncached input

  • 70M cache reads + 5M output, ignoring cache-write surcharges for simplicity (Anthropic charges them, OpenAI doesn't).
ModelUncached inputCache readsOutputTotal
Claude Opus 4.830M × $5 = $150.0070M × $0.50 = $35.005M × $25 = $125.00$310.00
GPT-5.530M × $5 = $150.0070M × $0.50 = $35.005M × $30 = $150.00$335.00
Claude Sonnet 530M × $2 = $60.0070M × $0.20 = $14.005M × $10 = $50.00$124.00
Gemini 2.5 Pro30M × $1.25 = $37.5070M × $0.125 = $8.755M × $10 = $50.00$96.25

The spread between the flagship and mid tiers is 2–3× at identical volume — which is why model routing beats prompt-tweaking as a cost lever.

Use TokenCheat's live table

Our price comparison tracks the TokenCheat catalog — the same numbers used in our calculators and audits, each with a source URL and a last-verified date (currently 2026-07-02). When vendors move rates, the catalog moves with them; this post is a snapshot.