OpenAI vs Anthropic vs Google: 2026 AI Coding Cost Breakdown
TokenCheat Team
4/19/2026

Model pricing is a moving target — this post has been re-verified against provider pricing pages as of 2026-07-02. Prices change; check the live price comparison before you budget. All rates are USD per 1M tokens, standard pay-as-you-go list price.
How to compare apples
- Input vs output — coding agents skew input-heavy (file reads, tool schemas, history), but output rates are 4–6× input on most tiers, so both matter.
- Cache reads — all three providers now price cache reads at 10% of the input rate. The differences are in write charges and TTL control, not the discount.
- Context window — bigger windows enable fewer round trips, but tempt larger prompts. Google bills a higher rate above 200K prompt tokens on Pro tiers.
Anthropic
| Model | Tier | Input | Output | Cache read |
|---|---|---|---|---|
| Claude Fable 5 | Frontier | $10.00 | $50.00 | $1.00 |
| Claude Opus 4.8 | Deep coding / agentic | $5.00 | $25.00 | $0.50 |
| Claude Sonnet 5 | Coding default | $2.00* | $10.00* | $0.20 |
| Claude Sonnet 4.6 | Coding, 1M context | $3.00 | $15.00 | $0.30 |
| Claude Haiku 4.5 | Fast subtasks | $1.00 | $5.00 | $0.10 |
* Sonnet 5 is intro pricing through 2026-08-31 ($3/$15 after), and its new tokenizer produces roughly 30% more tokens per text than Sonnet 4.6 — compare per-task cost, not just per-token rate.
Anthropic charges for cache writes: 1.25× input for the 5-minute TTL (Opus 4.8: $6.25/M), 2× input for the 1-hour TTL ($10/M on Opus 4.8).
Note the retirement story: the old $15/$75 Opus tier is gone. Opus 4.8 at $5/$25 is a 3× input price cut versus what "Opus pricing" meant in early 2025 — any cost model still carrying the old Opus input rate is stale.
OpenAI
| Model | Tier | Input | Output | Cache read |
|---|---|---|---|---|
| GPT-5.5 Pro | Max reasoning | $30.00 | $180.00 | — (no caching) |
| GPT-5.5 | Flagship | $5.00 | $30.00 | $0.50 |
| GPT-5.4 | Balanced | $2.50 | $15.00 | $0.25 |
| GPT-5.3 Codex | Coding agents | $1.75 | $14.00 | $0.175 |
| GPT-5.4 mini | High-volume subtasks | $0.75 | $4.50 | $0.075 |
| GPT-5.4 nano | Bulk / classification | $0.20 | $1.25 | $0.02 |
OpenAI does not charge for cache writes — caching is free to populate, which makes their effective caching economics slightly better than the identical read discount suggests. OpenAI's previous generation — the GPT-4-class chat models and the o-series reasoning models — is no longer sold; the GPT-5.5/5.4 family replaced it.
| Model | Tier | Input | Output | Cache read |
|---|---|---|---|---|
| Gemini 3.1 Pro (Preview) | Flagship | $2.00* | $12.00* | $0.20 |
| Gemini 3.5 Flash | Speed flagship | $1.50 | $9.00 | $0.15 |
| Gemini 2.5 Pro | Long-context analysis | $1.25* | $10.00* | $0.125 |
| Gemini 2.5 Flash | Fast general-purpose | $0.30 | $2.50 | $0.03 |
| Gemini 2.5 Flash-Lite | Cheapest tier | $0.10 | $0.40 | $0.01 |
* Base tier for prompts ≤200K tokens; above that, 3.1 Pro bills $4/$18 and 2.5 Pro bills $2.50/$15. Google's explicit context caching also bills hourly storage (e.g., $1.00/hr on Gemini 3.5 Flash per Google's pricing page).
Beyond the big three
The value tier is where the price war actually lives:
| Model | Provider | Input | Output |
|---|---|---|---|
| Grok 4.3 | xAI | $1.25 | $2.50 |
| Kimi K2.7 Code | Moonshot AI | $0.95 | $4.00 |
| Qwen3-Coder-Plus | Alibaba | $1.00* | $5.00* |
| GLM-4.7 | Zhipu (Z.ai) | $0.60 | $2.20 |
| Mistral Large 3 | Mistral | $0.50 | $1.50 |
| DeepSeek V4 Pro | DeepSeek | $0.435 | $0.87 |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 |
* Qwen3-Coder-Plus is heavily tiered by context length (input $1–$6, output $5–$60); base tier shown.
A worked month of agent traffic
Assume a team pushes 100M input tokens (70% served from cache) and 5M output tokens through one model in a month. Arithmetic: 30M uncached input
- 70M cache reads + 5M output, ignoring cache-write surcharges for simplicity (Anthropic charges them, OpenAI doesn't).
| Model | Uncached input | Cache reads | Output | Total |
|---|---|---|---|---|
| Claude Opus 4.8 | 30M × $5 = $150.00 | 70M × $0.50 = $35.00 | 5M × $25 = $125.00 | $310.00 |
| GPT-5.5 | 30M × $5 = $150.00 | 70M × $0.50 = $35.00 | 5M × $30 = $150.00 | $335.00 |
| Claude Sonnet 5 | 30M × $2 = $60.00 | 70M × $0.20 = $14.00 | 5M × $10 = $50.00 | $124.00 |
| Gemini 2.5 Pro | 30M × $1.25 = $37.50 | 70M × $0.125 = $8.75 | 5M × $10 = $50.00 | $96.25 |
The spread between the flagship and mid tiers is 2–3× at identical volume — which is why model routing beats prompt-tweaking as a cost lever.
Use TokenCheat's live table
Our price comparison tracks the TokenCheat catalog — the same numbers used in our calculators and audits, each with a source URL and a last-verified date (currently 2026-07-02). When vendors move rates, the catalog moves with them; this post is a snapshot.