Model prices
Every rate used by the inference cost calculator, in USD per 1 million tokens on the standard API tier. Verified September 23, 2026 against the provider pricing pages linked below.
~~~
OpenAI
pricing source ↗| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| GPT-6 Astra | flagship | $10.00 | $1.00 | $50.00 | 1.05M | OpenAI flagship for the hardest agents, coding and research |
| GPT-6 Sol | workhorse | $2.00 | $0.20 | $10.00 | 1.05M | Complex coding and agentic workflows at a mid-tier price |
| GPT-5.4 mini | workhorse | $0.75 | $0.075 | $4.50 | 400K | High-volume chat, RAG, extraction, support bots |
| GPT-5.4 nano | budget | $0.20 | $0.020 | $1.25 | 400K | Routing, tagging, classification, simple summaries |
| GPT-6 Luna | budget | $0.10 | $0.010 | $0.50 | 1.05M | Cheapest OpenAI model for focused, high-volume tasks |
Anthropic
pricing source ↗| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | flagship | $10.00 | $0.25 | $50.00 | 1M | Long-running agents that need the most capable ClaudeCache reads cost 2.5% of the input price. |
| Claude Opus 5.5 | flagship | $4.00 | $0.20 | $20.00 | 1M | Complex agentic coding and enterprise workCache reads cost 5% of the input price. |
| Claude Sonnet 5 | workhorse | $2.00 | $0.20 | $10.00 | 1M | Best speed/intelligence balance; daily-driver agents |
| Claude Haiku 4.5 | workhorse | $1.00 | $0.10 | $5.00 | 200K | Fast near-frontier chat and tool use |
| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| Gemini 3.1 Pro | flagship | $2.00 | $0.20 | $12.00 | 1M | Flagship reasoning with long contextPrompts over 200K tokens are billed at $4/$18. |
| Gemini 3.8 Flash | workhorse | $0.75 | $0.075 | $3.75 | 1M | Agents and coding at Flash pricesIntro price through Dec 31, 2026, then $1.50/$7.50. |
| Gemini 3.5 Flash-Lite | budget | $0.30 | $0.030 | $2.50 | 1M | Bulk simple tasks with a newer model |
| Gemini 3.1 Flash-Lite | budget | $0.25 | $0.025 | $1.50 | 1M | Bulk simple tasks, translation, extraction |
| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| Grok 4.7 | workhorse | $2.00 | $0.50 | $6.00 | 500K | Latest Grok for chat and tool usePrompts over 200K tokens are billed at $4/$12. |
| Grok 4.3 | workhorse | $1.25 | $0.20 | $2.50 | 1M | Aggressively priced frontier-class chat and tool use |
Mistral
pricing source ↗| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| Mistral Large 3open | workhorse | $0.50 | $0.050 | $1.50 | 256K | Open-weight European frontier; multimodal MoE |
| Mistral Small 4open | budget | $0.15 | $0.015 | $0.60 | 256K | Cheap general tasks; open-weight |
| Ministral 3 8Bopen | budget | $0.15 | $0.015 | $0.15 | 256K | Edge and embedded deployments |
OpenRouter
pricing source ↗Aggregator. Listed price is the default routing; actual price varies by underlying provider.
| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| Kimi K3 | flagship | $3.00 | $0.30 | $15.00 | 1M | Moonshot's top model for agents and writing |
| Kimi K2.6open | workhorse | $0.95 | $0.16 | $4.00 | 262K | Strong open-weight agentic and writing model |
| DeepSeek V4 Proopen | workhorse | $0.955 | $0.080 | $1.911 | 1M | Frontier-class reasoning and coding at budget rates |
| DeepSeek V4 Flashopen | budget | $0.089 | $0.018 | $0.177 | 1M | The open-weight price leader for chat and agents |
| Qwen3.8 Flash | budget | $0.15 | $0.016 | $0.47 | 1M | Cheap long-context chat and extraction |
Workers AI
pricing source ↗Cloudflare edge inference. Billed in neurons ($0.011 per 1k); per-token equivalents shown. 10k neurons/day free.
| Model | Tier | Input / 1M | Cached / 1M | Output / 1M | Context | Best for |
|---|---|---|---|---|---|---|
| GPT-OSS 120Bopen | workhorse | $0.35 | — | $0.75 | 128K | OpenAI's open-weight model at the edge, no cold starts |
| Llama 3.3 70Bopen | workhorse | $0.293 | — | $2.253 | 131K | Capable open-weight default on Cloudflare's edge |
| Llama 3.1 8B (fast)open | budget | $0.045 | — | $0.384 | 131K | Very cheap edge inference for simple tasks |
| Llama 3.2 1Bopen | budget | $0.027 | — | $0.201 | 128K | The cheapest credible model for routing/tagging |
~~~
Free books, courses, and software
Get my programming books, course editions, and complete software.
Get the download library →