Model prices

← Back to calculator

Every rate used by the inference cost calculator, in USD per 1 million tokens on the standard API tier. Verified September 23, 2026 against the provider pricing pages linked below.

~~~
ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
GPT-6 Astraflagship$10.00$1.00$50.001.05MOpenAI flagship for the hardest agents, coding and research
GPT-6 Solworkhorse$2.00$0.20$10.001.05MComplex coding and agentic workflows at a mid-tier price
GPT-5.4 miniworkhorse$0.75$0.075$4.50400KHigh-volume chat, RAG, extraction, support bots
GPT-5.4 nanobudget$0.20$0.020$1.25400KRouting, tagging, classification, simple summaries
GPT-6 Lunabudget$0.10$0.010$0.501.05MCheapest OpenAI model for focused, high-volume tasks
ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
Claude Fable 5.1flagship$10.00$0.25$50.001MLong-running agents that need the most capable ClaudeCache reads cost 2.5% of the input price.
Claude Opus 5.5flagship$4.00$0.20$20.001MComplex agentic coding and enterprise workCache reads cost 5% of the input price.
Claude Sonnet 5workhorse$2.00$0.20$10.001MBest speed/intelligence balance; daily-driver agents
Claude Haiku 4.5workhorse$1.00$0.10$5.00200KFast near-frontier chat and tool use
ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
Gemini 3.1 Proflagship$2.00$0.20$12.001MFlagship reasoning with long contextPrompts over 200K tokens are billed at $4/$18.
Gemini 3.8 Flashworkhorse$0.75$0.075$3.751MAgents and coding at Flash pricesIntro price through Dec 31, 2026, then $1.50/$7.50.
Gemini 3.5 Flash-Litebudget$0.30$0.030$2.501MBulk simple tasks with a newer model
Gemini 3.1 Flash-Litebudget$0.25$0.025$1.501MBulk simple tasks, translation, extraction
ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
Grok 4.7workhorse$2.00$0.50$6.00500KLatest Grok for chat and tool usePrompts over 200K tokens are billed at $4/$12.
Grok 4.3workhorse$1.25$0.20$2.501MAggressively priced frontier-class chat and tool use
ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
Mistral Large 3openworkhorse$0.50$0.050$1.50256KOpen-weight European frontier; multimodal MoE
Mistral Small 4openbudget$0.15$0.015$0.60256KCheap general tasks; open-weight
Ministral 3 8Bopenbudget$0.15$0.015$0.15256KEdge and embedded deployments

Aggregator. Listed price is the default routing; actual price varies by underlying provider.

ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
Kimi K3flagship$3.00$0.30$15.001MMoonshot's top model for agents and writing
Kimi K2.6openworkhorse$0.95$0.16$4.00262KStrong open-weight agentic and writing model
DeepSeek V4 Proopenworkhorse$0.955$0.080$1.9111MFrontier-class reasoning and coding at budget rates
DeepSeek V4 Flashopenbudget$0.089$0.018$0.1771MThe open-weight price leader for chat and agents
Qwen3.8 Flashbudget$0.15$0.016$0.471MCheap long-context chat and extraction

Cloudflare edge inference. Billed in neurons ($0.011 per 1k); per-token equivalents shown. 10k neurons/day free.

ModelTierInput / 1MCached / 1MOutput / 1MContextBest for
GPT-OSS 120Bopenworkhorse$0.35—$0.75128KOpenAI's open-weight model at the edge, no cold starts
Llama 3.3 70Bopenworkhorse$0.293—$2.253131KCapable open-weight default on Cloudflare's edge
Llama 3.1 8B (fast)openbudget$0.045—$0.384131KVery cheap edge inference for simple tasks
Llama 3.2 1Bopenbudget$0.027—$0.201128KThe cheapest credible model for routing/tagging
~~~

Free books, courses, and software

Get my programming books, course editions, and complete software.

Get the download library →