AI Model Pricing Database

Current LLM pricing from every major provider — GPT, Claude, Gemini, DeepSeek, Qwen, Kimi. Updated every two weeks because prices don't sit still

aipricing.tatpenguin.xyz
Last updated: Sep 12, 2026
ProviderModelInput $/MOutput $/MContextStrengthsSubscriptionFree Tier
OpenAIGPT-6 Astra$10.00$50.001.05MCurrent OpenAI flagship (Sep 3-4 2026) — ends the GPT-5 generation; $10/$50 matches Fable 5.1 exactly; 1,050,000-token context, 128K max output, knowledge cutoff Apr 30 2026; cached input $1.00, cache write $12.50; Batch & Flex 50% off, Fast mode 2x; reasoning.effort adds xhigh and max above high; no fine-tuning; first model to trigger OpenAI's 'critical' cyber tier — production version has restricted cyber caps gated behind verificationChatGPT Pro $200/moNo free API tier
OpenAIGPT-5.6 Sol$4.00$20.001.05MNOW THE VALUE FLAGSHIP — promo $4/$20 through Nov 21 2026 (list $5/$30); cached input $0.40; >272K input bills $8 in / $30 out; Batch halves; still the reasoning leader (GPQA-D 94.6%) at 1/5 of Astra's priceChatGPT Plus $20/moLimited on ChatGPT Free
OpenAIGPT-5.6 Terra$2.00$12.001.05MMid tier after the Jul 30 2026 cut ($2.50/$15 → $2/$12); 1.05M ctx; Terminal-Bench 2.1 87.4%; the $2 input tier is the current price war zoneChatGPT Plus $20/moLimited on ChatGPT Free
OpenAIGPT-5.6 Luna$0.20$1.201.05MPrice FLOOR of the mainstream market — cut 80% on Jul 30 2026 ($1/$6 → $0.20/$1.20); a budget price on a current-generation flagship family, not a legacy modelChatGPT Plus $20/moLimited on ChatGPT Free
OpenAIGPT-4.1$2.00$8.001MLegacy workhorse; 1M ctx; being retired — gpt-4.1-nano, gpt-4o-2024-05-13, o1 and o3-mini snapshots all retire Oct 23 2026 (users migrate to gpt-5.6-sol / luna)ChatGPT Plus $20/moChatGPT Free limited
OpenAIGPT-4o$2.50$10.00128KLegacy, price grandfathered; the gpt-4o-2024-05-13 snapshot retires Oct 23 2026 (→ gpt-5.6-sol); gpt-4o-transcribe still $2.50/$10ChatGPT Plus $20/moChatGPT Free limited
OpenAIo3 (Reasoning)$2.00$8.00200KLegacy reasoning tier at post-cut rates; o3-pro $20/$80 above it; cost-sensitive reasoning now goes to GPT-5.6 Sol's thinking mode or TerraChatGPT Pro $200/moLimited on ChatGPT Free
OpenAIo4-mini$1.10$4.40200KLegacy cheap reasoning; fast at coding & STEM; superseded on price by GPT-5.6 Terra ($2/$12) and Luna ($0.20/$1.20)ChatGPT Plus $20/moLimited
AnthropicClaude Fable 5.1$10.00$50.001MCurrent #1 frontier model (Sep 1 2026); CACHE READS CUT 75%: $1.00 → $0.25/MTok (0.025x input, vs 0.1x industry norm) — ~25% cheaper typical, up to ~45% agentic; 5m cache write $12.50, 1h $20; no long-context surcharge on the full 1M window; Batch $5/$25; AA Intelligence Index 66; gated twin Mythos 5.1 (same weights, $10/$50, cyber defenders only)Claude Pro $20/moClaude Free limited
AnthropicClaude Opus 5$5.00$25.001M3x cheaper than Opus 4 ($15/$75); cache hits $0.50; fast mode $10/$50; matches Opus 4.8 rate; AA Index 63 — the value pick for hard reasoningClaude Pro $20/moLimited
AnthropicClaude Sonnet 5$2.00$10.001M$2/$10 made PERMANENT on Aug 11 2026 — the scheduled Sep 1 step-up to $3/$15 was cancelled (first headline increase to evaporate before taking effect); cache read $0.20; the coding/agentic defaultClaude Pro $20/moClaude Free limited
AnthropicClaude Haiku 4.5$1.00$5.00200KCheapest Claude; cache read $0.10; Haiku 3.5 ($0.80/$4) retiredClaude Pro $20/moClaude Free limited
GoogleGemini 3.8 Flash$0.75$3.751MNEWEST Flash (Sep 2 2026) — 4th Flash in under 4 months; Terminal-Bench 2.1 jumped 81.6% → 90.8%; intro rate $0.75/$3.75 (thinking included) through Dec 31 2026 then DOUBLES to $1.50/$7.50 on Jan 1 2027; cache $0.075; Batch/Flex 50% off; Priority $1.35/$6.75; works harder per task (~40% higher cost/task than 3.7 at same token price)Google AI Pro $19.99/moGemini Free limited
GoogleGemini 3.7 Flash$0.75$3.751MSame intro rate as 3.8 through Dec 31 2026 ($1.50/$7.50 from Jan 1 2027); still supported — the cheaper-per-task fallback when 3.8's token hunger bites; cache $0.075Google AI Pro $19.99/moGemini Free limited
GoogleGemini 3.1 Pro (Preview)$2.00$12.001MTop Pro tier; >200K input bills $4 in / $18 out; cache $0.20 (storage $4.50/1M/hr); one of the three models sitting exactly at the $2 input price pointGoogle AI Pro $19.99/mo (was Gemini Advanced)Gemini Free limited
GoogleGemini 2.5 Pro$1.25$10.001MLegacy Pro; >200K input $2.50 in / $15 out; cache $0.125; video understandingGoogle AI Pro $19.99/moGemini Free limited
GoogleGemini 2.5 Flash$0.30$2.501MLegacy cheap Flash; $0.30 text/image/video ($1.00 audio) in; cache $0.03; Flash-Lite tiers below itGoogle AI Pro $19.99/moGemini Free limited
DeepSeekDeepSeek V4.1 Flash$0.15$0.601MREPLACES V4 Flash (GA Sep 10 2026, model name deepseek-flash); OFF-PEAK $0.15 in / $0.60 out / cache hit $0.003 — PEAK ×2 = $0.30/$1.20/$0.006 (peak = Mon-Fri 01:00-04:00 & 06:00-10:00 UTC); 552B multimodal MoE, native vision, 384K max output, MIT open weights; KV-cache memory cut to 25% of V4-Flash for long agent sessions; legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp now billed at Flash priceDeepSeek Chat Free5M free tokens for new users
DeepSeekDeepSeek V4 Pro$0.66$1.981MV4-Pro-0813: off-peak $0.66 in / $1.98 out / cache $0.022; peak ×2 = $1.32/$3.96/$0.044 — the counter-trend rise of up to 14x on Aug 13 2026 while everyone else cut; official docs (checked Sep 12) confirm V4 Pro CONTINUES past Sep 14 with billing unchanged, superseding reports that it would auto-route to V4.1 Flash at Flash rates; no vision on ProDeepSeek Chat Free5M free tokens for new users
DeepSeekDeepSeek R1 (Legacy)$0.55$2.1964KRETIRED Jul 24 2026 along with V3.2 and the deepseek-chat / deepseek-reasoner aliases — reasoning now served by V4.1 Flash thinking modeDeepSeek Chat FreeFree with rate limits
xAIGrok 4.6$2.00$6.00500KCurrent xAI flagship (Aug 12 2026, on Bedrock since Aug 19); cached input $0.50; ≥200K prompt bills the whole request at $4 in / $12 out; long-running agents, coding, visual workSuperGrok $30/moLimited on X
xAIGrok 4.3$1.25$2.501MBudget tier; cached input $0.20; >200K rate $2.50/$5.00; Grok 4.1 Fast sits below at $0.20/$0.50. Legacy Grok 3 was $3/$15SuperGrok $30/moLimited on X
MetaMuse Spark 1.3$1.25$4.251MMeta's current frontier line (Muse Spark replaced Llama as the flagship in Apr 2026); standard tier $1.25/$4.25 unchanged from 1.2; ~20% fewer tool calls and ~25% fewer tokens than 1.2; CONTRIBUTOR TIER $0.10/$0.20 (~$0.10/1M blended — cheapest model in the top-5 leaderboard) but Meta may train on your prompts and outputs; max-reasoning mode held back at launchN/A (API + Meta AI app)Limited free on Meta AI
MetaLlama 4 Maverick (open weights)$0.20$0.601MStill the cheap hosted open-weight option ($0.20 via Vercel, $0.15-$0.35 across 7 providers; DeepInfra $0.15/$0.60); no longer Meta's frontier model — see Muse Spark 1.3; Intelligence Index 14.5 (45th pct) shows how far the frontier has movedN/A (open model)Free to self-host
MistralMistral Large 3$0.50$1.50262KOpen-weight flagship; cached input $0.05; vision + function calling; best price-to-quality in the mid tier; Medium 3.5 above it at $1.50/$7.50, Small 4 below at $0.10/$0.30Le Chat Pro $14.99/moLe Chat Free limited
AlibabaQwen3.8-Max$2.00$6.001MAlibaba flagship (GA Aug 3 2026); flat rate across the full 1M context (no >200K surcharge); cached input $0.25; 2.4T MoE with 1M ctx; third-party hosts undercut it (AIHubMix $1.69/$5.07 on the 0902 snapshot)Token Plan discount1M free tokens
AlibabaQwen3.8-Flash$0.15$0.47262K (→1M)Ultra-cheap multimodal MoE (live Aug 27 2026, QwenCloud); cache hit $0.016; NO peak surcharge; 262K native context extensible to 1M; open weights ship separately as Qwen3.8-Flash-Next (Qwen Community 1.0 license)Token Plan discount1M free tokens
AlibabaQwen3.7-Max$1.25$3.751MPrevious flagship, still cheapest route to Max-class quality — 50% promo off $2.50/$7.50 list; frontier coding agent, 35h autonomous runsToken Plan discount1M free tokens
AlibabaQwen-Flash$0.05$0.40256KBudget tier that REPLACED Qwen-Turbo (Turbo no longer updated); $0.05/$0.40 up to 256K, rising to $0.25/$2.00 on longer inputs; mid-tier Qwen-Plus starts at $0.40/$1.20 (aggregator-reported rates — check Alibaba's own rate card before committing)Token Plan discount1M free tokens
Zhipu (Z.ai)GLM-5.3-Flash$0.15$0.501MPRICE ROSE Sep 9 2026 — the 50% launch promo ($0.075/$0.25) ended at 24:00 UTC+8, so z.ai now bills list: $0.15 in / $0.50 out / $0.03 cached (storage limited-time free); 320B MoE (18B active), native vision + video, MIT weights, 1M ctx; AA Index 57 ≈ Gemini 3.8 Flash at medium effort for ~1/6 the price; cheapest top-tier blend on the board at $0.17/1MGLM Coding Plan20M tokens on signup
Zhipu (Z.ai)GLM-5.3$1.40$4.401MText flagship, same price as GLM-5.2/5.1; cached input $0.26; open weights expected mid-to-late Sep 2026; CyberGym 84.5% SOTA with 2,436 real vulnerabilities found; thinking cannot be disabledGLM Coding Plan20M tokens on signup
ByteDanceDoubao Seed 2.1 Pro$0.88$4.41256KVolcano Ark rate card (Jun 24 2026): ¥6 in / ¥30 out, cache hit ¥1.2 (~$0.18); Turbo is exactly half (¥3/¥15 ≈ $0.44/$2.21); text+image+video in; USD figures are conversions at 6.80 CNY — not a published dollar price card; supersedes the old Doubao 1.5 Pro (32K, ¥0.8/¥2)Free tier on Doubao appFree quotas on Volcano Engine
MoonshotKimi K3$3.00$15.001MMost powerful open-weights model by GPQA (93.5%), released Jul 16 2026 with weights in late Jul under Modified MIT; 2.8T MoE with native vision; FLAT $3/$15 across the full 1M window — no long-context surcharge; cache hit $0.30; coding sibling K2.7 Code $0.95/$4.00 (cache $0.19); Batch not available on K3Kimi free app; $1 min top-up on APINo free API tier
01.AIYi-Lightning$0.14$0.14128KUNVERIFIED — earlier data claimed this API was shutting down Sep 3 2026 with refunds to Dec 3; no such notice could be found on 01.ai, in the 2026 deprecation trackers, or in the press as of Sep 12 2026. What IS verifiable: 01.AI stopped pre-training large models in Mar 2025 and now sells enterprise/agent solutions (万策/万智 platforms), so Yi-Lightning API availability is doubtful but NOT confirmed discontinued. Treat this row as legacy/unconfirmedN/AFree tier (historical)
Inception LabsMercury 2$0.25$0.75128KDiffusion (non-autoregressive) decoder, not a transformer-with-next-token — 1,009+ tok/s on Blackwell; cached input $0.025; native tool use, schema-aligned JSON, OpenAI-compatible (switch is a base-URL change); live since Mar 4 2026, #18/134 on AA intelligence; Mercury 2.5 Preview listed at $0.04/$0.15 (tracker-reported)N/A10M free tokens on signup
Sakana AIFugu Max$2.00$6.00NEW Sep 11 2026 — multi-agent ORCHESTRATION product rather than a single model (OpenAI-compatible API); cost tier of the Fugu line, undercutting Sonnet 5 / GPT-5.6 Terra / Kimi K3 output by 40-60% at the same $6 output as Qwen3.8-Max and Grok 4.6; capability twin Fugu Ultra v2.0 is $5/$30 (cached $0.50; >272K input $10/$45/$1.00) — the price war has moved from per-model inference to orchestration economicsN/ANo free tier reported