Current LLM pricing from every major provider — GPT, Claude, Gemini, DeepSeek, Qwen, Kimi. Updated every two weeks because prices don't sit still
| Provider↕ | Model↕ | Input $/M↕ | Output $/M↕ | Context↕ | Strengths↕ | Subscription↕ | Free Tier↕ |
|---|---|---|---|---|---|---|---|
| OpenAI | GPT-6 Astra | $10.00 | $50.00 | 1.05M | Current OpenAI flagship (Sep 3-4 2026) — ends the GPT-5 generation; $10/$50 matches Fable 5.1 exactly; 1,050,000-token context, 128K max output, knowledge cutoff Apr 30 2026; cached input $1.00, cache write $12.50; Batch & Flex 50% off, Fast mode 2x; reasoning.effort adds xhigh and max above high; no fine-tuning; first model to trigger OpenAI's 'critical' cyber tier — production version has restricted cyber caps gated behind verification | ChatGPT Pro $200/mo | No free API tier |
| OpenAI | GPT-5.6 Sol | $4.00 | $20.00 | 1.05M | NOW THE VALUE FLAGSHIP — promo $4/$20 through Nov 21 2026 (list $5/$30); cached input $0.40; >272K input bills $8 in / $30 out; Batch halves; still the reasoning leader (GPQA-D 94.6%) at 1/5 of Astra's price | ChatGPT Plus $20/mo | Limited on ChatGPT Free |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | 1.05M | Mid tier after the Jul 30 2026 cut ($2.50/$15 → $2/$12); 1.05M ctx; Terminal-Bench 2.1 87.4%; the $2 input tier is the current price war zone | ChatGPT Plus $20/mo | Limited on ChatGPT Free |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | 1.05M | Price FLOOR of the mainstream market — cut 80% on Jul 30 2026 ($1/$6 → $0.20/$1.20); a budget price on a current-generation flagship family, not a legacy model | ChatGPT Plus $20/mo | Limited on ChatGPT Free |
| OpenAI | GPT-4.1 | $2.00 | $8.00 | 1M | Legacy workhorse; 1M ctx; being retired — gpt-4.1-nano, gpt-4o-2024-05-13, o1 and o3-mini snapshots all retire Oct 23 2026 (users migrate to gpt-5.6-sol / luna) | ChatGPT Plus $20/mo | ChatGPT Free limited |
| OpenAI | GPT-4o | $2.50 | $10.00 | 128K | Legacy, price grandfathered; the gpt-4o-2024-05-13 snapshot retires Oct 23 2026 (→ gpt-5.6-sol); gpt-4o-transcribe still $2.50/$10 | ChatGPT Plus $20/mo | ChatGPT Free limited |
| OpenAI | o3 (Reasoning) | $2.00 | $8.00 | 200K | Legacy reasoning tier at post-cut rates; o3-pro $20/$80 above it; cost-sensitive reasoning now goes to GPT-5.6 Sol's thinking mode or Terra | ChatGPT Pro $200/mo | Limited on ChatGPT Free |
| OpenAI | o4-mini | $1.10 | $4.40 | 200K | Legacy cheap reasoning; fast at coding & STEM; superseded on price by GPT-5.6 Terra ($2/$12) and Luna ($0.20/$1.20) | ChatGPT Plus $20/mo | Limited |
| Anthropic | Claude Fable 5.1 | $10.00 | $50.00 | 1M | Current #1 frontier model (Sep 1 2026); CACHE READS CUT 75%: $1.00 → $0.25/MTok (0.025x input, vs 0.1x industry norm) — ~25% cheaper typical, up to ~45% agentic; 5m cache write $12.50, 1h $20; no long-context surcharge on the full 1M window; Batch $5/$25; AA Intelligence Index 66; gated twin Mythos 5.1 (same weights, $10/$50, cyber defenders only) | Claude Pro $20/mo | Claude Free limited |
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | 1M | 3x cheaper than Opus 4 ($15/$75); cache hits $0.50; fast mode $10/$50; matches Opus 4.8 rate; AA Index 63 — the value pick for hard reasoning | Claude Pro $20/mo | Limited |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1M | $2/$10 made PERMANENT on Aug 11 2026 — the scheduled Sep 1 step-up to $3/$15 was cancelled (first headline increase to evaporate before taking effect); cache read $0.20; the coding/agentic default | Claude Pro $20/mo | Claude Free limited |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K | Cheapest Claude; cache read $0.10; Haiku 3.5 ($0.80/$4) retired | Claude Pro $20/mo | Claude Free limited |
| Gemini 3.8 Flash | $0.75 | $3.75 | 1M | NEWEST Flash (Sep 2 2026) — 4th Flash in under 4 months; Terminal-Bench 2.1 jumped 81.6% → 90.8%; intro rate $0.75/$3.75 (thinking included) through Dec 31 2026 then DOUBLES to $1.50/$7.50 on Jan 1 2027; cache $0.075; Batch/Flex 50% off; Priority $1.35/$6.75; works harder per task (~40% higher cost/task than 3.7 at same token price) | Google AI Pro $19.99/mo | Gemini Free limited | |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1M | Same intro rate as 3.8 through Dec 31 2026 ($1.50/$7.50 from Jan 1 2027); still supported — the cheaper-per-task fallback when 3.8's token hunger bites; cache $0.075 | Google AI Pro $19.99/mo | Gemini Free limited | |
| Gemini 3.1 Pro (Preview) | $2.00 | $12.00 | 1M | Top Pro tier; >200K input bills $4 in / $18 out; cache $0.20 (storage $4.50/1M/hr); one of the three models sitting exactly at the $2 input price point | Google AI Pro $19.99/mo (was Gemini Advanced) | Gemini Free limited | |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1M | Legacy Pro; >200K input $2.50 in / $15 out; cache $0.125; video understanding | Google AI Pro $19.99/mo | Gemini Free limited | |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1M | Legacy cheap Flash; $0.30 text/image/video ($1.00 audio) in; cache $0.03; Flash-Lite tiers below it | Google AI Pro $19.99/mo | Gemini Free limited | |
| DeepSeek | DeepSeek V4.1 Flash | $0.15 | $0.60 | 1M | REPLACES V4 Flash (GA Sep 10 2026, model name deepseek-flash); OFF-PEAK $0.15 in / $0.60 out / cache hit $0.003 — PEAK ×2 = $0.30/$1.20/$0.006 (peak = Mon-Fri 01:00-04:00 & 06:00-10:00 UTC); 552B multimodal MoE, native vision, 384K max output, MIT open weights; KV-cache memory cut to 25% of V4-Flash for long agent sessions; legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp now billed at Flash price | DeepSeek Chat Free | 5M free tokens for new users |
| DeepSeek | DeepSeek V4 Pro | $0.66 | $1.98 | 1M | V4-Pro-0813: off-peak $0.66 in / $1.98 out / cache $0.022; peak ×2 = $1.32/$3.96/$0.044 — the counter-trend rise of up to 14x on Aug 13 2026 while everyone else cut; official docs (checked Sep 12) confirm V4 Pro CONTINUES past Sep 14 with billing unchanged, superseding reports that it would auto-route to V4.1 Flash at Flash rates; no vision on Pro | DeepSeek Chat Free | 5M free tokens for new users |
| DeepSeek | DeepSeek R1 (Legacy) | $0.55 | $2.19 | 64K | RETIRED Jul 24 2026 along with V3.2 and the deepseek-chat / deepseek-reasoner aliases — reasoning now served by V4.1 Flash thinking mode | DeepSeek Chat Free | Free with rate limits |
| xAI | Grok 4.6 | $2.00 | $6.00 | 500K | Current xAI flagship (Aug 12 2026, on Bedrock since Aug 19); cached input $0.50; ≥200K prompt bills the whole request at $4 in / $12 out; long-running agents, coding, visual work | SuperGrok $30/mo | Limited on X |
| xAI | Grok 4.3 | $1.25 | $2.50 | 1M | Budget tier; cached input $0.20; >200K rate $2.50/$5.00; Grok 4.1 Fast sits below at $0.20/$0.50. Legacy Grok 3 was $3/$15 | SuperGrok $30/mo | Limited on X |
| Meta | Muse Spark 1.3 | $1.25 | $4.25 | 1M | Meta's current frontier line (Muse Spark replaced Llama as the flagship in Apr 2026); standard tier $1.25/$4.25 unchanged from 1.2; ~20% fewer tool calls and ~25% fewer tokens than 1.2; CONTRIBUTOR TIER $0.10/$0.20 (~$0.10/1M blended — cheapest model in the top-5 leaderboard) but Meta may train on your prompts and outputs; max-reasoning mode held back at launch | N/A (API + Meta AI app) | Limited free on Meta AI |
| Meta | Llama 4 Maverick (open weights) | $0.20 | $0.60 | 1M | Still the cheap hosted open-weight option ($0.20 via Vercel, $0.15-$0.35 across 7 providers; DeepInfra $0.15/$0.60); no longer Meta's frontier model — see Muse Spark 1.3; Intelligence Index 14.5 (45th pct) shows how far the frontier has moved | N/A (open model) | Free to self-host |
| Mistral | Mistral Large 3 | $0.50 | $1.50 | 262K | Open-weight flagship; cached input $0.05; vision + function calling; best price-to-quality in the mid tier; Medium 3.5 above it at $1.50/$7.50, Small 4 below at $0.10/$0.30 | Le Chat Pro $14.99/mo | Le Chat Free limited |
| Alibaba | Qwen3.8-Max | $2.00 | $6.00 | 1M | Alibaba flagship (GA Aug 3 2026); flat rate across the full 1M context (no >200K surcharge); cached input $0.25; 2.4T MoE with 1M ctx; third-party hosts undercut it (AIHubMix $1.69/$5.07 on the 0902 snapshot) | Token Plan discount | 1M free tokens |
| Alibaba | Qwen3.8-Flash | $0.15 | $0.47 | 262K (→1M) | Ultra-cheap multimodal MoE (live Aug 27 2026, QwenCloud); cache hit $0.016; NO peak surcharge; 262K native context extensible to 1M; open weights ship separately as Qwen3.8-Flash-Next (Qwen Community 1.0 license) | Token Plan discount | 1M free tokens |
| Alibaba | Qwen3.7-Max | $1.25 | $3.75 | 1M | Previous flagship, still cheapest route to Max-class quality — 50% promo off $2.50/$7.50 list; frontier coding agent, 35h autonomous runs | Token Plan discount | 1M free tokens |
| Alibaba | Qwen-Flash | $0.05 | $0.40 | 256K | Budget tier that REPLACED Qwen-Turbo (Turbo no longer updated); $0.05/$0.40 up to 256K, rising to $0.25/$2.00 on longer inputs; mid-tier Qwen-Plus starts at $0.40/$1.20 (aggregator-reported rates — check Alibaba's own rate card before committing) | Token Plan discount | 1M free tokens |
| Zhipu (Z.ai) | GLM-5.3-Flash | $0.15 | $0.50 | 1M | PRICE ROSE Sep 9 2026 — the 50% launch promo ($0.075/$0.25) ended at 24:00 UTC+8, so z.ai now bills list: $0.15 in / $0.50 out / $0.03 cached (storage limited-time free); 320B MoE (18B active), native vision + video, MIT weights, 1M ctx; AA Index 57 ≈ Gemini 3.8 Flash at medium effort for ~1/6 the price; cheapest top-tier blend on the board at $0.17/1M | GLM Coding Plan | 20M tokens on signup |
| Zhipu (Z.ai) | GLM-5.3 | $1.40 | $4.40 | 1M | Text flagship, same price as GLM-5.2/5.1; cached input $0.26; open weights expected mid-to-late Sep 2026; CyberGym 84.5% SOTA with 2,436 real vulnerabilities found; thinking cannot be disabled | GLM Coding Plan | 20M tokens on signup |
| ByteDance | Doubao Seed 2.1 Pro | $0.88 | $4.41 | 256K | Volcano Ark rate card (Jun 24 2026): ¥6 in / ¥30 out, cache hit ¥1.2 (~$0.18); Turbo is exactly half (¥3/¥15 ≈ $0.44/$2.21); text+image+video in; USD figures are conversions at 6.80 CNY — not a published dollar price card; supersedes the old Doubao 1.5 Pro (32K, ¥0.8/¥2) | Free tier on Doubao app | Free quotas on Volcano Engine |
| Moonshot | Kimi K3 | $3.00 | $15.00 | 1M | Most powerful open-weights model by GPQA (93.5%), released Jul 16 2026 with weights in late Jul under Modified MIT; 2.8T MoE with native vision; FLAT $3/$15 across the full 1M window — no long-context surcharge; cache hit $0.30; coding sibling K2.7 Code $0.95/$4.00 (cache $0.19); Batch not available on K3 | Kimi free app; $1 min top-up on API | No free API tier |
| 01.AI | Yi-Lightning | $0.14 | $0.14 | 128K | UNVERIFIED — earlier data claimed this API was shutting down Sep 3 2026 with refunds to Dec 3; no such notice could be found on 01.ai, in the 2026 deprecation trackers, or in the press as of Sep 12 2026. What IS verifiable: 01.AI stopped pre-training large models in Mar 2025 and now sells enterprise/agent solutions (万策/万智 platforms), so Yi-Lightning API availability is doubtful but NOT confirmed discontinued. Treat this row as legacy/unconfirmed | N/A | Free tier (historical) |
| Inception Labs | Mercury 2 | $0.25 | $0.75 | 128K | Diffusion (non-autoregressive) decoder, not a transformer-with-next-token — 1,009+ tok/s on Blackwell; cached input $0.025; native tool use, schema-aligned JSON, OpenAI-compatible (switch is a base-URL change); live since Mar 4 2026, #18/134 on AA intelligence; Mercury 2.5 Preview listed at $0.04/$0.15 (tracker-reported) | N/A | 10M free tokens on signup |
| Sakana AI | Fugu Max | $2.00 | $6.00 | — | NEW Sep 11 2026 — multi-agent ORCHESTRATION product rather than a single model (OpenAI-compatible API); cost tier of the Fugu line, undercutting Sonnet 5 / GPT-5.6 Terra / Kimi K3 output by 40-60% at the same $6 output as Qwen3.8-Max and Grok 4.6; capability twin Fugu Ultra v2.0 is $5/$30 (cached $0.50; >272K input $10/$45/$1.00) — the price war has moved from per-model inference to orchestration economics | N/A | No free tier reported |