Home/LLM API pricing

LLM API Pricing: Every Model per 1M Tokens

LLM APIs cost $0.10 to $10.00 per million input tokens and $0.50 to $50.00 per million output tokens across the 26 current models below. One typical prompt (500 tokens in, 500 out) costs 0.03¢ on GPT-6 Luna up to 3¢ on Claude Fable 5.1. Official prices, verified October 2026.

Current Models

Sorted by cost per typical prompt, cheapest first. Tap a column header to sort by it; tap again to reverse. Prices in US dollars per million tokens.

Current LLM API prices per 1M tokens
ModelProviderLong-context rate
GPT-6 LunaOpenAI$0.10$0.50$0.01−50%1M0.03¢272K+ input: $0.20 / $0.75
Mistral Small 4Mistral$0.15$0.60$0.015−50%262K0.037¢—
gpt-oss-120BOpenAI (via Together)$0.15$0.60——131K0.037¢—
GPT-5.6 LunaOpenAI$0.20$1.20$0.02−50%1M0.07¢272K+ input: $0.40 / $1.80
DeepSeek V4.1 Flash3DeepSeek$0.30$1.20$0.006—1M0.075¢—
Muse Glimmer 30BMeta (via Together)$0.35$1.50$0.04—131K0.092¢—
Mistral Large 3Mistral$0.50$1.50$0.05−50%262K0.1¢—
Gemini 3.5 Flash-LiteGoogle$0.30$2.50$0.03−50%1M0.14¢—
Grok 4.3xAI$1.25$2.50$0.20−20%1M0.19¢200K+ input: $2.50 / $5.00
Gemini 3.8 Flash2Google$0.75$3.75$0.075−50%1M0.22¢—
DeepSeek V4 Pro4DeepSeek$1.32$3.96$0.044—1M0.26¢—
Muse Spark 1.3Meta$1.25$4.25$0.15—1M0.27¢—
Claude Haiku 4.5Anthropic$1.00$5.00$0.10−50%200K0.3¢—
Grok 4.7xAI$2.00$6.00$0.50—500K0.4¢200K+ input: $4.00 / $12.00
Qwen3.8 MaxAlibaba (Qwen)$2.00$6.00——1M0.4¢—
Mistral Medium 3.5Mistral$1.50$7.50$0.15−50%262K0.45¢—
GPT-6.1 SolOpenAI$2.00$10.00$0.10−50%1M0.6¢272K+ input: $4.00 / $15.00
GPT-6 SolOpenAI$2.00$10.00$0.20−50%1M0.6¢272K+ input: $4.00 / $15.00
Claude Sonnet 5.5Anthropic$2.00$10.00$0.20−50%1M0.6¢—
GPT-5.6 TerraOpenAI$2.00$12.00$0.20−50%1M0.7¢272K+ input: $4.00 / $18.00
Gemini 3.1 ProGoogle$2.00$12.00$0.20−50%1M0.7¢200K+ input: $4.00 / $18.00
Kimi K3Moonshot AI$3.00$15.00$0.30—1M0.9¢—
GPT-5.6 Sol1OpenAI$4.00$20.00$0.40−50%1M1.2¢272K+ input: $8.00 / $30.00
Claude Opus 5.5Anthropic$4.00$20.00$0.20−50%1M1.2¢—
GPT-6 AstraOpenAI$10.00$50.00$1.00−50%1M3¢272K+ input: $20.00 / $75.00
Claude Fable 5.1Anthropic$10.00$50.00$0.25−50%1M3¢—

Older Models Still Available

Superseded models you can still call, including those with an announced shutdown date.

Older LLM API prices per 1M tokens
ModelProviderLong-context rate
GPT-4.1 NanoShuts down October 23, 2026OpenAI$0.10$0.40$0.025−50%1M0.025¢—
GPT-4o MiniOlderOpenAI$0.15$0.60$0.075−50%128K0.037¢—
GPT-4.1 MiniOlderOpenAI$0.40$1.60$0.10−50%1M0.1¢—
Gemini 2.5 FlashOlderGoogle$0.30$2.50$0.03−50%1M0.14¢—
Grok 4.20OlderxAI$1.25$2.50$0.20−20%1M0.19¢200K+ input: $2.50 / $5.00
o4-miniShuts down October 23, 2026OpenAI$1.10$4.40$0.275−50%200K0.27¢—
GPT-4.1OlderOpenAI$2.00$8.00$0.50−50%1M0.5¢—
o3Shuts down December 11, 2026OpenAI$2.00$8.00$0.50−50%200K0.5¢—
GPT-5Shuts down December 11, 2026OpenAI$1.25$10.00$0.125−50%400K0.56¢—
Gemini 2.5 ProOlderGoogle$1.25$10.00$0.125−50%1M0.56¢200K+ input: $2.50 / $15.00
Claude Sonnet 5OlderAnthropic$2.00$10.00$0.20−50%1M0.6¢—
GPT-4oOlderOpenAI$2.50$10.00$1.25−50%128K0.63¢—
Claude Sonnet 4.6OlderAnthropic$3.00$15.00$0.30−50%1M0.9¢—
Claude Opus 4.6OlderAnthropic$5.00$25.00$0.50−50%1M1.5¢—

Retired, no longer on the provider’s API (last published rates on their pages): Claude Opus 4, Claude Sonnet 4, Gemini 2.0 Flash, DeepSeek V3, DeepSeek R1, Grok 4.1 Fast, Llama 4 Maverick.

How to Read This Table

  • Per prompt = 500 input + 500 output tokens at standard real-time rates, shown in cents. Hover or long-press for the exact dollar value.
  • Cached inputis the rate for prompt tokens served from the provider’s cache. Cache-write surcharges are not included.
  • Batch discountapplies to input and output sent through the provider’s asynchronous Batch API. “—” means no batch pricing.
  • Long-context rate: once the prompt reaches the threshold, the whole request is billed at the higher input / output rate.
  • 1 GPT-5.6 Sol: Promotional price, available at least through November 21, 2026.
  • 2 Gemini 3.8 Flash: Promotional price through December 31, 2026. From January 1, 2027 it rises to $1.50 input / $7.50 output per 1M tokens.
  • 3 DeepSeek V4.1 Flash: Peak-hour price shown. DeepSeek charges 50% less off-peak — every hour except 01:00–04:00 and 06:00–10:00 UTC on weekdays (about 79% of the week). Off-peak: $0.15 input / $0.60 output per 1M tokens.
  • 4 DeepSeek V4 Pro: Peak-hour price shown. DeepSeek charges 50% less off-peak — every hour except 01:00–04:00 and 06:00–10:00 UTC on weekdays (about 79% of the week). Off-peak: $0.66 input / $1.98 output per 1M tokens. DeepSeek first announced V4 Pro would be phased out after V4.1 Flash launched, then confirmed it keeps serving V4 Pro at these prices until further notice.

Want two models side by side? Browse all price comparisons. Pricing your own text? Count its tokens or use the prompt cost calculator.

LLM API Pricing — FAQ

Across the 26 current models tracked here, input tokens cost $0.10 to $10.00 per million and output tokens $0.50 to $50.00 per million (official prices, verified October 2026). One typical prompt of 500 input and 500 output tokens costs 0.03¢ on GPT-6 Luna up to 3¢ on Claude Fable 5.1. A mid-priced model such as Grok 4.7 costs 0.4¢.

By the cost of one typical prompt, the cheapest current model is GPT-6 Luna (OpenAI) at $0.10 input / $0.50 output per 1M tokens, $0.000300 per prompt. The next cheapest are Mistral Small 4 ($0.000375) and gpt-oss-120B ($0.000375).

The priciest current model per prompt is Claude Fable 5.1 (Anthropic) at $10.00 input / $50.00 output per 1M tokens, $0.0300 per prompt. That is about 100 times the cost of GPT-6 Luna.

Output tokens cost 2× to 8.3× the input price across current models, 5× at the median. Long answers therefore raise the bill faster than long prompts.

24 of 26 current models list a cached-input price, 75% to 98% below their normal input price. It applies only to the part of the prompt the provider has cached, such as a repeated system prompt. Cache-write surcharges, where a provider charges them, are not included in the table.

18 of 26 current models have Batch API pricing, at 20% and 50% off input and output, from OpenAI, Anthropic, Google, xAI and Mistral. Batch requests are processed asynchronously instead of in real time.

GPT-6 Astra, GPT-6.1 Sol, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna accept up to 1,050,000 tokens (1M) of prompt and answer combined, the most among current models.

10 current models charge a higher rate for the whole request once the prompt reaches a size threshold: GPT-6 Astra from 272K input tokens ($20.00 / $75.00); GPT-6.1 Sol from 272K input tokens ($4.00 / $15.00); GPT-6 Sol from 272K input tokens ($4.00 / $15.00); GPT-6 Luna from 272K input tokens ($0.20 / $0.75); GPT-5.6 Sol from 272K input tokens ($8.00 / $30.00); GPT-5.6 Terra from 272K input tokens ($4.00 / $18.00); GPT-5.6 Luna from 272K input tokens ($0.40 / $1.80); Gemini 3.1 Pro from 200K input tokens ($4.00 / $18.00); Grok 4.7 from 200K input tokens ($4.00 / $12.00); Grok 4.3 from 200K input tokens ($2.50 / $5.00).