As of September 23, 2026, across the 13 models compared here, a million input tokens costs anywhere from $0.10 (OpenAI’s GPT-6 Luna) to $10 (Claude Fable 5.1 and GPT-6 Astra), and a million output tokens from $0.50 to $50. That is a 100-fold spread between the cheapest and the most expensive models.
LLM API pricing is only half the story, though. What you actually pay depends on how many tokens a task uses, how much of each prompt is served from cache, and whether the job can wait for a batch discount. This guide covers all three, with prices from each vendor’s official pricing page.
- The priciest models here cost $10 per million input tokens and $50 per million output tokens; budget models cost 10 to 100 times less.
- Output tokens cost 3 to 6 times as much as input tokens, and hidden reasoning tokens are billed as output.
- Cached input costs 75% to 98% less, and batch jobs cost half price at Anthropic, OpenAI and Google.
- A chatbot reply costs between a small fraction of a cent and about 3.5 cents, depending on the model.
- Estimate with your own token counts, then recheck the vendor pages, because prices change often.
LLM API pricing at a glance (September 2026)
Prices are in US dollars per million tokens, for standard requests under each vendor’s long-prompt threshold. We checked every number on the official pricing pages on September 23, 2026.
| Model | Input | Cached input | Output | Good to know |
|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $0.25 | $50 | For demanding reasoning and agent work |
| Claude Opus 5.5 | $4 | $0.20 | $20 | Anthropic’s suggested default |
| Claude Sonnet 5 | $2 | $0.20 | $10 | Launch price made permanent |
| Claude Haiku 4.5 | $1 | $0.10 | $5 | 200K context window |
| GPT-6 Astra | $10 | $1.00 | $50 | Prompts over 272K tokens cost more |
| GPT-6 Sol | $2 | $0.20 | $10 | Built for coding and agent work |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | Cheapest model in this table |
| Gemini 3.1 Pro (preview) | $2 | $0.20 | $12 | $4 and $18 for prompts over 200K |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 | Rises to $1.50 and $7.50 on January 1, 2027 |
| Gemini 3.1 Flash-Lite | $0.25 | $0.025 | $1.50 | Audio input costs $0.50 |
| Grok 4.7 | $2 | $0.50 | $6 | $4 and $12 for prompts of 200K or more |
| DeepSeek V4-Pro | $1.32 | $0.044 | $3.96 | Half price at off-peak hours |
| DeepSeek V4.1 Flash | $0.30 | $0.006 | $1.20 | Half price at off-peak hours |
Two warnings before you compare rows. A token is not the same size everywhere: Anthropic says its current tokenizer produces about 30% more tokens for the same text than its previous one. And OpenAI and Anthropic charge extra to write to the cache, covered below. If tokens themselves are new to you, start with our guide to tokens and context windows.
Why output costs more than input
Reading your prompt is fast, because the model processes all of it at once. Writing is slow, because it generates one token at a time. Vendors price that difference in: across this table, output costs 3 to 6 times as much as input.
Reasoning makes the gap bigger. When a model thinks before it answers, those hidden thinking tokens are billed as output. Google’s pricing page says so outright (“including thinking tokens”), and xAI lists reasoning tokens as a billed type. A short answer after a long think can cost more than a long answer without one.
The discounts that change the math
Every vendor discounts repeated input and patient jobs, but the rules differ. This is where two models with the same list price can end up with very different bills:
| Anthropic | OpenAI | xAI | DeepSeek | ||
|---|---|---|---|---|---|
| Cached input | 10% of input price (5% on Opus 5.5, 2.5% on Fable 5.1) | 10% | 10% | 25% on Grok 4.7 | About 2% to 3% |
| Writing to cache | 1.25x for 5 minutes, 2x for 1 hour | 1.25x | Free and automatic; explicit caches add hourly storage | Free and automatic | Free and automatic |
| Batch jobs | 50% off | 50% off, plus a Flex tier at 50% | 50% off, plus a Flex tier at 50% | 20% off older models, none on Grok 4.7 | No batch price listed; 50% off at off-peak hours |
| Very long prompts | Same price up to 1M tokens | 2x input and 1.5x output past 272K | Pro: 2x input and 1.5x output past 200K | 2x past 200K | Same price |
| Faster tier | Fast mode on Opus 5.5, 2x | Fast mode, 2x | Priority, 1.8x | Priority, 2x | None listed |
DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays other than Chinese public holidays; everything else, including weekends, is off-peak. Caching only kicks in above a minimum prompt size, such as 1,024 tokens for OpenAI’s newest models and 4,096 for Gemini 3.8 Flash and 3.1 Pro. Our guide to cutting your AI bill shows how to structure prompts so the cache actually hits.
What typical tasks cost, worked out
Here are three everyday jobs, costed with the prices above. The token counts are our assumptions, chosen to be typical rather than extreme:
A chatbot reply: 2,000 input tokens (instructions, history and a help article) and 300 output tokens, no caching. Shown per 1,000 replies.
Sorting reviews overnight: 100,000 product reviews at 400 input tokens and 10 output tokens each.
A long coding-agent session: 2.5 million input tokens across many calls, 90% of them cache reads, plus 50,000 output tokens. New input is billed at the cache-write rate where one applies.
| Model | 1,000 chat replies | 100,000 reviews | Same, batch or off-peak | Agent session, no cache | Agent session, cached |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $35.00 | $450.00 | $225.00 | $27.50 | $6.19 |
| Claude Opus 5.5 | $14.00 | $180.00 | $90.00 | $11.00 | $2.70 |
| Claude Sonnet 5 | $7.00 | $90.00 | $45.00 | $5.50 | $1.58 |
| Claude Haiku 4.5 | $3.50 | $45.00 | $22.50 | $2.75 | $0.79 |
| GPT-6 Astra | $35.00 | $450.00 | $225.00 | $27.50 | $7.88 |
| GPT-6 Sol | $7.00 | $90.00 | $45.00 | $5.50 | $1.58 |
| GPT-6 Luna | $0.35 | $4.50 | $2.25 | $0.28 | $0.08 |
| Gemini 3.1 Pro | $7.60 | $92.00 | $46.00 | $5.60 | $1.55 |
| Gemini 3.8 Flash | $2.63 | $33.75 | $16.88 | $2.06 | $0.54 |
| Gemini 3.1 Flash-Lite | $0.95 | $11.50 | $5.75 | $0.70 | $0.19 |
| Grok 4.7 | $5.80 | $86.00 | $86.00 | $5.30 | $1.93 |
| DeepSeek V4-Pro | $3.83 | $56.76 | $28.38 | $3.50 | $0.63 |
| DeepSeek V4.1 Flash | $0.96 | $13.20 | $6.60 | $0.81 | $0.15 |
Three lessons stand out. A single chat reply costs between 0.035 cents and 3.5 cents, so model choice matters more than prompt trimming at small scale. Batch discounts halve bulk jobs at most vendors, but not on Grok 4.7. And caching cuts the agent session by 64% to 82%, a big part of why long agent sessions stay affordable.
Hidden costs that are not on the price list
These extras show up on real bills too:
Tool definitions. Every tool you offer the model adds input tokens to each request. On Claude, enabling tool use adds a few hundred tokens of system prompt, and the computer use toolset adds about 4,500.
Search. Web search costs $10 per 1,000 searches at Anthropic and OpenAI and $5 per 1,000 at xAI. Google includes 5,000 grounded searches a month, then charges $14 per 1,000.
Long prompts. Crossing 200,000 or 272,000 tokens can double the input price for the whole request on some models.
Data location. Keeping processing in one region adds about 10% at Anthropic (US-only), OpenAI (regional endpoints) and xAI (its US endpoint).
Growing conversations. Each turn resends the history, so a 30-turn chat costs far more than 30 separate questions. Caching softens this.
How to estimate your own bill
You can get within a reasonable margin in an hour. Here is the method:
Measure real token counts
Send 20 typical requests and read the input, cached and output token counts that every API returns. Real averages beat any rule of thumb.
Multiply by the price
Cost per request is input tokens times the input price, plus output tokens times the output price, divided by one million. Add cached tokens at the cached price.
Add overhead
Add 10% to 20% for retries, tool definitions and odd long requests until your logs tell you the real figure.
Apply the discounts you will really use
Batch only what can wait up to a day. Count on cache hits only for prefixes that repeat within the cache lifetime.
Set a hard limit
Set a monthly spend limit and an alert in the vendor’s console before launch, not after the first surprise.
The arithmetic fits in a few lines. This version handles caching too:
# Prices are dollars per million tokens. Token counts are per request.
def monthly_cost(requests, input_tokens, output_tokens, input_price, output_price,
cached_share=0.0, cached_price=0.0):
fresh = input_tokens * (1 - cached_share) * input_price
cached = input_tokens * cached_share * cached_price
output = output_tokens * output_price
return requests * (fresh + cached + output) / 1_000_000
# 30,000 support replies a month on a $2 in / $10 out model
print(f"No caching: ${monthly_cost(30_000, 2_000, 300, 2.00, 10.00):,.2f}")
# Same job with 80% of each prompt served from cache at $0.20
print(f"80% cached: ${monthly_cost(30_000, 2_000, 300, 2.00, 10.00, 0.8, 0.20):,.2f}")We ran it with Python 3.10 and 3.14:
No caching: $210.00
80% cached: $123.60It leaves out cache-write premiums, which the overhead allowance above covers.
FAQ
What is the cheapest LLM API?
Among the models above, GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output. DeepSeek V4.1 Flash drops to $0.15 and $0.60 at off-peak hours. The cheapest per token is not always cheapest per task, so test quality first.
Is there a free LLM API?
Google’s Gemini API has a free tier with free input and output tokens on some models, but lower rate limits, and your content may be used to improve Google’s products. The paid tier does not use your content that way.
How many words is a million tokens?
It depends on the tokenizer and the language. Anthropic puts 1 million tokens at roughly 555,000 English words on its current tokenizer, and about 750,000 on older models.
How often do API prices change?
Often. Anthropic dropped a planned September 2026 price rise for Sonnet 5, Google has scheduled Gemini 3.8 Flash to double in price on January 1, 2027, and OpenAI’s GPT-5.6 Sol is on promotional pricing through at least November 21, 2026.
- The priciest tokens here cost 100 times more than the cheapest, so match the model to the job.
- Output, including hidden reasoning, costs 3 to 6 times as much as input.
- Caching and batching are the biggest discounts, and their rules differ by vendor.
- Cost per task beats cost per token as a way to compare models.
Next, learn how to cut your AI bill without worse results, or find out when a small model beats a giant one.
- Pricing, Anthropic, accessed September 23, 2026
- Models overview, Anthropic, accessed September 23, 2026
- API pricing, OpenAI, accessed September 23, 2026
- GPT-6 Sol and GPT-6 Astra model pages, OpenAI, accessed September 23, 2026
- Prompt caching, OpenAI, accessed September 23, 2026
- Gemini Developer API pricing, Google, accessed September 23, 2026
- Context caching, Google, accessed September 23, 2026
- Pricing, xAI, accessed September 23, 2026
- Prompt caching, xAI, accessed September 23, 2026
- Models and pricing and Context caching, DeepSeek, accessed September 23, 2026




