What a million tokens really costs.

Prices from five official pricing pages, the discounts that change the math, and three everyday jobs costed out model by model.

Blank price tags hanging on strings
Photo by Angèle Kamp on Unsplashdithered by Cyborb

As of September 23, 2026, across the 13 models compared here, a million input tokens costs anywhere from $0.10 (OpenAI’s GPT-6 Luna) to $10 (Claude Fable 5.1 and GPT-6 Astra), and a million output tokens from $0.50 to $50. That is a 100-fold spread between the cheapest and the most expensive models.

LLM API pricing is only half the story, though. What you actually pay depends on how many tokens a task uses, how much of each prompt is served from cache, and whether the job can wait for a batch discount. This guide covers all three, with prices from each vendor’s official pricing page.

The short version
  • The priciest models here cost $10 per million input tokens and $50 per million output tokens; budget models cost 10 to 100 times less.
  • Output tokens cost 3 to 6 times as much as input tokens, and hidden reasoning tokens are billed as output.
  • Cached input costs 75% to 98% less, and batch jobs cost half price at Anthropic, OpenAI and Google.
  • A chatbot reply costs between a small fraction of a cent and about 3.5 cents, depending on the model.
  • Estimate with your own token counts, then recheck the vendor pages, because prices change often.

LLM API pricing at a glance (September 2026)

Prices are in US dollars per million tokens, for standard requests under each vendor’s long-prompt threshold. We checked every number on the official pricing pages on September 23, 2026.

ModelInputCached inputOutputGood to know
Claude Fable 5.1$10$0.25$50For demanding reasoning and agent work
Claude Opus 5.5$4$0.20$20Anthropic’s suggested default
Claude Sonnet 5$2$0.20$10Launch price made permanent
Claude Haiku 4.5$1$0.10$5200K context window
GPT-6 Astra$10$1.00$50Prompts over 272K tokens cost more
GPT-6 Sol$2$0.20$10Built for coding and agent work
GPT-6 Luna$0.10$0.01$0.50Cheapest model in this table
Gemini 3.1 Pro (preview)$2$0.20$12$4 and $18 for prompts over 200K
Gemini 3.8 Flash$0.75$0.075$3.75Rises to $1.50 and $7.50 on January 1, 2027
Gemini 3.1 Flash-Lite$0.25$0.025$1.50Audio input costs $0.50
Grok 4.7$2$0.50$6$4 and $12 for prompts of 200K or more
DeepSeek V4-Pro$1.32$0.044$3.96Half price at off-peak hours
DeepSeek V4.1 Flash$0.30$0.006$1.20Half price at off-peak hours

Two warnings before you compare rows. A token is not the same size everywhere: Anthropic says its current tokenizer produces about 30% more tokens for the same text than its previous one. And OpenAI and Anthropic charge extra to write to the cache, covered below. If tokens themselves are new to you, start with our guide to tokens and context windows.

Why output costs more than input

Reading your prompt is fast, because the model processes all of it at once. Writing is slow, because it generates one token at a time. Vendors price that difference in: across this table, output costs 3 to 6 times as much as input.

Reasoning makes the gap bigger. When a model thinks before it answers, those hidden thinking tokens are billed as output. Google’s pricing page says so outright (“including thinking tokens”), and xAI lists reasoning tokens as a billed type. A short answer after a long think can cost more than a long answer without one.

The discounts that change the math

Every vendor discounts repeated input and patient jobs, but the rules differ. This is where two models with the same list price can end up with very different bills:

AnthropicOpenAIGooglexAIDeepSeek
Cached input10% of input price (5% on Opus 5.5, 2.5% on Fable 5.1)10%10%25% on Grok 4.7About 2% to 3%
Writing to cache1.25x for 5 minutes, 2x for 1 hour1.25xFree and automatic; explicit caches add hourly storageFree and automaticFree and automatic
Batch jobs50% off50% off, plus a Flex tier at 50%50% off, plus a Flex tier at 50%20% off older models, none on Grok 4.7No batch price listed; 50% off at off-peak hours
Very long promptsSame price up to 1M tokens2x input and 1.5x output past 272KPro: 2x input and 1.5x output past 200K2x past 200KSame price
Faster tierFast mode on Opus 5.5, 2xFast mode, 2xPriority, 1.8xPriority, 2xNone listed

DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays other than Chinese public holidays; everything else, including weekends, is off-peak. Caching only kicks in above a minimum prompt size, such as 1,024 tokens for OpenAI’s newest models and 4,096 for Gemini 3.8 Flash and 3.1 Pro. Our guide to cutting your AI bill shows how to structure prompts so the cache actually hits.

What typical tasks cost, worked out

Here are three everyday jobs, costed with the prices above. The token counts are our assumptions, chosen to be typical rather than extreme:

  • A chatbot reply: 2,000 input tokens (instructions, history and a help article) and 300 output tokens, no caching. Shown per 1,000 replies.

  • Sorting reviews overnight: 100,000 product reviews at 400 input tokens and 10 output tokens each.

  • A long coding-agent session: 2.5 million input tokens across many calls, 90% of them cache reads, plus 50,000 output tokens. New input is billed at the cache-write rate where one applies.

Model1,000 chat replies100,000 reviewsSame, batch or off-peakAgent session, no cacheAgent session, cached
Claude Fable 5.1$35.00$450.00$225.00$27.50$6.19
Claude Opus 5.5$14.00$180.00$90.00$11.00$2.70
Claude Sonnet 5$7.00$90.00$45.00$5.50$1.58
Claude Haiku 4.5$3.50$45.00$22.50$2.75$0.79
GPT-6 Astra$35.00$450.00$225.00$27.50$7.88
GPT-6 Sol$7.00$90.00$45.00$5.50$1.58
GPT-6 Luna$0.35$4.50$2.25$0.28$0.08
Gemini 3.1 Pro$7.60$92.00$46.00$5.60$1.55
Gemini 3.8 Flash$2.63$33.75$16.88$2.06$0.54
Gemini 3.1 Flash-Lite$0.95$11.50$5.75$0.70$0.19
Grok 4.7$5.80$86.00$86.00$5.30$1.93
DeepSeek V4-Pro$3.83$56.76$28.38$3.50$0.63
DeepSeek V4.1 Flash$0.96$13.20$6.60$0.81$0.15

Three lessons stand out. A single chat reply costs between 0.035 cents and 3.5 cents, so model choice matters more than prompt trimming at small scale. Batch discounts halve bulk jobs at most vendors, but not on Grok 4.7. And caching cuts the agent session by 64% to 82%, a big part of why long agent sessions stay affordable.

Hidden costs that are not on the price list

These extras show up on real bills too:

  • Tool definitions. Every tool you offer the model adds input tokens to each request. On Claude, enabling tool use adds a few hundred tokens of system prompt, and the computer use toolset adds about 4,500.

  • Search. Web search costs $10 per 1,000 searches at Anthropic and OpenAI and $5 per 1,000 at xAI. Google includes 5,000 grounded searches a month, then charges $14 per 1,000.

  • Long prompts. Crossing 200,000 or 272,000 tokens can double the input price for the whole request on some models.

  • Data location. Keeping processing in one region adds about 10% at Anthropic (US-only), OpenAI (regional endpoints) and xAI (its US endpoint).

  • Growing conversations. Each turn resends the history, so a 30-turn chat costs far more than 30 separate questions. Caching softens this.

How to estimate your own bill

You can get within a reasonable margin in an hour. Here is the method:

  1. Measure real token counts

    Send 20 typical requests and read the input, cached and output token counts that every API returns. Real averages beat any rule of thumb.

  2. Multiply by the price

    Cost per request is input tokens times the input price, plus output tokens times the output price, divided by one million. Add cached tokens at the cached price.

  3. Add overhead

    Add 10% to 20% for retries, tool definitions and odd long requests until your logs tell you the real figure.

  4. Apply the discounts you will really use

    Batch only what can wait up to a day. Count on cache hits only for prefixes that repeat within the cache lifetime.

  5. Set a hard limit

    Set a monthly spend limit and an alert in the vendor’s console before launch, not after the first surprise.

The arithmetic fits in a few lines. This version handles caching too:

estimate.py
# Prices are dollars per million tokens. Token counts are per request.
def monthly_cost(requests, input_tokens, output_tokens, input_price, output_price,
                 cached_share=0.0, cached_price=0.0):
    fresh = input_tokens * (1 - cached_share) * input_price
    cached = input_tokens * cached_share * cached_price
    output = output_tokens * output_price
    return requests * (fresh + cached + output) / 1_000_000

# 30,000 support replies a month on a $2 in / $10 out model
print(f"No caching:  ${monthly_cost(30_000, 2_000, 300, 2.00, 10.00):,.2f}")
# Same job with 80% of each prompt served from cache at $0.20
print(f"80% cached:  ${monthly_cost(30_000, 2_000, 300, 2.00, 10.00, 0.8, 0.20):,.2f}")

We ran it with Python 3.10 and 3.14:

Text
No caching:  $210.00
80% cached:  $123.60

It leaves out cache-write premiums, which the overhead allowance above covers.

FAQ

What is the cheapest LLM API?

Among the models above, GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output. DeepSeek V4.1 Flash drops to $0.15 and $0.60 at off-peak hours. The cheapest per token is not always cheapest per task, so test quality first.

Is there a free LLM API?

Google’s Gemini API has a free tier with free input and output tokens on some models, but lower rate limits, and your content may be used to improve Google’s products. The paid tier does not use your content that way.

How many words is a million tokens?

It depends on the tokenizer and the language. Anthropic puts 1 million tokens at roughly 555,000 English words on its current tokenizer, and about 750,000 on older models.

How often do API prices change?

Often. Anthropic dropped a planned September 2026 price rise for Sonnet 5, Google has scheduled Gemini 3.8 Flash to double in price on January 1, 2027, and OpenAI’s GPT-5.6 Sol is on promotional pricing through at least November 21, 2026.

Key takeaways
  • The priciest tokens here cost 100 times more than the cheapest, so match the model to the job.
  • Output, including hidden reasoning, costs 3 to 6 times as much as input.
  • Caching and batching are the biggest discounts, and their rules differ by vendor.
  • Cost per task beats cost per token as a way to compare models.

Next, learn how to cut your AI bill without worse results, or find out when a small model beats a giant one.

Sources
  1. Pricing, Anthropic, accessed September 23, 2026
  2. Models overview, Anthropic, accessed September 23, 2026
  3. API pricing, OpenAI, accessed September 23, 2026
  4. GPT-6 Sol and GPT-6 Astra model pages, OpenAI, accessed September 23, 2026
  5. Prompt caching, OpenAI, accessed September 23, 2026
  6. Gemini Developer API pricing, Google, accessed September 23, 2026
  7. Context caching, Google, accessed September 23, 2026
  8. Pricing, xAI, accessed September 23, 2026
  9. Prompt caching, xAI, accessed September 23, 2026
  10. Models and pricing and Context caching, DeepSeek, accessed September 23, 2026
cyborb.ai

Stop reading about it. Build it.

Describe what you want in plain words. Cyborb plans the work, writes and runs the code, makes the assets, and puts the result online.

Download Cyborb

Free to start. No card required.