On this page9 sections
- 01How much are Opus, Sonnet and Haiku per million tokens?
- 02What do older Claude models cost?
- 03How much does prompt caching save?
- 04What everyday tasks cost on the Claude API
- 05How do batch, fast mode and long prompts change the price?
- 06What do Claude’s tools cost?
- 07Is a Claude Pro subscription the same as API access?
- 08How do I keep a Claude API bill under control?
- 09FAQ
The Claude API costs $1 to $10 per million input tokens and $5 to $50 per million output tokens, depending on the model. As of September 28, 2026, Claude Haiku 4.5 costs $1 in and $5 out, Sonnet 5.5 costs $2 and $10, Opus 5.5 costs $4 and $20, and Fable 5.1 costs $10 and $50.
Few bills are paid at those list prices. Batch jobs cost half, input read back from the cache costs 90% to 97.5% less, and a prompt of a million tokens costs the same per token as a short one. This guide lists every price on Anthropic’s own pages and works out what real jobs cost.
- Current models, per million input and output tokens: Haiku 4.5 $1 and $5, Sonnet 5.5 $2 and $10, Opus 5.5 $4 and $20, Fable 5.1 $10 and $50.
- Reading from the prompt cache costs 10% of the input price, 5% on Opus 5.5 and 2.5% on Fable 5.1. Writing to it costs 1.25 or 2 times the input price.
- The Batch API halves every price, and long prompts up to 1 million tokens carry no surcharge.
- Web search costs $10 per 1,000 searches on top of tokens. Web fetch costs only the tokens it adds.
- A Pro or Max plan does not include API access. The API is billed separately, from prepaid credits.
How much are Opus, Sonnet and Haiku per million tokens?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, Sonnet 5.5 costs $2 and $10, and Haiku 4.5 costs $1 and $5. Fable 5.1, Anthropic’s most capable model open to all customers, costs $10 and $50. Checked on September 28, 2026; each model name links to its official page.
| Model | Input | Output | Cache read | Batch input / output | Context window |
|---|---|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $0.25 | $5 / $25 | 1M tokens |
| Claude Opus 5.5 | $4 | $20 | $0.20 | $2 / $10 | 1M tokens |
| Claude Sonnet 5.5 | $2 | $10 | $0.20 | $1 / $5 | 1M tokens |
| Claude Haiku 4.5 | $1 | $5 | $0.10 | $0.50 / $2.50 | 200K tokens |
Prices are in US dollars per million tokens. Anthropic’s model guide suggests starting with Opus 5.5 for most work and moving up to Fable 5.1 when Opus 5.5 at a higher effort setting still falls short. Claude Mythos 5.1 offers the same capabilities at the same price, but only to invited organizations. Our guide to choosing an AI model helps you match the tier to the task.
Two details change what a task costs at the same price. Claude 4.7 and later models use a newer tokenizer that turns the same text into about 30% more tokens than older models such as Haiku 4.5. And Opus 5.5 and Fable 5.1 always think before they answer, with those hidden tokens billed at the output price. Our guide to tokens and context windows explains both.
What do older Claude models cost?
Older Claude models stay on the price list until they retire, and most cost more than the current ones. Checked on September 28, 2026, on Anthropic’s pricing page (linked in each row):
| Model | Input | Output | Cache read |
|---|---|---|---|
| Claude Fable 5 | $10 | $50 | $1 |
| Claude Opus 5, 4.8, 4.7, 4.6 and 4.5 | $5 | $25 | $0.50 |
| Claude Sonnet 5 | $2 | $10 | $0.20 |
| Claude Sonnet 4.6 and 4.5 | $3 | $15 | $0.30 |
Sonnet 5.5 replaced Sonnet 5 as the current Sonnet model on September 28, 2026, at the same price.
If you still run Opus 5 or Opus 4.5 to 4.8, switching to Opus 5.5 cuts the input and output price by 20% and cache reads by 60%. Retired models (Opus 4.1, Opus 4, Sonnet 4 and Haiku 3.5) no longer run on Anthropic’s own API, though the page still lists them for Amazon Bedrock or Google Cloud users.
How much does prompt caching save?
Prompt caching cuts the price of repeated input by 90% on most Claude models, 95% on Opus 5.5 and 97.5% on Fable 5.1. You pay a little extra the first time a prompt is stored, then a small fraction each time a later request reads it.
Checked on September 28, 2026, on each model’s page (linked) and Anthropic’s prompt caching guide:
| Model | 5-minute cache write | 1-hour cache write | Cache read | Shortest prompt it caches |
|---|---|---|---|---|
| Fable 5.1 | $12.50 | $20 | $0.25 | 512 tokens |
| Opus 5.5 | $5 | $8 | $0.20 | 512 tokens |
| Sonnet 5.5 | $2.50 | $4 | $0.20 | 512 tokens |
| Haiku 4.5 | $1.25 | $2 | $0.10 | 4,096 tokens |
A 5-minute write costs 1.25 times the input price and pays for itself after one cache read. A 1-hour write costs twice the input price and pays off after two reads. Each read also resets the timer at no charge. The simplest way in is automatic caching: add one cache_control field to the request, and the API moves the cache point forward as a conversation grows.
Here is the effect on a real job: 10 questions about a 300,000-token document, all asked within five minutes. This short script prices it with the four token counts that every Claude API response reports in its usage block:
# Claude Sonnet 5.5 prices in dollars per million tokens, checked September 28, 2026
INPUT, CACHE_WRITE_5M, CACHE_READ, OUTPUT = 2.00, 2.50, 0.20, 10.00
def request_cost(input_tokens, cache_creation_input_tokens, cache_read_input_tokens, output_tokens):
"""Pass the four counts from the usage block of a Messages API response."""
return (input_tokens * INPUT
+ cache_creation_input_tokens * CACHE_WRITE_5M
+ cache_read_input_tokens * CACHE_READ
+ output_tokens * OUTPUT) / 1_000_000
# First question about a 300,000-token document writes it to the cache
first = request_cost(100, 300_000, 0, 2_000)
# Nine more questions within five minutes read it back
later = 9 * request_cost(100, 0, 300_000, 2_000)
print(f"10 questions with caching: ${first + later:.2f}")
print(f"10 questions without: ${10 * request_cost(300_100, 0, 0, 2_000):.2f}")We ran it with Python 3.10 and 3.14:
10 questions with caching: $1.49
10 questions without: $6.20That is 76% less on Sonnet 5.5, and the next section shows the same job on every model. Prompts shorter than the minimum in the table are simply processed at the normal price, with no error, so check the cache fields in usage to confirm you are getting hits.
What everyday tasks cost on the Claude API
We costed two jobs with a Python script on September 28, 2026, using the prices above. Before it calculates anything, the script checks every price against a saved copy of Anthropic’s pricing page. The token counts are our assumptions:
Reading receipts: 1,000 photos of 1000 by 1000 pixels, each with a 200-token instruction and a 150-token answer. Claude bills one token per 28 by 28 pixel square, so each photo is 1,296 tokens.
A long document: one question about a 300,000-token document (about 166,000 words) with a 2,000-token answer, then 10 questions within five minutes using the cache.
| Model | 1,000 receipt photos | Same, Batch API | One question, 300K-token document | 10 questions, cached |
|---|---|---|---|---|
| Claude Fable 5.1 | $22.46 | $11.23 | $3.10 | $5.44 |
| Claude Opus 5.5 | $8.98 | $4.49 | $1.24 | $2.44 |
| Claude Sonnet 5.5 | $4.49 | $2.25 | $0.62 | $1.49 |
| Claude Haiku 4.5 | $2.25 | $1.12 | Too long (200K limit) | Too long (200K limit) |
Caching matters most on the top models: 10 questions cost $31.01 without it on Fable 5.1 and $5.44 with it, because Fable 5.1 charges only 2.5% of its input price for cache reads. Haiku 4.5 cannot take the document at all, because its context window holds 200,000 tokens.
For chatbot replies, overnight review sorting and a long coding-agent session on every Claude model, set side by side with OpenAI, Google, xAI and DeepSeek, see our LLM API pricing comparison. Our script reproduces its Claude numbers exactly.
How do batch, fast mode and long prompts change the price?
The Batch API halves both input and output prices if you can wait for the answer. Most batches finish within an hour, and requests still unprocessed after 24 hours expire without being billed. One batch holds up to 100,000 requests, and the discount stacks with caching.
Fast mode goes the other way. It runs Opus 5.5 up to 2.5 times faster for double the price, $8 in and $40 out, with caching discounts applied on top. It is a research preview on Anthropic’s own API only: you request access, and it cannot be combined with batches. In our script, fast mode doubles the comparison’s cached coding-agent session on Opus 5.5, from $2.70 to $5.40.
Long prompts cost nothing extra. Claude 4.6 and later models bill a 900,000-token request at the same rate per token as a 9,000-token one, up to the 1 million token context window.
Two smaller multipliers apply in special cases. Pinning inference to the US adds 10% to every token price on Claude 4.6 and later. On Amazon Bedrock and Google Cloud the cloud sets the price, and their regional endpoints cost 10% more than global ones for Claude 4.5 and later models.
What do Claude’s tools cost?
Most Claude tools cost only the tokens they add. The two exceptions are web search, at $10 per 1,000 searches, and code execution time beyond a free monthly allowance. Checked on September 28, 2026, on Anthropic’s pricing page:
Any tool. Turning on tool use adds a hidden instruction block of 286 tokens on Opus 5.5 and Sonnet 5.5 and 496 on Haiku 4.5, plus your own tool descriptions.
Web search. $10 per 1,000 searches, a cent each, plus the results billed as input tokens. A search that fails is not billed.
Web fetch. No fee beyond tokens. Anthropic estimates an average 10 kB web page at about 2,500 tokens.
Code execution. Free when the request also uses web search or web fetch. Otherwise each organization gets 1,550 free hours a month, then pays $0.05 per container hour, with a five-minute minimum.
Computer use and browser use. The toolsets add about 4,500 and 6,600 input tokens per request, and every screenshot is billed as an image.
Managed Agents. $0.08 per session-hour while an agent is running, plus tokens.
Is a Claude Pro subscription the same as API access?
No. A paid Claude plan covers Claude in the web, desktop and mobile apps and in Claude Code, but Anthropic’s help center says it “doesn’t include access to the Claude API or Console.” The API is a separate product with separate billing. Checked on September 28, 2026, on Anthropic’s plans page (linked in each row):
| Plan | Price | What you get |
|---|---|---|
| Free | $0 | Chat on web, desktop and mobile |
| Pro | $17 a month billed yearly ($200), or $20 monthly | More usage, plus Claude Code |
| Max | From $100 a month | 5 or 20 times the usage of Pro |
| Team | $20 or $100 a seat billed yearly ($25 or $125 monthly) | Standard or premium seats, 2 to 150 people |
| Enterprise | $20 a seat a month billed yearly, plus usage at API rates | Admin and security controls |
Plans come with usage limits that reset every five hours, and paid plans add weekly caps. When you hit them, paid plans can keep going with usage credits billed at API rates. Our guide to AI usage limits explains how those windows work.
The API works the other way round. You buy prepaid credits in the Claude Console and pay per token; credits expire a year after purchase and are non-refundable. For scale, $20 of credit buys about 2,857 short chatbot replies on Sonnet 5.5 (2,000 tokens in, 300 out), 5,714 on Haiku 4.5 or 1,428 on Opus 5.5. To compare Claude’s plans with other AI subscriptions, see AI subscription prices in 2026.
How do I keep a Claude API bill under control?
Set a monthly spend limit in the Claude Console before you launch. Each standard tier has a monthly cap, from $500 on the Start tier to $200,000 on the Scale tier, and you can set your own lower limit. Count the tokens in a few real requests first: Anthropic’s token counting endpoint is free to use, subject to rate limits.
Then work the levers. Cache what you resend, batch what can wait a day, and try a lower effort setting before you switch models; Anthropic’s model guide says effort is often the better lever. Our guide to cutting your AI bill covers each one, and keeping API keys safe matters too, because a leaked key spends your credits.
FAQ
Is there a free tier for the Claude API?
No. New users get a small amount of free credit to test the API, and after that you prepay. A few things cost nothing extra: counting tokens, the web fetch tool beyond its tokens, and the first 1,550 hours a month of code execution.
Does Claude charge for thinking tokens?
Yes. Thinking is billed as output tokens, even when the thinking text is not returned to you. Opus 5.5 and Fable 5.1 always think, so use the effort setting to control how much.
Do I pay for failed requests?
No. Anthropic bills only successful calls. One catch: if your client disconnects or times out partway through a request that would have succeeded, that request is still charged.
Is Claude cheaper on Amazon Bedrock or Google Cloud?
Those clouds set their own prices, so check their pricing pages. Anthropic notes that their regional endpoints cost 10% more than their global endpoints for Claude 4.5 and later models.
- Claude API prices run from $1 to $10 per million input tokens and $5 to $50 per million output tokens.
- Opus 5.5 costs 20% less than Opus 5 and is Anthropic’s suggested starting point.
- Caching cuts repeated input by 90% or more, and batching halves everything.
- Long prompts carry no surcharge, but thinking tokens count as output.
- Claude subscriptions and the Claude API are billed separately.
Next, read the same breakdown for the OpenAI API and the Gemini API.
- Pricing, Anthropic, accessed September 28, 2026
- Models overview, Anthropic, accessed September 28, 2026
- Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 4.5 model pages, Anthropic, accessed September 28, 2026
- Choosing the right model, Anthropic, accessed September 28, 2026
- Prompt caching and Batch processing, Anthropic, accessed September 28, 2026
- Fast mode (research preview), Anthropic, accessed September 28, 2026
- Vision and Thinking, Anthropic, accessed September 28, 2026
- Token counting and Rate limits, Anthropic, accessed September 28, 2026
- Claude plans and pricing, Anthropic, accessed September 28, 2026
- Why do I have to pay separately to use the Claude API and Console?, Claude Help Center, March 2026
- How do I pay for my Claude API usage?, Claude Help Center, August 2026




