On this page10 sections
- 01How much are the GPT models per million tokens?
- 02What do older GPT models cost?
- 03Batch, Flex or Fast mode: which should you use?
- 04How does prompt caching work on the OpenAI API?
- 05What happens to the price above 272K tokens?
- 06Is GPT-6 cheaper than GPT-5.6?
- 07What do tools, images and audio cost?
- 08Is ChatGPT Plus the same as API access?
- 09How do I keep an OpenAI API bill under control?
- 10FAQ
The OpenAI API costs $0.10 to $10 per million input tokens and $0.50 to $50 per million output tokens for its newest models, the GPT-6 family. As of September 28, 2026, GPT-6 Luna costs $0.10 in and $0.50 out, GPT-6 Sol costs $2 and $10, and GPT-6 Astra costs $10 and $50.
Discounts and surcharges move the bill a long way from those list prices. Batch and Flex requests cost half, cached input costs 90% less, Fast mode costs double, and a prompt over 272,000 tokens costs up to twice as much. This guide covers every price on OpenAI’s own pages and works out real jobs with a script.
- GPT-6 prices per million input and output tokens: Luna $0.10 and $0.50, Sol $2 and $10, Astra $10 and $50.
- GPT-6 Sol and Luna cost half or less of the GPT-5.6 models with the same names.
- Batch and Flex cost 50% of the standard price. Fast mode costs twice as much on GPT-6 and GPT-5.6 and runs up to 2.5 times faster.
- On GPT-5.6 and later, cached input costs 10% of the input price and writing to the cache costs 1.25 times.
- A prompt over 272K tokens costs twice as much for input and 1.5 times for output, on the whole request.
- ChatGPT Plus and Pro do not include API use. The API is billed separately, from prepaid credits.
How much are the GPT models per million tokens?
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, GPT-6 Sol costs $2 and $10, and GPT-6 Luna costs $0.10 and $0.50. The previous GPT-5.6 models are still sold alongside them. Checked on September 28, 2026; each model name links to its official page.
| Model | Input | Cached input | Output | Over 272K tokens, input / output |
|---|---|---|---|---|
| GPT-6 Astra | $10 | $1 | $50 | $20 / $75 |
| GPT-6 Sol | $2 | $0.20 | $10 | $4 / $15 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | $0.20 / $0.75 |
| GPT-5.6 Sol | $4 | $0.40 | $20 | $8 / $30 |
| GPT-5.6 Terra | $2 | $0.20 | $12 | $4 / $18 |
| GPT-5.6 Luna | $0.20 | $0.02 | $1.20 | $0.40 / $1.80 |
Prices are in US dollars per million tokens. All six models accept up to 922,000 input tokens and write up to 128,000. GPT-5.6 Sol is on promotional pricing until at least November 21, 2026. OpenAI’s own guide pitches Astra for the hardest work, Sol for demanding reasoning and Luna for repeatable work at scale, while its model catalog still suggests GPT-5.6 Terra when you want to balance quality and cost.
All three GPT-6 models are reasoning models: they think before answering, and those hidden reasoning tokens are billed as output. Sol and Luna can switch reasoning off with an effort setting of none; Astra’s lowest setting is low. Our guide to reasoning models explains when the extra thinking is worth paying for.
What do older GPT models cost?
Older GPT models cost anywhere from $0.05 per million input tokens (GPT-5 nano) to $150 (o1-pro, which shuts down on October 23, 2026), so the newest model is not always the cheapest. Here is a selection, checked on September 28, 2026, on OpenAI’s pricing page (standard prices, prompts under 272K tokens):
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-5.5 | $5 | $0.50 | $30 |
| GPT-5.5 Pro | $30 | Not offered | $180 |
| GPT-5.4 | $2.50 | $0.25 | $15 |
| GPT-5.4 mini | $0.75 | $0.075 | $4.50 |
| GPT-5.4 nano | $0.20 | $0.02 | $1.25 |
| GPT-5 nano | $0.05 | $0.005 | $0.40 |
| GPT-4.1 | $2 | $0.50 | $8 |
| GPT-4o mini | $0.15 | $0.075 | $0.60 |
GPT-5 nano, at $0.05 in and $0.40 out, is the cheapest text model on the list and the only one that undercuts GPT-6 Luna on both input and output. OpenAI plans to shut it down on December 11, 2026. GPT-5.5 leaves ChatGPT on October 14, 2026, but OpenAI says the API is not affected.
Batch, Flex or Fast mode: which should you use?
Use Batch for jobs that can wait up to 24 hours, Flex for slower requests you still want answered in a single call, and Fast mode only where speed earns its double price. Checked on September 28, 2026, on OpenAI’s pricing page and its Batch, Flex and Fast mode guides (linked in each row):
| Tier | Price | Speed | Good for |
|---|---|---|---|
| Standard | List price | Normal | Most apps |
| Batch | 50% of Standard | Within 24 hours, often sooner | Bulk jobs and evaluations |
| Flex | 50% of Standard | Slower, sometimes unavailable | Background work in a normal API call |
| Fast mode | 2x Standard on GPT-6 and GPT-5.6 | Up to 2.5 times faster | Features where waiting costs you users |
Batch also comes with its own, much higher rate limits. Flex is in beta for a limited set of models, including GPT-6 and GPT-5.6. Fast mode was called Priority processing until July 30, 2026, and for GPT-6 Astra it carries no latency guarantee.
Two surcharges apply to specific setups. Regional data residency adds 10% for models released on or after March 5, 2026, and on GPT-6, EU data residency works only with standard processing. FedRAMP endpoints also cost 10% more than the standard rates.
How does prompt caching work on the OpenAI API?
On GPT-5.6 and later models, reading a cached prompt costs 10% of the input price and writing one costs 125%, so a prompt reused even once costs less than sending it twice. OpenAI’s own example: one write plus one read costs 1.35 times the normal input price, against 2 times without caching, and ten requests cost 2.15 times instead of 10.
On those models, caching starts at 1,024 tokens and keeps a prompt for at least 30 minutes after its last use. It works automatically, and you can also mark exactly where the cached part should end. Models before GPT-5.6 charge nothing extra to write to the cache but each set their own read price. Our guide to cutting your AI bill shows how to order a prompt so the cache actually hits.
What happens to the price above 272K tokens?
A prompt over 272,000 input tokens is billed at twice the input and cache prices and 1.5 times the output price, for the whole request, not only the tokens past the line. That makes 272K the most expensive line in OpenAI’s pricing to cross by accident.
We checked the effect with a Python script on September 28, 2026. Before it calculates anything, it checks every price against a saved copy of OpenAI’s pricing page. The job: one question about a long document with a 2,000-token answer, then 10 questions about the larger document with the cache.
| Model | 250K-token document | 300K-token document | 10 questions on the 300K document, cached |
|---|---|---|---|
| GPT-6 Astra | $2.60 | $6.15 | $14.42 |
| GPT-6 Sol | $0.52 | $1.23 | $2.88 |
| GPT-6 Luna | $0.03 | $0.06 | $0.14 |
A document 20% longer costs 2.37 times as much. Caching still helps above the line: 10 questions cost $61.52 on Astra without it and $14.42 with it. If your prompts sit near 272K tokens, trimming or splitting them is the cheapest fix available.
Is GPT-6 cheaper than GPT-5.6?
Yes: GPT-6 Sol costs exactly half of GPT-5.6 Sol, and GPT-6 Luna costs half as much for input and 58% less for output than GPT-5.6 Luna. On September 28, 2026, our script priced three everyday jobs on both generations, using the same token counts as our LLM API pricing comparison, which has the GPT-6 figures next to other labs:
| Job | GPT-5.6 Sol | GPT-6 Sol | GPT-5.6 Luna | GPT-6 Luna |
|---|---|---|---|---|
| 1,000 chatbot replies | $14.00 | $7.00 | $0.76 | $0.35 |
| 100,000 reviews, Batch | $90.00 | $45.00 | $4.60 | $2.25 |
| Coding-agent session, cached | $3.15 | $1.58 | $0.17 | $0.08 |
A chatbot reply here is 2,000 tokens in and 300 out, a review is 400 in and 10 out, and the session is 2.5 million input tokens (90% cached) plus 50,000 output. GPT-5.6 Terra costs $7.60, $46.00 and $1.68 on the same jobs. Price is only half the decision, so test the new model on your own tasks before you switch.
What do tools, images and audio cost?
Built-in tools add a fee per call on top of the tokens they pull in. Checked on September 28, 2026, on OpenAI’s pricing page and image generation guide:
Web search. $10 per 1,000 calls, plus the search results billed as input tokens at the model’s rates. The older preview tool on non-reasoning models costs $25 per 1,000 calls, with the result tokens free.
File search. $2.50 per 1,000 calls, plus $0.10 per GB of storage per day after the first free GB.
Code Interpreter and hosted shell. From $0.03 per 20-minute session for a 1 GB container to $1.92 for 64 GB, billed by the minute with a five-minute minimum.
Images. GPT Image 2.5 and GPT Image 2 cost $8 per million image input tokens and $30 per million image output tokens. A 1024 by 1024 GPT Image 2 picture costs about $0.006 at low quality, $0.053 at medium and $0.211 at high.
Voice. GPT-Live 1 costs $0.05 a minute, billed by the second, plus the model and tools behind it. Transcription with GPT-Transcribe costs about $0.0045 a minute.
Embeddings and moderation. text-embedding-3-small costs $0.02 per million tokens, text-embedding-3-large $0.13, and the moderation model is free.
In our script, 100 product pictures from GPT Image 2 at 1024 by 1024, each with a 100-token prompt, cost $0.65 at low quality, $5.35 at medium and $21.15 at high. Fine-tuning is winding down and closed to new users.
Is ChatGPT Plus the same as API access?
No. ChatGPT and the API platform have separate billing systems, and OpenAI’s help center says “API usage is billed separately from your ChatGPT subscription.” Checked on September 28, 2026, on OpenAI’s plans page (linked in each row):
| Plan | Price | Codex on this plan |
|---|---|---|
| Free | $0 | GPT-6 Luna in the desktop app, subject to rollout |
| Go | $8 a month | GPT-6 Luna in the desktop app, subject to rollout |
| Plus | $20 a month | GPT-6 Sol and Luna on the web, CLI, IDE and iOS |
| Pro | From $100 a month | 5 or 20 times the Codex usage of Plus |
| Business | $20 per user a month billed yearly, $25 monthly | For two or more users, with admin controls |
The API runs on prepaid credits. The first purchase is at least $5, credits expire after a year, and auto-reload is switched on by default when you set up billing. Your usage tier rises with total payments, from Tier 1 after $5 to Tier 5 after $1,000, and each tier has a monthly usage limit.
For scale, $20 of credit buys about 2,857 of our example chatbot replies on GPT-6 Sol or 57,142 on GPT-6 Luna. Codex can also run on an API key, billed at API rates. For every consumer AI plan in one table, see AI subscription prices in 2026.
How do I keep an OpenAI API bill under control?
Set a spend limit on your organization or project, and turn off auto-reload if you want spending to stop when your credit runs out. Even then, OpenAI warns that a prepaid balance is not an instant cutoff: usage processed during a short delay can push it below zero, and that amount comes off your next purchase.
Then pick the cheapest tier that fits each job, keep prompts under 272K tokens, and keep a stable prompt prefix so the cache hits. A leaked key spends your credits too, so read our guide to keeping API keys safe before you ship. Our guide to tokens and context windows explains what you are counting.
FAQ
Is there a free tier for the OpenAI API?
Not for tokens. OpenAI’s pricing page lists no free allowance for its GPT models, and new API accounts prepay at least $5. The moderation model is free, and file search includes 1 GB of free storage.
Why is my output token count higher than the answer I got?
Reasoning. GPT-6 models think before they answer, and those hidden reasoning tokens are billed as output. Lower the reasoning effort, or set it to none on Sol and Luna, for simple tasks.
What is the cheapest OpenAI model?
GPT-5 nano, at $0.05 per million input tokens and $0.40 per million output tokens, until it shuts down on December 11, 2026. In the GPT-6 family, Luna is cheapest at $0.10 and $0.50.
Are OpenAI models cheaper on Amazon Bedrock?
No. OpenAI says Bedrock pricing in commercial regions matches its direct prices for the same services. Bedrock bills you through AWS instead of OpenAI.
- GPT-6 prices run from $0.10 to $10 per million input tokens and $0.50 to $50 per million output tokens.
- GPT-6 Sol and Luna cost half or less of their GPT-5.6 namesakes.
- Batch and Flex halve the price; Fast mode doubles it on GPT-6 and GPT-5.6.
- Crossing 272K tokens raises the price of the whole request.
- ChatGPT plans and the API are billed separately.
Next, read the same breakdown for the Claude API and the Gemini API.
- API pricing, OpenAI, accessed September 28, 2026
- GPT-6 Astra, GPT-6 Sol and GPT-6 Luna model pages, OpenAI, accessed September 28, 2026
- GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna model pages, OpenAI, accessed September 28, 2026
- Models and Using GPT-6, OpenAI, accessed September 28, 2026
- Prompt caching, OpenAI, accessed September 28, 2026
- Batch API, Flex processing and Fast mode, OpenAI, accessed September 28, 2026
- Reasoning models and Image generation, OpenAI, accessed September 28, 2026
- Rate limits and Deprecations, OpenAI, accessed September 28, 2026
- Pricing, ChatGPT documentation, OpenAI, accessed September 28, 2026
- Managing billing for ChatGPT and the API platform, OpenAI Help Center, accessed September 28, 2026
- Setting up and managing prepaid API billing, OpenAI Help Center, accessed September 28, 2026




