On this page11 sections
- 01How much does each Gemini model cost per million tokens?
- 02Is the Gemini API free?
- 03What are the Gemini API free tier limits?
- 04What happens to Gemini prices in January 2027?
- 05How does context caching work on the Gemini API?
- 06Batch, Flex and Priority: what do they cost?
- 07What does grounding with Google Search cost?
- 08What do images, video and music cost?
- 09Is Google AI Pro the same as the Gemini API?
- 10How do I keep a Gemini API bill under control?
- 11FAQ
The Gemini API costs $0.25 to $2 per million input tokens and $1.50 to $12 per million output tokens on Google’s current Gemini 3 text models, and most of them also have a free tier. As of September 28, 2026, Gemini 3.8 Flash costs $0.75 in and $3.75 out, and Gemini 3.1 Pro Preview costs $2 and $12 for prompts up to 200,000 tokens.
Two details shape the bill. Gemini 3.8, 3.7 and 3.6 Flash double in price on January 1, 2027. And the free tier really is free, but Google may use what you send it to improve its products. This guide covers every price on Google’s pricing page and works out real jobs with a script.
- Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens until December 31, 2026, then $1.50 and $7.50.
- The cheapest current model, Gemini 3.1 Flash-Lite, costs $0.25 and $1.50. Gemini 3.1 Pro Preview costs $2 and $12, or $4 and $18 above 200K tokens.
- The free tier covers most Flash and Flash-Lite models at no cost, with lower rate limits, and your prompts may be used to improve Google’s products.
- Batch and Flex cost half, cached input costs 90% less, and Priority costs 1.8 times the standard price.
- Grounding with Google Search is free for 5,000 queries a month on Gemini 3 models, then $14 per 1,000.
How much does each Gemini model cost per million tokens?
Gemini 3.8 Flash, which Google recommends for new projects along with 3.5 Flash-Lite, costs $0.75 per million input tokens and $3.75 per million output tokens on the paid tier. Output prices include the model’s thinking tokens. Checked on September 28, 2026; each model name links to its row on Google’s pricing page.
| Model | Input | Output | Cached input | Free tier |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | Yes |
| Gemini 3.7 Flash | $0.75 | $3.75 | $0.075 | Yes |
| Gemini 3.6 Flash | $0.75 | $3.75 | $0.075 | Yes |
| Gemini 3.5 Flash | $1.50 | $9 | $0.15 | Yes |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.03 | Yes |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $0.025 | Yes |
| Gemini 3.1 Pro Preview | $2 ($4 over 200K) | $12 ($18 over 200K) | $0.20 ($0.40 over 200K) | No |
Prices are in US dollars per million tokens, and all seven models read up to about 1 million input tokens. Gemini 3.1 Flash-Lite charges $0.50 for audio input. Because 3.8, 3.7 and 3.6 Flash cost the same, the newest one costs nothing extra. Our guide to choosing an AI model helps you pick between Flash, Flash-Lite and Pro.
Older and preview models are still listed. Gemini 3 Flash Preview costs $0.50 and $3. The Gemini 2.5 models (2.5 Pro at $1.25 and $10, 2.5 Flash at $0.30 and $2.50, 2.5 Flash-Lite at $0.10 and $0.40) are now limited to users who have used them before. Gemma 4 runs on the free tier only.
Is the Gemini API free?
Yes, for most models: the free tier charges nothing for input or output tokens on the Flash and Flash-Lite models, and Google AI Studio itself is free to use. The catch is the trade you make for it:
- Free input and output tokens on Gemini 3.8, 3.7, 3.6 and 3.5 Flash and both Flash-Lite models
- Free use of Google AI Studio in every supported region
- Free code execution and URL context on models that support them
- Your prompts and responses may be used to improve Google’s products
- No Gemini 3.1 Pro Preview, no image, video or music models
- No Search grounding on Gemini 3 models, and Batch or Flex only on the two Flash-Lite models
- Lower rate limits than any paid tier
Moving to the paid tier means linking a billing account and, for most new users, prepaying at least $5. From then on, Google does not use your prompts to improve its products. Keep sensitive or customer data off the free tier.
What are the Gemini API free tier limits?
Google’s docs do not print the free tier’s numbers model by model: your exact limits appear on the rate limit page in Google AI Studio. They are counted in requests per minute, input tokens per minute and requests per day, per project rather than per API key, and the daily count resets at midnight Pacific time. Preview and experimental models get tighter limits.
Paid limits rise through three tiers, each with a monthly spend cap. Checked on September 28, 2026, on Google’s rate limits and billing pages (linked in each row):
The 10-minute spend limits apply depending on your billing history, and hitting one returns a 429 error until you wait or slow down. Batch jobs have separate limits, including up to 100 batches running at once.
What happens to Gemini prices in January 2027?
On January 1, 2027, Gemini 3.8, 3.7 and 3.6 Flash double in price: input goes from $0.75 to $1.50 per million tokens and output from $3.75 to $7.50. Cached input rises from $0.075 to $0.15, and the explicit cache storage fee doubles too. The other models on the list show no scheduled change.
We priced three everyday jobs with a Python script on September 28, 2026. Before it calculates anything, the script checks every price against a saved copy of Google’s pricing page. It uses the same token counts as our LLM API pricing comparison, which already covers 3.1 Pro and 3.1 Flash-Lite, so this table adds the 2027 price and the 3.5 models:
| Model | 1,000 chatbot replies | 100,000 reviews, Batch | Coding-agent session, 90% cache hits |
|---|---|---|---|
| Gemini 3.8 Flash, 2026 price | $2.63 | $16.88 | $0.54 |
| Gemini 3.8 Flash, from January 1, 2027 | $5.25 | $33.75 | $1.09 |
| Gemini 3.5 Flash | $5.70 | $34.50 | $1.16 |
| Gemini 3.5 Flash-Lite | $1.35 | $7.25 | $0.27 |
A chatbot reply here is 2,000 tokens in and 300 out, a review is 400 in and 10 out, and the session is 2.5 million input tokens plus 50,000 output. Google does not guarantee implicit cache hits, so the last column is a best case. Even after the rise, 3.8 Flash stays cheaper than 3.5 Flash on all three jobs.
How does context caching work on the Gemini API?
Gemini caches repeated input in two ways. Implicit caching is on by default and passes on savings automatically, but hits are not guaranteed. Explicit caching guarantees the discount on content you store yourself, and charges an hourly storage fee for as long as you keep it.
Cached input costs 10% of the normal input price on every model above. Caching starts at 4,096 tokens on Gemini 3.5 to 3.8 Flash and 3.1 Pro Preview, and an explicit cache lasts one hour unless you set another time.
Storage costs $0.50 per million tokens per hour on 3.8 Flash until the end of 2026, $1 on 3.5 Flash and the Flash-Lite models, and $4.50 on 3.1 Pro. Keeping a 100,000-token cache for an hour costs $0.05 on 3.8 Flash and $0.45 on 3.1 Pro.
Here is the effect on one job, from the same script run on September 28, 2026: questions about a 300,000-token document with a 2,000-token answer each. For the cached column we stored the document for one hour. Google’s caching guide does not say what creating a cache costs, so to be safe we also counted one full-price pass to create it.
| Model | One question, 300K-token document | 10 questions, explicit cache for an hour |
|---|---|---|
| Gemini 3.8 Flash | $0.23 | $0.68 |
| Gemini 3.5 Flash-Lite | $0.10 | $0.53 |
| Gemini 3.1 Flash-Lite | $0.08 | $0.48 |
| Gemini 3.1 Pro Preview | $1.24 | $4.11 |
Without the cache, 10 questions cost $2.33 on 3.8 Flash and $12.36 on 3.1 Pro, so caching saves 71% and 67%. On the Flash-Lite models it saves only about 40%, because the storage fee is large next to their low input price. Watch the 200K line on 3.1 Pro as well: in our script a 190,000-token document costs $0.40 and a 210,000-token one $0.88.
Batch, Flex and Priority: what do they cost?
Batch and Flex both cost half the standard price, and Priority costs 1.8 times as much. Checked on September 28, 2026, on Google’s pricing page and its guides (linked in each row):
Batch fits overnight work such as tagging or summarizing thousands of records. Flex fits background tasks that still need an answer in the same call. Priority suits business-critical traffic, and its default rate limits are 0.3 times the standard ones. Our guide to cutting your AI bill shows how to combine these with caching.
What does grounding with Google Search cost?
On Gemini 3 models, Grounding with Google Search is free for the first 5,000 search queries a month, shared across all Gemini 3 models, then costs $14 per 1,000 queries. You pay per query the model decides to run, so one prompt that triggers two searches counts twice.
Google’s pricing page says the retrieved search results are not charged as input tokens, though it states this in its image model and agent notes rather than in the text model tables.
In our script, 10,000 grounded answers on 3.8 Flash, each running two searches with a 300-token prompt and a 500-token answer, cost $21 in tokens (counting none for the search results) and $210 in searches. Searches, not tokens, dominate a grounded app’s bill.
The other tools, checked on September 28, 2026, on the same page:
Gemini 2.5 models: Search grounding is free for 1,500 prompts a day, then $35 per 1,000 grounded prompts.
Grounding with Google Maps: free for 5,000 prompts a month on Gemini 3 models, then $14 per 1,000.
URL context: the fetched page is billed as input tokens.
Code execution: billed only as tokens, with no charge for runtime.
File search: $0.15 per million tokens to index your files, plus retrieved text billed as input tokens.
What do images, video and music cost?
Google’s image, video and music models are paid-only, priced per image, second or song. Checked on September 28, 2026, on Google’s pricing page:
Nano Banana 2 (Gemini 3.1 Flash Image): $0.067 per 1K image, from $0.045 at 512 pixels to $0.151 at 4K. Batch halves it.
Nano Banana 2 Lite: $0.0336 per 1K image. Nano Banana Pro: $0.134 per 1K or 2K image and $0.24 at 4K.
Veo 3.1 video: $0.40 a second at 720p or 1080p, $0.10 to $0.30 on Veo 3.1 Fast and $0.05 to $0.08 on Veo 3.1 Lite. You pay only for videos that generate successfully.
Gemini Omni Flash video: about $0.10 a second at 720p.
Lyria 3.5 music: $0.08 per song.
Gemini Embedding 2: $0.20 per million text tokens on the paid tier, and unlike the models above it also has a free tier.
In our script, 1,000 product pictures at 1K cost $67 on Nano Banana 2, $33.60 on Nano Banana 2 Lite and $134 on Nano Banana Pro, or half that through Batch.
Is Google AI Pro the same as the Gemini API?
No. Google AI Plus, Pro and Ultra are subscriptions for the Gemini app and other Google products, while the Gemini API has its own prepaid billing in AI Studio.
The Pro and Ultra plans do include a monthly developer credit, which Google’s Developer Program page says you can use to start building in AI Studio. Google’s Gemini API billing docs do not say whether it pays for API usage, and they note that not all Google Cloud credits do.
Checked on September 28, 2026, on Google’s plans page (linked in each row) and its Developer Program page for the credits:
| Plan | Price | Monthly developer credit |
|---|---|---|
| Free | $0 | None |
| Google AI Plus | $4.99 a month | None listed |
| Google AI Pro | $19.99 a month | $10 |
| Google AI Ultra | $99.99 a month (5 times Pro’s usage) | $40 |
| Google AI Ultra, 20x | $199.99 a month | $100 |
On a prepaid account you must buy API credit before eligible Google Cloud credits apply. Prepaid credit starts at $5, expires after 12 months and is non-refundable. For scale, $20 covers about 7,619 of our example chatbot replies on 3.8 Flash at today’s price, or 3,809 from January 2027. For every consumer AI plan in one table, see AI subscription prices in 2026.
How do I keep a Gemini API bill under control?
Set a monthly spend cap on each project on AI Studio’s Spend page. Billing data can lag by about 10 minutes, so a burst of traffic, a batch job or an agent can run past the cap before it takes effect.
Prepay also acts as a backstop: when the balance reaches $0, every API key on that billing account stops with an HTTP 402 error until you add credit. Build and test on the free tier, but only with data you are happy for Google to use in improving its products. Our guide to keeping API keys safe covers the other common way bills run away, and tokens and context windows explains what you are counting.
FAQ
Does Google use my data on the Gemini API free tier?
Yes, it may. Google’s pricing page marks free-tier content as used to improve its products, and paid-tier content as not used.
Can I use the $300 Google Cloud free trial credit on the Gemini API?
Not on newer accounts. If your Cloud billing account was opened after March 2, 2026, the welcome credit does not cover Gemini API or AI Studio usage.
Are thinking tokens billed on Gemini?
Yes. Google’s output prices include thinking tokens, so a model that thinks longer costs more per answer.
Is Gemini cheaper through Google Cloud?
Not necessarily. Google says prices on its Gemini Enterprise Agent Platform may differ from the Gemini API prices, so compare both pages for your models.
- Gemini 3.8 Flash costs $0.75 in and $3.75 out per million tokens until the end of 2026, then twice that.
- The free tier is real, but Google may use what you send it.
- Batch and Flex halve the price, and cached input costs 90% less.
- Search grounding is free up to 5,000 queries a month, then $14 per 1,000.
- Google AI subscriptions and the Gemini API are billed separately.
Next, read the same breakdown for the Claude API and the OpenAI API.
- Gemini Developer API pricing, Google, accessed September 28, 2026
- Gemini models, Google, accessed September 28, 2026
- Gemini 3.8 Flash and Gemini 3.1 Pro Preview model pages, Google, accessed September 28, 2026
- Rate limits and Billing, Google, accessed September 28, 2026
- Context caching and explicit caching, Google, accessed September 28, 2026
- Batch API, Flex inference and Priority inference, Google, accessed September 28, 2026
- Grounding with Google Search, Google, accessed September 28, 2026
- Gemini subscriptions, Google, accessed September 28, 2026
- Google AI plans, Google One, accessed September 28, 2026
- Google Developer Program plans and pricing, Google, accessed September 28, 2026




