How much does the Gemini API cost? Free tier included.

Every price from Google’s own pages, what the free tier really covers, the January 2027 price rise, and real jobs costed out with a script.

A vending machine standing in a quiet hallway
Photo by Petr on Unsplashdithered by Cyborb

The Gemini API costs $0.25 to $2 per million input tokens and $1.50 to $12 per million output tokens on Google’s current Gemini 3 text models, and most of them also have a free tier. As of September 28, 2026, Gemini 3.8 Flash costs $0.75 in and $3.75 out, and Gemini 3.1 Pro Preview costs $2 and $12 for prompts up to 200,000 tokens.

Two details shape the bill. Gemini 3.8, 3.7 and 3.6 Flash double in price on January 1, 2027. And the free tier really is free, but Google may use what you send it to improve its products. This guide covers every price on Google’s pricing page and works out real jobs with a script.

The short version
  • Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens until December 31, 2026, then $1.50 and $7.50.
  • The cheapest current model, Gemini 3.1 Flash-Lite, costs $0.25 and $1.50. Gemini 3.1 Pro Preview costs $2 and $12, or $4 and $18 above 200K tokens.
  • The free tier covers most Flash and Flash-Lite models at no cost, with lower rate limits, and your prompts may be used to improve Google’s products.
  • Batch and Flex cost half, cached input costs 90% less, and Priority costs 1.8 times the standard price.
  • Grounding with Google Search is free for 5,000 queries a month on Gemini 3 models, then $14 per 1,000.

How much does each Gemini model cost per million tokens?

Gemini 3.8 Flash, which Google recommends for new projects along with 3.5 Flash-Lite, costs $0.75 per million input tokens and $3.75 per million output tokens on the paid tier. Output prices include the model’s thinking tokens. Checked on September 28, 2026; each model name links to its row on Google’s pricing page.

ModelInputOutputCached inputFree tier
Gemini 3.8 Flash$0.75$3.75$0.075Yes
Gemini 3.7 Flash$0.75$3.75$0.075Yes
Gemini 3.6 Flash$0.75$3.75$0.075Yes
Gemini 3.5 Flash$1.50$9$0.15Yes
Gemini 3.5 Flash-Lite$0.30$2.50$0.03Yes
Gemini 3.1 Flash-Lite$0.25$1.50$0.025Yes
Gemini 3.1 Pro Preview$2 ($4 over 200K)$12 ($18 over 200K)$0.20 ($0.40 over 200K)No

Prices are in US dollars per million tokens, and all seven models read up to about 1 million input tokens. Gemini 3.1 Flash-Lite charges $0.50 for audio input. Because 3.8, 3.7 and 3.6 Flash cost the same, the newest one costs nothing extra. Our guide to choosing an AI model helps you pick between Flash, Flash-Lite and Pro.

Older and preview models are still listed. Gemini 3 Flash Preview costs $0.50 and $3. The Gemini 2.5 models (2.5 Pro at $1.25 and $10, 2.5 Flash at $0.30 and $2.50, 2.5 Flash-Lite at $0.10 and $0.40) are now limited to users who have used them before. Gemma 4 runs on the free tier only.

Is the Gemini API free?

Yes, for most models: the free tier charges nothing for input or output tokens on the Flash and Flash-Lite models, and Google AI Studio itself is free to use. The catch is the trade you make for it:

What the free tier gives you
  • Free input and output tokens on Gemini 3.8, 3.7, 3.6 and 3.5 Flash and both Flash-Lite models
  • Free use of Google AI Studio in every supported region
  • Free code execution and URL context on models that support them
What it costs you
  • Your prompts and responses may be used to improve Google’s products
  • No Gemini 3.1 Pro Preview, no image, video or music models
  • No Search grounding on Gemini 3 models, and Batch or Flex only on the two Flash-Lite models
  • Lower rate limits than any paid tier

Moving to the paid tier means linking a billing account and, for most new users, prepaying at least $5. From then on, Google does not use your prompts to improve its products. Keep sensitive or customer data off the free tier.

What are the Gemini API free tier limits?

Google’s docs do not print the free tier’s numbers model by model: your exact limits appear on the rate limit page in Google AI Studio. They are counted in requests per minute, input tokens per minute and requests per day, per project rather than per API key, and the daily count resets at midnight Pacific time. Preview and experimental models get tighter limits.

Paid limits rise through three tiers, each with a monthly spend cap. Checked on September 28, 2026, on Google’s rate limits and billing pages (linked in each row):

TierHow you qualifyMonthly spend capSpend limit per 10 minutes
FreeAn active project or free trialNoneNone
Tier 1Link a billing account and prepay$250$10
Tier 2$100 paid, 3 days after the first payment$2,000$50
Tier 3$1,000 paid, 30 days after the first payment$20,000 to over $100,000$200

The 10-minute spend limits apply depending on your billing history, and hitting one returns a 429 error until you wait or slow down. Batch jobs have separate limits, including up to 100 batches running at once.

What happens to Gemini prices in January 2027?

On January 1, 2027, Gemini 3.8, 3.7 and 3.6 Flash double in price: input goes from $0.75 to $1.50 per million tokens and output from $3.75 to $7.50. Cached input rises from $0.075 to $0.15, and the explicit cache storage fee doubles too. The other models on the list show no scheduled change.

We priced three everyday jobs with a Python script on September 28, 2026. Before it calculates anything, the script checks every price against a saved copy of Google’s pricing page. It uses the same token counts as our LLM API pricing comparison, which already covers 3.1 Pro and 3.1 Flash-Lite, so this table adds the 2027 price and the 3.5 models:

Model1,000 chatbot replies100,000 reviews, BatchCoding-agent session, 90% cache hits
Gemini 3.8 Flash, 2026 price$2.63$16.88$0.54
Gemini 3.8 Flash, from January 1, 2027$5.25$33.75$1.09
Gemini 3.5 Flash$5.70$34.50$1.16
Gemini 3.5 Flash-Lite$1.35$7.25$0.27

A chatbot reply here is 2,000 tokens in and 300 out, a review is 400 in and 10 out, and the session is 2.5 million input tokens plus 50,000 output. Google does not guarantee implicit cache hits, so the last column is a best case. Even after the rise, 3.8 Flash stays cheaper than 3.5 Flash on all three jobs.

How does context caching work on the Gemini API?

Gemini caches repeated input in two ways. Implicit caching is on by default and passes on savings automatically, but hits are not guaranteed. Explicit caching guarantees the discount on content you store yourself, and charges an hourly storage fee for as long as you keep it.

Cached input costs 10% of the normal input price on every model above. Caching starts at 4,096 tokens on Gemini 3.5 to 3.8 Flash and 3.1 Pro Preview, and an explicit cache lasts one hour unless you set another time.

Storage costs $0.50 per million tokens per hour on 3.8 Flash until the end of 2026, $1 on 3.5 Flash and the Flash-Lite models, and $4.50 on 3.1 Pro. Keeping a 100,000-token cache for an hour costs $0.05 on 3.8 Flash and $0.45 on 3.1 Pro.

Here is the effect on one job, from the same script run on September 28, 2026: questions about a 300,000-token document with a 2,000-token answer each. For the cached column we stored the document for one hour. Google’s caching guide does not say what creating a cache costs, so to be safe we also counted one full-price pass to create it.

ModelOne question, 300K-token document10 questions, explicit cache for an hour
Gemini 3.8 Flash$0.23$0.68
Gemini 3.5 Flash-Lite$0.10$0.53
Gemini 3.1 Flash-Lite$0.08$0.48
Gemini 3.1 Pro Preview$1.24$4.11

Without the cache, 10 questions cost $2.33 on 3.8 Flash and $12.36 on 3.1 Pro, so caching saves 71% and 67%. On the Flash-Lite models it saves only about 40%, because the storage fee is large next to their low input price. Watch the 200K line on 3.1 Pro as well: in our script a 190,000-token document costs $0.40 and a 210,000-token one $0.88.

Batch, Flex and Priority: what do they cost?

Batch and Flex both cost half the standard price, and Priority costs 1.8 times as much. Checked on September 28, 2026, on Google’s pricing page and its guides (linked in each row):

TierPriceHow it works
StandardList priceNormal requests
Batch50% of StandardSubmit a job, results within 24 hours (Google’s target)
Flex50% of StandardA normal call with slower, best-effort service (preview)
Priority1.8 times StandardServed ahead of other traffic (preview)

Batch fits overnight work such as tagging or summarizing thousands of records. Flex fits background tasks that still need an answer in the same call. Priority suits business-critical traffic, and its default rate limits are 0.3 times the standard ones. Our guide to cutting your AI bill shows how to combine these with caching.

What does grounding with Google Search cost?

On Gemini 3 models, Grounding with Google Search is free for the first 5,000 search queries a month, shared across all Gemini 3 models, then costs $14 per 1,000 queries. You pay per query the model decides to run, so one prompt that triggers two searches counts twice.

Google’s pricing page says the retrieved search results are not charged as input tokens, though it states this in its image model and agent notes rather than in the text model tables.

In our script, 10,000 grounded answers on 3.8 Flash, each running two searches with a 300-token prompt and a 500-token answer, cost $21 in tokens (counting none for the search results) and $210 in searches. Searches, not tokens, dominate a grounded app’s bill.

The other tools, checked on September 28, 2026, on the same page:

  • Gemini 2.5 models: Search grounding is free for 1,500 prompts a day, then $35 per 1,000 grounded prompts.

  • Grounding with Google Maps: free for 5,000 prompts a month on Gemini 3 models, then $14 per 1,000.

  • URL context: the fetched page is billed as input tokens.

  • Code execution: billed only as tokens, with no charge for runtime.

  • File search: $0.15 per million tokens to index your files, plus retrieved text billed as input tokens.

What do images, video and music cost?

Google’s image, video and music models are paid-only, priced per image, second or song. Checked on September 28, 2026, on Google’s pricing page:

  • Nano Banana 2 (Gemini 3.1 Flash Image): $0.067 per 1K image, from $0.045 at 512 pixels to $0.151 at 4K. Batch halves it.

  • Nano Banana 2 Lite: $0.0336 per 1K image. Nano Banana Pro: $0.134 per 1K or 2K image and $0.24 at 4K.

  • Veo 3.1 video: $0.40 a second at 720p or 1080p, $0.10 to $0.30 on Veo 3.1 Fast and $0.05 to $0.08 on Veo 3.1 Lite. You pay only for videos that generate successfully.

  • Gemini Omni Flash video: about $0.10 a second at 720p.

  • Lyria 3.5 music: $0.08 per song.

  • Gemini Embedding 2: $0.20 per million text tokens on the paid tier, and unlike the models above it also has a free tier.

In our script, 1,000 product pictures at 1K cost $67 on Nano Banana 2, $33.60 on Nano Banana 2 Lite and $134 on Nano Banana Pro, or half that through Batch.

Is Google AI Pro the same as the Gemini API?

No. Google AI Plus, Pro and Ultra are subscriptions for the Gemini app and other Google products, while the Gemini API has its own prepaid billing in AI Studio.

The Pro and Ultra plans do include a monthly developer credit, which Google’s Developer Program page says you can use to start building in AI Studio. Google’s Gemini API billing docs do not say whether it pays for API usage, and they note that not all Google Cloud credits do.

Checked on September 28, 2026, on Google’s plans page (linked in each row) and its Developer Program page for the credits:

PlanPriceMonthly developer credit
Free$0None
Google AI Plus$4.99 a monthNone listed
Google AI Pro$19.99 a month$10
Google AI Ultra$99.99 a month (5 times Pro’s usage)$40
Google AI Ultra, 20x$199.99 a month$100

On a prepaid account you must buy API credit before eligible Google Cloud credits apply. Prepaid credit starts at $5, expires after 12 months and is non-refundable. For scale, $20 covers about 7,619 of our example chatbot replies on 3.8 Flash at today’s price, or 3,809 from January 2027. For every consumer AI plan in one table, see AI subscription prices in 2026.

How do I keep a Gemini API bill under control?

Set a monthly spend cap on each project on AI Studio’s Spend page. Billing data can lag by about 10 minutes, so a burst of traffic, a batch job or an agent can run past the cap before it takes effect.

Prepay also acts as a backstop: when the balance reaches $0, every API key on that billing account stops with an HTTP 402 error until you add credit. Build and test on the free tier, but only with data you are happy for Google to use in improving its products. Our guide to keeping API keys safe covers the other common way bills run away, and tokens and context windows explains what you are counting.

FAQ

Does Google use my data on the Gemini API free tier?

Yes, it may. Google’s pricing page marks free-tier content as used to improve its products, and paid-tier content as not used.

Can I use the $300 Google Cloud free trial credit on the Gemini API?

Not on newer accounts. If your Cloud billing account was opened after March 2, 2026, the welcome credit does not cover Gemini API or AI Studio usage.

Are thinking tokens billed on Gemini?

Yes. Google’s output prices include thinking tokens, so a model that thinks longer costs more per answer.

Is Gemini cheaper through Google Cloud?

Not necessarily. Google says prices on its Gemini Enterprise Agent Platform may differ from the Gemini API prices, so compare both pages for your models.

Key takeaways
  • Gemini 3.8 Flash costs $0.75 in and $3.75 out per million tokens until the end of 2026, then twice that.
  • The free tier is real, but Google may use what you send it.
  • Batch and Flex halve the price, and cached input costs 90% less.
  • Search grounding is free up to 5,000 queries a month, then $14 per 1,000.
  • Google AI subscriptions and the Gemini API are billed separately.

Next, read the same breakdown for the Claude API and the OpenAI API.

Sources
  1. Gemini Developer API pricing, Google, accessed September 28, 2026
  2. Gemini models, Google, accessed September 28, 2026
  3. Gemini 3.8 Flash and Gemini 3.1 Pro Preview model pages, Google, accessed September 28, 2026
  4. Rate limits and Billing, Google, accessed September 28, 2026
  5. Context caching and explicit caching, Google, accessed September 28, 2026
  6. Batch API, Flex inference and Priority inference, Google, accessed September 28, 2026
  7. Grounding with Google Search, Google, accessed September 28, 2026
  8. Gemini subscriptions, Google, accessed September 28, 2026
  9. Google AI plans, Google One, accessed September 28, 2026
  10. Google Developer Program plans and pricing, Google, accessed September 28, 2026
cyborb.ai

Stop reading about it. Build it.

Describe what you want in plain words. Cyborb plans the work, writes and runs the code, makes the assets, and puts the result online.

Download Cyborb

Free to start. No card required.