Why your AI tool says wait.

What the vendors’ own help pages say about resets, why agents drain an allowance so fast, and the habits that make it last.

A traffic light hanging from a metal pole
Photo by CARTER SAUNDERS on Unsplashdithered by Cyborb

AI usage limits are a budget of computing work that refills on a timer. Claude, ChatGPT’s Codex and Google’s Gemini refill it every 5 hours and can also stop you at a weekly ceiling, while Cursor and GitHub Copilot hand out a monthly allowance. When the budget runs out, you wait, switch to a lighter model or pay for more.

So when your tool tells you to come back in a few hours, you have used up the 5-hour window. If the reset is days away, you have usually hit the weekly cap. Below are each vendor’s own rules as of September 2026, why agents drain limits so fast, and how to make an allowance last.

The short version
  • A usage limit measures work, not messages. Long chats, big files, bigger models, tools and images all spend it faster.
  • Claude, Codex and Gemini refill every 5 hours and can also apply a weekly cap. Cursor and GitHub Copilot reset monthly.
  • Agents drain limits because every step re-sends the whole session. In our model of a 30-turn session, clearing it every 10 turns cut the tokens processed by 59%.
  • When you run out, you can wait, switch to a lighter model, pay per use with credits or an API key, or upgrade.
  • Vendors mostly publish multiples such as “5x Pro,” not absolute numbers. OpenAI’s Codex estimates are the main exception.

What is an AI usage limit?

An AI usage limit caps how much work a model does for you over a period of time, and it is measured in compute, not in messages. Anthropic says its plans have “no fixed message count,” and Google calls Gemini’s limits “compute-based,” weighing your prompt’s complexity, the features you use and the length of your chat.

In practice, a short question barely dents your allowance. A long chat with attached files, web searches and the biggest model uses far more of it.

Tokens sit underneath all three: chunks of text about three quarters of a word long. Our guide to tokens and context windows shows how they add up. This guide is about the first kind.

How does the 5-hour limit work?

The 5-hour limit is a short allowance that refills on its own every five hours, so one heavy morning cannot use up your whole week. Anthropic calls it a rolling five-hour session, and Settings > Usage shows how much of the current session you have used and when it resets.

Gemini’s limit “refreshes every 5 hours until you reach your weekly limit,” and OpenAI states Codex allowances per five-hour period. OpenAI is also the only vendor that puts numbers on the window.

Checked on September 28, 2026. All rows come from OpenAI’s Codex pricing page: estimated local Codex messages per 5 hours.

ModelPlus, $20Pro, $100Pro, $200
GPT-6 Astra5 to 4525 to 225100 to 900
GPT-6 Sol15 to 15070 to 700300 to 3,000
GPT-6 Luna350 to 3,0001,750 to 14,0007,000 to 56,000

The ranges span roughly tenfold because the work varies. OpenAI says model choice, context, reasoning, tool use, retrieval and caching all change what a message costs, and that cloud chats may use more of the allowance than local messages.

How do weekly limits work in Claude Code and Codex?

A weekly limit is a second, bigger ceiling that stops sustained heavy use even while the 5-hour windows keep refilling.

At Anthropic, Pro and Max plans have a weekly limit that applies across all models. It resets at a fixed day and time assigned to your account, which does not move when you start working. Claude on the web, desktop and mobile and Claude Code all draw from the same pool, so a long chat leaves less for coding.

Max plans may spend up to half of the weekly limit on Anthropic’s Fable models, while Pro users pay for Fable with usage credits. Anthropic also reserves the right to cap usage “in other ways, such as weekly and monthly caps,” to manage capacity.

OpenAI says only that weekly limits “may also apply” to Codex, and that local messages, cloud chats and ChatGPT Work share one allowance. Google applies a weekly limit to both the Gemini app and its Antigravity coding agent.

AI usage limits by tool

Most major tools meter work in one of two ways: a 5-hour window with a weekly cap, or a monthly allowance.

Checked on September 28, 2026. Each tool name links to the page we read. The last row is Cyborb, a desktop AI agent from Orbioom and our product, which meters usage over a 5-hour window and a weekly rolling window that both reset on their own.

ToolShort windowLonger capWhen you run out
ClaudeRolling 5 hoursWeekly on paid plans, at a fixed reset timeWait, upgrade, or usage credits at API rates
ChatGPT and CodexPer 5 hoursWeekly limits may applyCurrent turn finishes; buy credits or use an API key
Gemini appRefreshes every 5 hoursWeekly limitSubscribers continue on Flash-Lite; others wait or upgrade
Google Antigravity5 hours on AI Pro and UltraWeekly; the free tier refreshes weekly onlyAI credits on Pro and Ultra, or wait
CursorNoneMonthly, in two poolsOn-demand usage at API rates, or upgrade
GitHub CopilotNoneMonthly AI CreditsPaid credits within a budget, a cheaper model, or wait
Devin DesktopDaily refreshWeekly refreshExtra usage at API pricing
KiroNoneMonthly creditsAdd-on credits at $0.04 each
Cyborb (ours)Rolling 5 hoursRolling weeklyWait for the reset or move up a plan

Plans on the same ladder differ by multiples. Claude Max gives 5x or 20x Pro’s per-session usage, ChatGPT Pro 5x or 20x Plus, and Google AI Ultra 5x or 20x AI Pro. Our AI coding plans comparison lines up what each step costs.

Why did I hit my limit so fast?

Because an AI agent re-sends its whole working history on every step, so each turn costs more than the one before. Anthropic’s Claude Code guide lists what every turn sends: the conversation so far, your project context (instructions and every file the agent has read) and your new prompt. The conversation “grows the fastest.”

To see how fast that adds up, we modeled a 30-turn coding session in Python. A 6,000-token project context rides along on every turn, and each turn adds 4,000 tokens of file reads, diffs and replies. Real sessions vary, but the shape does not:

limit_math.py, part 1 (our model, not a vendor measurement)
Part 1: input tokens processed in a 30-turn session
  one session, never cleared :  2,040,000
  cleared every 10 turns     :    840,000
  saving                     : 59%
  turn 30 alone sends        :    126,000 tokens; turn 1 sends 10,000

Caching softens the curve, since Anthropic says cached content such as project files counts less against your limits. The history still grows with every turn, though. Clearing has a cost too: the agent may need to reread a few files, so real savings are smaller than the model’s.

Four other things drain an allowance quickly:

  • Bigger models. Anthropic says Opus “costs several times more per turn than Sonnet.” On OpenAI’s Codex rate card, the same message costs 5 times as much on GPT-6 Astra as on GPT-6 Sol, and 20 times as much on Sol as on GPT-6 Luna.

  • Speed and images. Codex fast mode uses 2.5 times the standard credit rate, and image generation uses included limits 3 to 5 times faster on average.

  • Tools and connectors. OpenAI says every MCP server adds context to your messages, and Anthropic calls tools and connectors “token-intensive.”

  • Reasoning effort. Higher effort settings use more tokens, so Anthropic suggests a lower level for routine tasks.

Here is the model comparison, using OpenAI’s published credits per million tokens:

limit_math.py, part 2 (OpenAI's Codex credit rates)
Part 2: one message with 10,000 new input, 50,000 cached input, 3,000 output tokens
  GPT-6 Astra     7.500 credits
  GPT-6 Sol       1.500 credits
  GPT-6 Luna      0.075 credits
  Astra / Sol  = 5x
  Sol / Luna   = 20x
  GPT-6 Sol in fast mode (2.5x the Standard rate) = 3.750 credits

How can I make my usage limit last longer?

Keep sessions short, use the smallest model that does the job, and stop sending context the model does not need. Every tip below comes from Anthropic’s or OpenAI’s own guidance.

Make an AI allowance last0 of 8
PromptPlan before you spend
Before you change anything, list each file you plan to edit and what you will change in it, one line per file. Then wait for my go-ahead before writing any code.

Anthropic also suggests timing intensive work around your 5-hour windows, so a big job starts with a fresh allowance.

What happens when you hit your limit?

You have five choices: wait for the reset, switch to a lighter model, pay per use, use an API key, or upgrade.

  • Wait. Claude, Codex and Copilot show when your limit resets. OpenAI says its support team does not reset ChatGPT or Codex limits. Anthropic occasionally gives eligible plans a free “limit reset” that refills a session or weekly limit on demand.

  • Switch models. Gemini subscribers can continue on Flash-Lite, Copilot suggests lighter models, and on Claude Max other models keep working after Fable’s share runs out.

  • Pay per use. Claude usage credits bill at standard API rates, with prepaid bundles 10% to 30% off. OpenAI’s credits last 12 months, Google sells AI credits for Antigravity and Flow, and Cursor bills on-demand use at API rates. Copilot credits cost $0.01 each and Kiro’s $0.04.

  • Use an API key. Codex and Claude Code can run on an API key billed per token, with no plan cap, only your bill.

  • Upgrade. Each step up buys a multiple of the allowance.

Is it cheaper to pay for extra usage or to upgrade?

As a rule of thumb, if you run out once or twice a month, paying for the overflow tends to cost less than moving from a $20 plan to a $100 one. If you run out most days, upgrade.

Cursor’s docs offer a yardstick: daily Agent users typically spend $60 to $100 a month in total usage, and power users running several agents often spend $200 or more. For more ways to trim the bill, see our guide to cutting AI costs, and for per-token rates, our LLM API pricing comparison.

FAQ

Why did I hit my Claude limit so fast?

Long sessions, large files, Opus or Fable instead of Sonnet, and tools such as Research all spend the allowance faster, and Claude chat and Claude Code share one pool. Settings > Usage shows whether you hit the 5-hour or the weekly limit.

Do Claude Code and the Claude app share usage limits?

Yes. Anthropic says usage across claude.ai, Claude Desktop and Claude Code counts toward the same limit, including Claude Code in VS Code and JetBrains IDEs.

Do unused AI usage limits roll over?

No. An unused window simply refills, and monthly allowances at Cursor and Copilot reset each cycle. Bought credits are different: Claude usage credits usually do not expire, OpenAI’s last 12 months, and Kiro’s add-on credits last 12 months.

Can support reset my usage limit?

OpenAI says its support does not reset ChatGPT or Codex limits. Anthropic occasionally gives eligible plans a one-time limit reset, which you redeem in Settings > Usage on the web or in Claude Desktop.

Is a usage limit the same as an API rate limit?

No. Plan usage limits cap a subscription over hours, weeks or months. API rate limits cap requests per minute for developers who pay per token, and paying through an API key has no plan cap.

Read next: what $20, $100 and $200 AI coding plans buy, what every AI subscription costs, and which AI model to use for each job.

Sources
  1. How do usage and length limits work?, Anthropic Help Center, accessed September 2026
  2. Usage limit best practices, Anthropic Help Center, September 2026
  3. What is the Pro plan?, Anthropic Help Center, September 2026
  4. What is the Max plan?, Anthropic Help Center, September 2026
  5. Models, usage, and limits in Claude Code, Anthropic Help Center, September 2026
  6. Use Claude Code with your Pro or Max plan, Anthropic Help Center, August 2026
  7. Claude Fable models on your plan, Anthropic Help Center, accessed September 2026
  8. Manage usage credits for paid Claude plans, Anthropic Help Center, September 2026
  9. Buy usage bundles, Anthropic Help Center, May 2026
  10. What is a limit reset?, Anthropic Help Center, September 2026
  11. Claude plans and pricing, Anthropic, accessed September 2026
  12. Codex pricing, OpenAI, accessed September 2026
  13. About ChatGPT Pro tiers, OpenAI Help Center, September 2026
  14. Using credits for flexible usage in ChatGPT, OpenAI Help Center, September 2026
  15. Gemini Apps limits and upgrades for Google AI subscribers, Google, accessed September 2026
  16. Manage your AI credits with Google One, Google, accessed September 2026
  17. Plans and AI credits, Google Antigravity Docs, accessed September 2026
  18. Models and pricing, Cursor Docs, accessed September 2026
  19. GitHub Copilot plans and pricing, GitHub, accessed September 2026
  20. Plans and pricing, Cognition, accessed September 2026
  21. Kiro pricing, Amazon Web Services, accessed September 2026
  22. Cyborb pricing, Orbioom, accessed September 2026
cyborb.ai

Stop reading about it. Build it.

Describe what you want in plain words. Cyborb plans the work, writes and runs the code, makes the assets, and puts the result online.

Download Cyborb

Free to start. No card required.