On this page9 sections
- 01What is an AI usage limit?
- 02How does the 5-hour limit work?
- 03How do weekly limits work in Claude Code and Codex?
- 04AI usage limits by tool
- 05Why did I hit my limit so fast?
- 06How can I make my usage limit last longer?
- 07What happens when you hit your limit?
- 08Is it cheaper to pay for extra usage or to upgrade?
- 09FAQ
AI usage limits are a budget of computing work that refills on a timer. Claude, ChatGPT’s Codex and Google’s Gemini refill it every 5 hours and can also stop you at a weekly ceiling, while Cursor and GitHub Copilot hand out a monthly allowance. When the budget runs out, you wait, switch to a lighter model or pay for more.
So when your tool tells you to come back in a few hours, you have used up the 5-hour window. If the reset is days away, you have usually hit the weekly cap. Below are each vendor’s own rules as of September 2026, why agents drain limits so fast, and how to make an allowance last.
- A usage limit measures work, not messages. Long chats, big files, bigger models, tools and images all spend it faster.
- Claude, Codex and Gemini refill every 5 hours and can also apply a weekly cap. Cursor and GitHub Copilot reset monthly.
- Agents drain limits because every step re-sends the whole session. In our model of a 30-turn session, clearing it every 10 turns cut the tokens processed by 59%.
- When you run out, you can wait, switch to a lighter model, pay per use with credits or an API key, or upgrade.
- Vendors mostly publish multiples such as “5x Pro,” not absolute numbers. OpenAI’s Codex estimates are the main exception.
What is an AI usage limit?
An AI usage limit caps how much work a model does for you over a period of time, and it is measured in compute, not in messages. Anthropic says its plans have “no fixed message count,” and Google calls Gemini’s limits “compute-based,” weighing your prompt’s complexity, the features you use and the length of your chat.
In practice, a short question barely dents your allowance. A long chat with attached files, web searches and the biggest model uses far more of it.
Tokens sit underneath all three: chunks of text about three quarters of a word long. Our guide to tokens and context windows shows how they add up. This guide is about the first kind.
How does the 5-hour limit work?
The 5-hour limit is a short allowance that refills on its own every five hours, so one heavy morning cannot use up your whole week. Anthropic calls it a rolling five-hour session, and Settings > Usage shows how much of the current session you have used and when it resets.
Gemini’s limit “refreshes every 5 hours until you reach your weekly limit,” and OpenAI states Codex allowances per five-hour period. OpenAI is also the only vendor that puts numbers on the window.
Checked on September 28, 2026. All rows come from OpenAI’s Codex pricing page: estimated local Codex messages per 5 hours.
| Model | Plus, $20 | Pro, $100 | Pro, $200 |
|---|---|---|---|
| GPT-6 Astra | 5 to 45 | 25 to 225 | 100 to 900 |
| GPT-6 Sol | 15 to 150 | 70 to 700 | 300 to 3,000 |
| GPT-6 Luna | 350 to 3,000 | 1,750 to 14,000 | 7,000 to 56,000 |
The ranges span roughly tenfold because the work varies. OpenAI says model choice, context, reasoning, tool use, retrieval and caching all change what a message costs, and that cloud chats may use more of the allowance than local messages.
How do weekly limits work in Claude Code and Codex?
A weekly limit is a second, bigger ceiling that stops sustained heavy use even while the 5-hour windows keep refilling.
At Anthropic, Pro and Max plans have a weekly limit that applies across all models. It resets at a fixed day and time assigned to your account, which does not move when you start working. Claude on the web, desktop and mobile and Claude Code all draw from the same pool, so a long chat leaves less for coding.
Max plans may spend up to half of the weekly limit on Anthropic’s Fable models, while Pro users pay for Fable with usage credits. Anthropic also reserves the right to cap usage “in other ways, such as weekly and monthly caps,” to manage capacity.
OpenAI says only that weekly limits “may also apply” to Codex, and that local messages, cloud chats and ChatGPT Work share one allowance. Google applies a weekly limit to both the Gemini app and its Antigravity coding agent.
AI usage limits by tool
Most major tools meter work in one of two ways: a 5-hour window with a weekly cap, or a monthly allowance.
Checked on September 28, 2026. Each tool name links to the page we read. The last row is Cyborb, a desktop AI agent from Orbioom and our product, which meters usage over a 5-hour window and a weekly rolling window that both reset on their own.
| Tool | Short window | Longer cap | When you run out |
|---|---|---|---|
| Claude | Rolling 5 hours | Weekly on paid plans, at a fixed reset time | Wait, upgrade, or usage credits at API rates |
| ChatGPT and Codex | Per 5 hours | Weekly limits may apply | Current turn finishes; buy credits or use an API key |
| Gemini app | Refreshes every 5 hours | Weekly limit | Subscribers continue on Flash-Lite; others wait or upgrade |
| Google Antigravity | 5 hours on AI Pro and Ultra | Weekly; the free tier refreshes weekly only | AI credits on Pro and Ultra, or wait |
| Cursor | None | Monthly, in two pools | On-demand usage at API rates, or upgrade |
| GitHub Copilot | None | Monthly AI Credits | Paid credits within a budget, a cheaper model, or wait |
| Devin Desktop | Daily refresh | Weekly refresh | Extra usage at API pricing |
| Kiro | None | Monthly credits | Add-on credits at $0.04 each |
| Cyborb (ours) | Rolling 5 hours | Rolling weekly | Wait for the reset or move up a plan |
Plans on the same ladder differ by multiples. Claude Max gives 5x or 20x Pro’s per-session usage, ChatGPT Pro 5x or 20x Plus, and Google AI Ultra 5x or 20x AI Pro. Our AI coding plans comparison lines up what each step costs.
Why did I hit my limit so fast?
Because an AI agent re-sends its whole working history on every step, so each turn costs more than the one before. Anthropic’s Claude Code guide lists what every turn sends: the conversation so far, your project context (instructions and every file the agent has read) and your new prompt. The conversation “grows the fastest.”
To see how fast that adds up, we modeled a 30-turn coding session in Python. A 6,000-token project context rides along on every turn, and each turn adds 4,000 tokens of file reads, diffs and replies. Real sessions vary, but the shape does not:
Part 1: input tokens processed in a 30-turn session
one session, never cleared : 2,040,000
cleared every 10 turns : 840,000
saving : 59%
turn 30 alone sends : 126,000 tokens; turn 1 sends 10,000Caching softens the curve, since Anthropic says cached content such as project files counts less against your limits. The history still grows with every turn, though. Clearing has a cost too: the agent may need to reread a few files, so real savings are smaller than the model’s.
Four other things drain an allowance quickly:
Bigger models. Anthropic says Opus “costs several times more per turn than Sonnet.” On OpenAI’s Codex rate card, the same message costs 5 times as much on GPT-6 Astra as on GPT-6 Sol, and 20 times as much on Sol as on GPT-6 Luna.
Speed and images. Codex fast mode uses 2.5 times the standard credit rate, and image generation uses included limits 3 to 5 times faster on average.
Tools and connectors. OpenAI says every MCP server adds context to your messages, and Anthropic calls tools and connectors “token-intensive.”
Reasoning effort. Higher effort settings use more tokens, so Anthropic suggests a lower level for routine tasks.
Here is the model comparison, using OpenAI’s published credits per million tokens:
Part 2: one message with 10,000 new input, 50,000 cached input, 3,000 output tokens
GPT-6 Astra 7.500 credits
GPT-6 Sol 1.500 credits
GPT-6 Luna 0.075 credits
Astra / Sol = 5x
Sol / Luna = 20x
GPT-6 Sol in fast mode (2.5x the Standard rate) = 3.750 creditsHow can I make my usage limit last longer?
Keep sessions short, use the smallest model that does the job, and stop sending context the model does not need. Every tip below comes from Anthropic’s or OpenAI’s own guidance.
Before you change anything, list each file you plan to edit and what you will change in it, one line per file. Then wait for my go-ahead before writing any code.
Anthropic also suggests timing intensive work around your 5-hour windows, so a big job starts with a fresh allowance.
What happens when you hit your limit?
You have five choices: wait for the reset, switch to a lighter model, pay per use, use an API key, or upgrade.
Wait. Claude, Codex and Copilot show when your limit resets. OpenAI says its support team does not reset ChatGPT or Codex limits. Anthropic occasionally gives eligible plans a free “limit reset” that refills a session or weekly limit on demand.
Switch models. Gemini subscribers can continue on Flash-Lite, Copilot suggests lighter models, and on Claude Max other models keep working after Fable’s share runs out.
Pay per use. Claude usage credits bill at standard API rates, with prepaid bundles 10% to 30% off. OpenAI’s credits last 12 months, Google sells AI credits for Antigravity and Flow, and Cursor bills on-demand use at API rates. Copilot credits cost $0.01 each and Kiro’s $0.04.
Use an API key. Codex and Claude Code can run on an API key billed per token, with no plan cap, only your bill.
Upgrade. Each step up buys a multiple of the allowance.
Is it cheaper to pay for extra usage or to upgrade?
As a rule of thumb, if you run out once or twice a month, paying for the overflow tends to cost less than moving from a $20 plan to a $100 one. If you run out most days, upgrade.
Cursor’s docs offer a yardstick: daily Agent users typically spend $60 to $100 a month in total usage, and power users running several agents often spend $200 or more. For more ways to trim the bill, see our guide to cutting AI costs, and for per-token rates, our LLM API pricing comparison.
FAQ
Why did I hit my Claude limit so fast?
Long sessions, large files, Opus or Fable instead of Sonnet, and tools such as Research all spend the allowance faster, and Claude chat and Claude Code share one pool. Settings > Usage shows whether you hit the 5-hour or the weekly limit.
Do Claude Code and the Claude app share usage limits?
Yes. Anthropic says usage across claude.ai, Claude Desktop and Claude Code counts toward the same limit, including Claude Code in VS Code and JetBrains IDEs.
Do unused AI usage limits roll over?
No. An unused window simply refills, and monthly allowances at Cursor and Copilot reset each cycle. Bought credits are different: Claude usage credits usually do not expire, OpenAI’s last 12 months, and Kiro’s add-on credits last 12 months.
Can support reset my usage limit?
OpenAI says its support does not reset ChatGPT or Codex limits. Anthropic occasionally gives eligible plans a one-time limit reset, which you redeem in Settings > Usage on the web or in Claude Desktop.
Is a usage limit the same as an API rate limit?
No. Plan usage limits cap a subscription over hours, weeks or months. API rate limits cap requests per minute for developers who pay per token, and paying through an API key has no plan cap.
Read next: what $20, $100 and $200 AI coding plans buy, what every AI subscription costs, and which AI model to use for each job.
- How do usage and length limits work?, Anthropic Help Center, accessed September 2026
- Usage limit best practices, Anthropic Help Center, September 2026
- What is the Pro plan?, Anthropic Help Center, September 2026
- What is the Max plan?, Anthropic Help Center, September 2026
- Models, usage, and limits in Claude Code, Anthropic Help Center, September 2026
- Use Claude Code with your Pro or Max plan, Anthropic Help Center, August 2026
- Claude Fable models on your plan, Anthropic Help Center, accessed September 2026
- Manage usage credits for paid Claude plans, Anthropic Help Center, September 2026
- Buy usage bundles, Anthropic Help Center, May 2026
- What is a limit reset?, Anthropic Help Center, September 2026
- Claude plans and pricing, Anthropic, accessed September 2026
- Codex pricing, OpenAI, accessed September 2026
- About ChatGPT Pro tiers, OpenAI Help Center, September 2026
- Using credits for flexible usage in ChatGPT, OpenAI Help Center, September 2026
- Gemini Apps limits and upgrades for Google AI subscribers, Google, accessed September 2026
- Manage your AI credits with Google One, Google, accessed September 2026
- Plans and AI credits, Google Antigravity Docs, accessed September 2026
- Models and pricing, Cursor Docs, accessed September 2026
- GitHub Copilot plans and pricing, GitHub, accessed September 2026
- Plans and pricing, Cognition, accessed September 2026
- Kiro pricing, Amazon Web Services, accessed September 2026
- Cyborb pricing, Orbioom, accessed September 2026




