EvoMap
Claude Usage Limits Agent: Fix Outages & Limits

Claude Usage Limits Agent: Fix Outages & Limits

April 15, 2026
635 views
claude rate-limits usage-caps outages agent-resilience circuit-breaker anthropic

Hello, Lena is here. There was a point last year when I was staring at a workflow that had been running fine for two weeks, and then suddenly it just… stopped. No useful error. No obvious reason. The task was 80% done, and the agent had gone quiet.

I spent the next few hours assuming it was my code. It wasn't. It was a Claude usage limit — but not the one I thought I'd hit.

That experience got me thinking. Most of the troubleshooting I'd seen treated "Claude is down" as a single problem with a single fix. But sitting with the logs, I started to notice that ​three completely different things were getting confused for one another​. And once I could tell them apart, the fixes became much more obvious.

This is that breakdown — the one I wish I'd had before I started.

What Actually Happens When Claude Hits a Limit or Goes Down

The first thing worth slowing down on: not all Claude failures are the same kind of failure.

Rate Limits vs Usage Caps vs Outages

These are three distinct problems with different causes, different error signatures, and different fixes. Mixing them up wastes time.

Rate limits are the per-minute constraints on requests and tokens. Your rate limit depends on your usage tier, and is measured in three key metrics: requests per minute (RPM), tokens per minute (TPM), and sometimes daily token quotas. When you exceed these, you get an HTTP 429 Too Many Requests. This is a throughput problem — you're asking for too much, too fast, within a short window.

Usage caps are different. Claude Code rate limits operate as a system of three independent, overlapping constraints, and the dashboard percentage reflects only one of them. You might look at your Anthropic console and see 6% daily usage remaining, assume everything's fine, and still hit a wall — because you've burned through the per-minute token ceiling, not the daily one. It's a subtle but important distinction.

Outages are another category entirely. A 529 Service Unavailable means Anthropic's servers are stressed at a system level. The 529 error is not your fault — it occurs when Anthropic's servers experience high traffic across all users, and rejected 529 requests don't count toward your billing. You can't fix a 529 through code optimization. You wait, you back off, and you check Anthropic's status page for incident updates.

The reason this distinction matters so much: ​the wrong diagnosis leads to the wrong fix​. I've watched people upgrade their API tier trying to fix what was actually a transient outage. And I've watched others wait patiently for a 429 to "resolve itself" when what they actually needed was a different request pattern.

How Each Failure Mode Affects Agent Workflows

A chat interface is pretty forgiving when Claude hits a limit. You see a message, you wait, you try again.

Agent workflows are not forgiving. Claude Code does not send a single prompt to the API and wait for a response — each interaction is a multi-turn conversation that includes the system prompt, the accumulated conversation history, file contents pulled into context, and tool-use tokens. A seemingly simple "edit this file" command might consume 50,000–150,000 tokens in a single API call once the full context is assembled.

Rate limit hits in this environment cause retry loops — the agent keeps firing requests, each one failing, each one burning tokens that count against your quota. Usage cap hits can stop execution mid-task, sometimes without a clear signal that it was a cap issue. And outages cause what I think of as silent failures: the tool call hangs, times out, and depending on your setup, may not log anything useful.

This is where I spent an uncomfortable few hours with that stalled workflow. The agent had hit a rate limit, entered a retry loop without backoff, and burned through my remaining token budget before I noticed.

Immediate Fixes for Each Failure Mode

Rate Limit Hits: Stop the Bleeding First

The most effective immediate fix is exponential backoff with jitter. The idea: each retry waits twice as long as the previous one, with a small random offset (jitter) to spread out burst retries from multiple clients.

The jitter part is easy to overlook, but it matters. Without it, when multiple workers or parallel agent threads all hit the same rate limit and then all retry at exactly the same interval, you recreate the burst problem at every retry cycle. You're not solving the issue — you're deferring it by a fixed offset.

A simple Python pattern that's served me well:

Python
import time, random
from anthropic import Anthropic, RateLimitError

def call_with_backoff(client, messages, max_retries=5):
    for attempt in range(max_retries):
        try:
            return client.messages.create(
                model="claude-sonnet-4-6",
                max_tokens=1024,
                messages=messages
            )
        except RateLimitError as e:
            if attempt == max_retries - 1:
                raise
            base_wait = min(2 ** attempt, 60)
            wait_time = base_wait + (random.random() * base_wait * 0.1)
            time.sleep(wait_time)

Worth noting: the official Anthropic Python SDK includes built-in retry logic by default. It retries 429 errors up to 2 times with exponential backoff. For production agent workflows with higher failure tolerance requirements, you'll usually want to configure that with a higher max_retries value and your own jitter logic.

Beyond retries, ​reduce ​parallelism​. If you have five agent threads all making Claude API calls simultaneously, and your rate limit is 50 RPM with a shared token pool, you're almost certainly going to collide. Stagger requests, batch where you can, and don't assume parallel execution scales linearly with throughput.

Usage Cap Hits: Don't Restart From Zero

A cap hit mid-task is one of the more frustrating failure modes because ​the work done before the cap hit is usually still valid​. The fix isn't to restart — it's to save state and resume.

If you're on a subscription plan and hitting caps frequently in production workflows, the Batch API allows asynchronous processing of large volumes of requests with a 50% discount on both input and output tokens — which extends your effective capacity significantly for non-latency-sensitive tasks. Prompt caching is worth implementing if you have large system prompts or repeated context: cached tokens cost a fraction of standard input tokens, and in long agent sessions where the same context is re-sent repeatedly, this adds up fast.

If the usage cap is a structural mismatch — you genuinely need more throughput than your current tier provides — verify current plan limits at docs.anthropic.com/en/api/rate-limits before making tier decisions, since these numbers change without notice.

Outage Handling: Degrade Gracefully

For outages, the core patterns are circuit breakers and graceful degradation.

Circuit breaker patterns have three states: Closed (normal operation), Open (failures detected — stop trying), and Half-Open (testing if service has recovered). Applied to a Claude API integration: after N consecutive failures, the circuit opens and stops sending requests. After a timeout period, it allows a single probe request. If that succeeds, the circuit closes again.

The key behavior difference from retries: ​retries handle individual request failures, circuit breakers handle systemic failures​. If Claude is down for 20 minutes, you don't want your agent making 400 failed API calls during that window. You want it to detect the pattern, stop attempting, queue work, and resume when the service recovers.

Building Resilience Into Your Agent Workflow

These patterns are where the interesting part is — at least for me. The immediate fixes above stop the bleeding. This section is about not bleeding in the first place.

Fallback Model Routing

When Claude is unavailable, having a fallback model endpoint means your agent can continue with reduced capability rather than stop entirely. The practical shape of this: your primary workflow routes to Claude; if you get a 529 or a circuit breaker trip, you route to a secondary model for lower-stakes subtasks while critical-path work is queued for Claude to resume.

This isn't a simple drop-in — different models have different tool call behavior, context formats, and output consistency. I'd suggest starting with a narrow fallback scope: identify the parts of your agent workflow that are genuinely model-agnostic (summarization, simple classification, format conversion) and route those to a fallback first. Keep complex reasoning and tool-heavy steps in the queue.

For multi-provider routing with proper fallback logic, Portkey's LLM gateway covers the patterns in detail — the combination of retries, fallbacks, and circuit breakers as a layered system is worth reading through if you're building for production reliability.

Checkpoint and Resume Patterns

This part I haven't fully figured out yet, if I'm being honest. But the direction is clear: agent state needs to be ​serializable​​​ at task boundaries​.

The basic shape: before a Claude API call, serialize your current agent state — task progress, completed steps, intermediate outputs — to persistent storage. If the call fails and the circuit breaker opens, you have a checkpoint to resume from rather than starting over.

The harder part is defining "task boundaries" in agentic workflows where steps are interdependent. I've found the cleanest approach is to treat each tool call as a potential checkpoint, even if that feels granular. It's easier to merge checkpoints than to untangle a partially-completed multi-step task with no recovery point.

Separating Critical Path From Background Tasks

Not all agent tasks need the same availability guarantees. Background tasks — logging, summarization, low-priority analysis — can tolerate Claude being unavailable for minutes or hours. Critical path tasks can't.

Modeling this explicitly lets you design different resilience profiles: critical path gets aggressive retry logic, fallback models, and circuit breaker monitoring; background tasks get queued with patience and no retry pressure. This alone reduces the noise from Claude usage limits significantly, because you're no longer treating every failed API call as equally urgent.

The Deeper Problem: Every Workaround Gets Rediscovered

Here's the thing that's been bothering me about this whole space.

I've talked to enough people building Claude-based workflows to notice a pattern. Someone solves the exponential backoff problem. They implement it well. It works. Then three months later, a teammate starts a new project, hits the same 429 errors, and solves it again — slightly differently, in a slightly different file, in a slightly different way. The first solution never transferred.

The same applies to checkpoint patterns, circuit breaker configurations, fallback routing logic. A global retry counter treats all tools as a single failure domain — when one tool degrades, it drains the budget for every other tool. Someone figures this out, builds per-tool circuit breakers, and it lives in one project's codebase, undocumented, unreferenced by the next project that runs into the exact same issue.

This isn't a code problem — it's a knowledge structure problem. A validated fallback strategy belongs in a reusable form that travels with the agent capability, not in a chat log or a one-off script that gets forgotten.

I don't have a complete answer to this. I keep thinking about it.

FAQ

What are Claude's current rate limits for ​API​​ users?

These change, so verify at the official Anthropic rate limits documentation before making architectural decisions. As a rough reference: limits are tier-based (Tier 1 through 4), measured in RPM, ITPM, and OTPM, with higher tiers unlocking after spending thresholds are met. Tier 1 starts conservative; Tier 4 is significantly more generous. Numbers from 2025 articles are likely already outdated.

​How do I handle Claude outages in a production agent ​workflow​?

Circuit breakers are the right pattern. Open the circuit after a threshold of consecutive failures, queue work during the open state, probe after a timeout, close when the service recovers. Check status.anthropic.com in your monitoring to distinguish a localized issue from a platform-wide incident.

What's the best fallback model when Claude is unavailable?

There's no universal answer here — it depends on what your agent is doing. For structured output tasks, most capable models handle it reasonably well. For complex tool-use chains and multi-step reasoning, fallback quality degrades more noticeably. Start by identifying the narrow slice of your workflow that's genuinely model-agnostic and route that to a fallback first.

How do I save agent state so a Claude failure doesn't restart my task from zero?

Serialize agent state at task boundaries before each Claude API call. At minimum: completed steps, intermediate outputs, current position in the task graph. The granularity question is harder — err toward more frequent checkpoints. Storage is cheap; re-running a two-hour agent task is not.

How do I stop my agent from burning tokens on failed retries?

Two things: exponential backoff with jitter (so retries don't create synchronized bursts), and circuit breakers (so systemic failures don't keep consuming retry budget). The Anthropic Python SDK has basic retry logic built in — configure it explicitly rather than relying on defaults. For production systems with parallelism, add per-tool or per-operation circuit breakers so one degraded endpoint doesn't drain your entire retry budget.

I'll probably keep watching how this evolves. The limit and resilience patterns feel like they're still being worked out in public, and I'm not sure I've gotten to the bottom of what the cleanest version looks like for long-running agentic workflows. But the distinction between the three failure modes — that part at least feels more settled.

Previous Posts:

Related Articles