The Agentic Token Explosion in CI/CD Pipelines, Explained

Meet TrueForge: The open-source, vendor-neutral agent harness. 50% lower cost. Explore Now→

Rate Limiting AI Agents: Preventing LLM API Exhaustion

By Boyu Wang

Published: September 11, 2026

Built for Speed: ~10ms Latency, Even Under Load

Blazingly fast way to build, track and deploy your models!

Get Started with Truefoundry Now Talk to the Expert

Why this matters

The most expensive AI incident most teams have ever had wasn't a wrong answer. It was a loop. An agent that decided to retry, and retry, and retry, with each retry appending more context, until the bill exceeded the monthly budget in hours. Rate limiting at the gateway is the layer that turns this from a budget-destroying incident into a paged on-call ticket. The three primitives — token bucket per identity, circuit breaker per pattern, fallback chain per route — are well-known SRE infrastructure adapted to a workload where the unit of waste is dollars per token instead of milliseconds per request.

TL;DR Layer 1: token bucket per (user, repo, model) — 429s tell well-behaved agents to back off. Layer 2: circuit breakers that trip on pattern (cost velocity, repeated prompts, error rate, growing context). Layer 3: declarative fallback chain — primary → cheaper model → semantic cache → 503. Goal isn't to eliminate runaways; it's bounded blast radius — a misbehaving caller never breaks other callers, the budget, or the on-call's sleep.

The runaway loop is the default failure mode

LLM agents fail in a specific shape. The most common production incident is not a model giving the wrong answer; it is an agent that decides to retry, and retry, and retry, and retry. Each retry is a full provider call. Each call appends to context. Context grows quadratically. Tokens are consumed at a rate the human in front of the keyboard would never produce — because there is no human in front of the keyboard.

The arithmetic is brutal. A 4,000-token initial context, doubling at each step because the previous step's output gets appended, reaches 128,000 tokens at step 5 and the per-step cost has gone up 32×. By step 15 the context has overflowed the model's window and every call is paying the maximum-context rate. By step 30 the loop has spent more than a competent engineer's monthly salary. The agent never noticed; the agent's job is to keep going.

The first time most teams see this, they see it on the next day's bill. The second time, they put rate limiting in. The right place for that rate limiting is not the application — it's the gateway, where a single layer protects every workload regardless of which framework launched the loop. A team that puts the limit inside each agent has to write it again for each agent, miss it in some, and discover the failure modes individually. A team that puts it at the gateway writes it once and inherits the protection across every workload that ever calls a model.

Three layers of enforcement

Figure 1 — The runaway agent's traffic ramps from 1 req/s to 110 req/s in two minutes. The gateway's three enforcement layers — token bucket, circuit breakers, fallback chain — turn what would be a budget-draining incident into a graceful degradation. The token bucket throttles volume; the circuit breakers catch pattern; the fallback chain preserves user experience while the breaker is open.

Layer 1 — Token bucket per identity

The bucket is the first line. Every (user, repo, model) tuple gets its own bucket — say, 100 requests per minute with a burst of 200. Refills happen continuously at the configured rate. When a request arrives and the bucket is empty, the gateway returns HTTP 429 immediately, before the request reaches the provider. The 429 carries a Retry-After header indicating how long the caller should wait.

The choice of granularity matters. A single bucket per user is too coarse — one rogue script blocks all the user's legitimate work, including the work they need to debug the rogue script. A bucket per request is too fine — there's nothing to throttle. The (user, repo, model) tuple is usually the right shape: it isolates the runaway repository from the user's other workloads, allows the platform to set different ceilings for different model tiers, and produces useful dimensional breakdowns in the rate-limit dashboard.

Most agent frameworks read 429 as a standard backoff signal. The agent pauses, waits for the suggested duration, and retries — exactly the behavior the gateway wants. Frameworks that don't handle 429 gracefully are broken in a different way, but the gateway has already done its job. The contract is HTTP-standard: 429 with Retry-After is what every HTTP client library has known how to handle for fifteen years. Teams that find their agent framework can't process 429 correctly should fix the framework, not the gateway.

Burst is the parameter that gets tuned. Set the burst too low and legitimate bursty workloads hit 429 falsely; set it too high and the rate limit doesn't actually constrain a runaway. The right starting point is empirical — look at the natural burst shape of legitimate traffic in the previous month, set the bucket size to roughly the P99 burst observed, and tune from there based on false-positive complaints.

Layer 2 — Circuit breakers

Token buckets handle volume. Circuit breakers handle pattern. Some runaways pass under the bucket's ceiling but exhibit other signatures of pathological behavior — an agent that obeys the 100-rpm bucket but is making 100 identical calls per minute in a tight loop is still a runaway, just a slower one. The gateway watches every identity's recent traffic and trips when one of the following holds:

Each breaker has its own trip and reopen logic. The combined effect is that pathological behavior gets stopped quickly without false-positive trips on legitimate bursty traffic.

Layer 3 — Fallback chain

When the primary path is unavailable — the token bucket is throttled, the circuit breaker is open, the provider itself is having an outage — hard-failing every request is the wrong outcome. A configurable fallback chain serves degraded output instead. The chain is declarative, per-route configuration:

# Per-route configuration
fallback_chain:
  - model: "claude-4-6-sonnet"        # primary
  - model: "claude-4-5-haiku"         # cheaper, still capable
  - source: "semantic-cache"          # cached response if similar
  - response: 503                     # last resort, descriptive error

The chain is a config knob, not hardcoded behavior. Different routes have different chains.

What this looks like in cost terms

Figure 2 — The same runaway agent under two architectures. Without enforcement, cost grows quadratically with the loop's context until something else (a provider rate limit, a finance review, an exhausted budget) stops it — typically hours later. With the three-layer enforcement, the cost is bounded within minutes.

How this looks to the caller

Caller receives Meaning Recommended response
HTTP 429 Bucket empty. Try again after Retry-After. Standard backoff.
HTTP 503 Fallback chain exhausted. Surface to user as "AI temporarily unavailable."
200 — primary model Normal path. —
200 — fallback model Header X-TFY-Resolved-Model identifies the actual model. Quality may differ; UI may indicate this for some routes.

Common configuration mistakes

Mistake 1: too-coarse bucket key. A single bucket per user means one runaway agent blocks the user's debugging work. The (user, repo, model) tuple is usually the right granularity.

Mistake 2: no fallback chain. A workload that fails hard on every 429 produces a worse user experience than necessary.

Mistake 3: paging on every 429. 429s are normal; they are the system working as designed.

Mistake 4: cost-velocity thresholds set as static dollar amounts. A static threshold becomes obsolete quickly; the threshold should be relative to the workload's planned budget.

How to roll this out without breaking things

Rate limiting is one of the few platform changes where the deployment risk is asymmetric: not enforcing is the status quo, enforcing too aggressively breaks legitimate workloads. The rollout sequence should bias toward observation first, enforcement second.

  1. Week 1 — Audit-only buckets, audit-only breakers. Log but don't enforce.
  2. Week 2 — Enable the token bucket only, on one workload class. Pick the highest-risk class.
  3. Week 3 — Add the cost-velocity breaker. Pages are configured but routed only to the platform team during week 3.
  4. Week 4 — Other breakers, fallback chain. Add loop-signature detectors and configure fallback chains.
  5. Week 5+ — Steady state. The pages route to product on-call.

The operating model around the rate-limit layer

The rate-limiting layer is owned by the platform team. The workload-specific thresholds are co-owned with the workload teams.

Multi-tenant rate limiting — the B2B-SaaS case

The rate-limiting layer needs to isolate tenants from each other while enforcing per-tenant pricing tiers.

FAQ

What's the right value for the token bucket limit?

Start from the natural traffic shape — adjust based on false-positive rate.

Does rate limiting at the gateway replace the provider's own rate limits?

No — they compose.

How do we handle cases where the agent needs many calls?

Configure a higher bucket for that workload.

What happens when the breaker is open?

The fallback chain takes over; if it's exhausted, the application gets a clean HTTP 503.