The short answer
Quick answer: A rate limiter caps how many requests a client can make in a period of time. For each client (identified by API key, user or IP address) it keeps a small piece of state, such as a counter or a bucket of tokens, and checks it on every request. If the client is within its allowance, the request passes. If not, the server rejects it, usually with HTTP status 429 Too Many Requests, and tells the client when to try again. The most common algorithm is the token bucket, which allows short bursts while enforcing an average rate.
Why limit requests at all
- Stop abuse. Password guessing, scraping, spam and denial-of-service attempts all involve sending many requests.
- Protect the system. One buggy client in a retry loop should not take the service down for everyone.
- Fairness. Shared capacity should be divided reasonably between customers.
- Cost control. Each request may trigger expensive work, such as a database query or a paid third-party call.
- Business tiers. Free plans get 100 requests an hour; paid plans get more.
Stripe's engineering post Scaling your API with rate limiters describes running several kinds of limiter together for these different purposes.
The algorithms
Fixed window counter
Divide time into fixed windows, such as each minute, and count requests per window. Reset at the start of the next.
- Simple and cheap: one counter per client.
- Boundary problem: with a limit of 100 per minute, a client can send 100 requests at 12:00:59 and 100 more at 12:01:00. That is 200 requests in two seconds.
Sliding window log
Store the timestamp of every request. On each new request, discard timestamps older than the window and count the rest.
- Exactly accurate.
- Memory hungry: it stores every request, which is costly for busy clients.
Sliding window counter
A compromise: keep counters for the current and previous fixed windows and estimate with a weighted sum.
estimate = current_count + previous_count × (share of previous window still in range)
Thirty seconds into the current minute, half of the previous minute's count is included.
- Smooths out the boundary problem with only two counters per client.
- An approximation, but close enough for nearly all uses.
Token bucket
Picture a bucket that holds up to N tokens. Tokens are added at a steady rate. Each request removes one token. If the bucket is empty, the request is rejected.
- The refill rate sets the long-term average.
- The bucket size sets how big a burst is allowed.
A bucket of 10 refilled at 1 per second lets a client make 10 requests at once after a quiet period, then 1 per second.
No background timer is needed. Store only the token count and the time of the last update, and calculate lazily:
def allow(bucket, now, rate, capacity):
elapsed = now - bucket.last
bucket.tokens = min(capacity, bucket.tokens + elapsed * rate)
bucket.last = now
if bucket.tokens >= 1:
bucket.tokens -= 1
return True
return False
The token bucket is the most widely used algorithm because it matches real usage: clients are bursty, and a short burst is usually harmless.
Leaky bucket
Requests enter a queue (the bucket) and are processed at a constant rate, like water dripping from a hole. If the queue is full, new requests are dropped.
- Output is perfectly smooth, which protects a fragile downstream system.
- Bursts are delayed, not served immediately.
| Algorithm | Memory per client | Allows bursts | Accuracy |
|---|---|---|---|
| Fixed window | One counter | Yes, up to 2x at boundaries | Low |
| Sliding window log | Every timestamp | No | Exact |
| Sliding window counter | Two counters | Limited | Good |
| Token bucket | Two numbers | Yes, by design | Good |
| Leaky bucket | A queue | No; smooths them | Good |
Where the limiter sits
- At the edge: a CDN, API gateway or load balancer rejects excess traffic before it reaches your servers. This is the most efficient place.
- In application middleware: allows limits based on business context, such as the user's plan or the cost of a specific endpoint.
- In the client: well-behaved clients throttle themselves to stay under a known limit.
Most systems use layers: coarse limits at the edge, finer ones in the application.
What to limit by
| Key | Pros | Cons |
|---|---|---|
| API key or user ID | Accurate and fair | Only works for authenticated requests |
| IP address | Works for anonymous traffic | Many users can share one IP behind NAT; attackers rotate IPs |
| Endpoint | Protects expensive operations | Needs per-route configuration |
| Combination | Most precise | More state |
Shared IP addresses are a real issue: a whole office or mobile network can appear as one address. See how NAT works. Login endpoints are usually limited per account and per IP to slow password guessing.
Rate limiting across many servers
With one server, a counter in memory is enough. With twenty servers behind a load balancer, each would keep its own count and a client could get twenty times the limit.
The usual fix is a shared, fast store, most often Redis. See why Redis is fast.
The subtle part is atomicity. "Read the counter, check it, then increment" is a race: two servers can both read 99 and both allow the hundredth request. Solutions:
- Use a single atomic command such as
INCRwith an expiry. - Put the check-and-update logic in a Lua script, which Redis runs as one uninterruptible step.
Each check adds a network round trip, so high-volume systems sometimes keep local counters and sync with the central store periodically, accepting a little inaccuracy for speed.
Decide in advance what happens if the store is unavailable. Failing open (allowing requests) keeps the service usable; failing closed (rejecting them) is safer for sensitive endpoints.
Responding to the client
The standard response is 429 Too Many Requests, defined in RFC 6585, with a Retry-After header:
HTTP/1.1 429 Too Many Requests
Retry-After: 30
RateLimit-Limit: 100
RateLimit-Remaining: 0
RateLimit-Reset: 30
Many APIs also send the limit and remaining allowance on every response, so clients can pace themselves.
Being a good client
When you receive a 429:
- Respect
Retry-After. - Otherwise, retry with exponential backoff: wait 1 second, then 2, then 4.
- Add jitter (a random variation), so that thousands of clients do not all retry at the same instant.
- Make retried writes safe with idempotency keys.
Related techniques
- Concurrency limits cap how many requests are in progress at once, which protects against slow requests piling up.
- Load shedding drops low-priority traffic when the system is overloaded, regardless of who sent it.
- Circuit breakers stop calling a failing dependency for a while.
- Autoscaling adds capacity for legitimate growth; see how auto-scaling works.
Rate limiting is not full protection against large denial-of-service attacks. Those need to be absorbed at the network edge.
Frequently asked questions
What does 429 Too Many Requests mean?
You have sent more requests than the server allows in a given time. Wait for the period given in Retry-After, then try again.
What is the difference between token bucket and leaky bucket?
A token bucket allows bursts up to the bucket size while enforcing an average rate. A leaky bucket releases requests at a constant rate, smoothing bursts out.
What is the difference between rate limiting and throttling?
The terms overlap. Rate limiting usually means rejecting excess requests; throttling often means slowing or queuing them.
How do I rate limit across multiple servers?
Keep the counters in a shared store such as Redis and update them atomically, so all servers see the same count.
Conclusion
A rate limiter is a small amount of state and a simple rule applied on every request. Token buckets cover most needs, sliding windows give stricter control, and a shared atomic counter makes it work across a fleet. Pair server-side limits with clear 429 responses and clients that back off politely, and one misbehaving caller stops being everyone's problem.
Related articles
- How Redis Is So Fast
- What Is a Load Balancer and How Does It Decide Where Traffic Goes?
- How Payment Systems Avoid Charging You Twice (Idempotency)
- How Auto-Scaling Handles Sudden Traffic Spikes
