API Rate Limiting: Token Buckets, Choosing the Right Key and Distributed Counter Pitfalls
Why fixed windows let twice the limit through, why the IP address is a poor key, how INCR plus EXPIRE in Redis can block a client forever, and why you should answer with 429 and Retry-After instead of 503.
Shipping an API without rate limiting is like leaving the door open and hoping nobody walks in. You do not need an attacker for it to hurt: a client stuck in a loop, a badly written retry policy or one large customer running a bulk import is enough to slow everyone else down. That makes rate limiting less a security feature and more a fairness and resilience mechanism. The trouble is that "100 requests per minute" turns into four separate decisions the moment you implement it: which algorithm, which key, where the counter lives and what you return when the limit is hit. Each of those has a common mistake attached.
Picking an algorithm
Fixed window is the simplest: keep one counter per minute and reject once it passes the limit. Its weakness is the window boundary. With a limit of 100 per minute, a client can send 100 requests at 00:59 and another 100 at 01:00, pushing twice the limit through in two seconds. It is fine for coarse quotas such as a daily allowance, not for protecting against bursts.
Sliding log stores a timestamp for every request and counts the ones in the last 60 seconds. It is exact, but it keeps as many entries per client as the limit allows, so memory grows quickly with high limits.
Sliding window counter sits in between. It keeps the current and previous window counters and weights the previous one by how much of it still overlaps. With two integers it behaves almost as smoothly as the sliding log.
Token bucket asks a different question: how much is the client allowed to save up? The bucket refills at a constant rate up to a capacity, and each request spends a token. That gives you two knobs, sustained rate and burst size. Real clients do not send evenly; a page load fires ten requests at once and then goes quiet. A token bucket tolerates that shape while still cutting off sustained abuse. GCRA (Generic Cell Rate Algorithm) produces the same behaviour with a single timestamp per client, which makes it cheaper to store.
If you want a recommendation: start with a token bucket or GCRA for user facing APIs. Use fixed windows only for billing related quotas, where a calendar limit such as "10,000 requests this month" is exactly what people expect.
What are you counting?
The key is wrong more often than the algorithm.
The IP address is the easiest key and the most misleading one. An office, a mobile carrier's carrier grade NAT or a university dorm can put hundreds of users behind one IPv4 address, and a strict per IP limit blocks all of them together. IPv6 fails the other way: a client usually gets at least a /64 prefix and can use a different address for every request. Key IPv6 clients by prefix, not by full address.
If you sit behind a reverse proxy or CDN, make sure you read the IP correctly. Use the socket's source address and all traffic appears to come from the proxy, so everyone shares one bucket. Trust the client supplied X-Forwarded-For header blindly and an attacker bypasses the limit by writing a new value on every request. Only read the header on connections from proxies you trust, and take the last hop in the chain that you trust.
For authenticated requests the key should be the user or the API key. In a multi tenant system add a second limit at tenant level; otherwise one tenant can create a hundred users and consume the whole shared budget.
Endpoints are not equally expensive. A search or report request can cost a hundred times more than a simple read. A nice property of the token bucket is that a request can spend more than one token, so give expensive endpoints a higher cost.
Login attempts are a separate problem
Limiting password attempts needs two counters for two different attacks: many passwords against one account (a per account counter) and a few common passwords against many accounts (a per IP or per prefix counter). Add only one and the other attack walks straight through.
The per account counter has a side effect: an attacker can deliberately enter wrong passwords to lock the real user out. Instead of locking the account permanently, apply a short backoff that grows with each failure, and reset it on a successful login. Do not reveal whether the account exists, and return the same generic response when the limit is hit.
When there is more than one server
If the counter lives in each application instance's memory, a service running four replicas has an effective limit four times what you configured, and it fluctuates with how the load balancer distributes traffic. Redis is the usual choice for a shared counter, and there are two classic mistakes.
The first is the read then write race. If you read the value, decide in application code and then write, two concurrent requests see the same stale value and both get through. The second is the INCR followed by EXPIRE pattern common in fixed window implementations: if the process dies between the two commands, the key never gets a TTL and that client is blocked forever. The fix for both is to move the logic into a single atomic Lua script.
The script below implements a token bucket. It takes the time from Redis's own clock rather than the application server, so clock skew between servers does not distort the refill rate. Calling TIME before a write inside a script is fine on Redis 7 and later, where scripts are replicated by their effects.
-- KEYS[1]: bucket key
-- ARGV[1]: capacity, ARGV[2]: refill per second, ARGV[3]: cost of this request
local capacity = tonumber(ARGV[1])
local rate = tonumber(ARGV[2])
local cost = tonumber(ARGV[3])
local t = redis.call('TIME')
local now = tonumber(t[1]) + tonumber(t[2]) / 1000000
local state = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(state[1]) or capacity
local ts = tonumber(state[2]) or now
tokens = math.min(capacity, tokens + (now - ts) * rate)
local allowed = 0
local wait = 0
if tokens >= cost then
tokens = tokens - cost
allowed = 1
else
wait = (cost - tokens) / rate
end
redis.call('HSET', KEYS[1], 'tokens', tokens, 'ts', now)
redis.call('EXPIRE', KEYS[1], math.ceil(capacity / rate) + 1)
return {allowed, tostring(wait)}
The wait time is returned as a string on purpose: Redis truncates floating point numbers returned from Lua to integers. The TTL is long enough for the bucket to refill completely, at which point a missing key and a full bucket mean the same thing, so idle clients cost no memory. A request whose cost exceeds the capacity can never pass; validate that when you load the configuration.
Decide in advance what happens when Redis is unreachable. For general API traffic, failing open (letting the request through) is usually right; a broken rate limiter should not take the whole service down. For login and password reset endpoints, failing closed is safer. Keep both behaviours configurable and emit a metric either way.
What to return when the limit is hit
The right status code is 429, together with a Retry-After header. That header is part of the HTTP standard and is the one stable signal a client can rely on. There is an IETF draft for RateLimit headers that report remaining quota, but the header names have changed between draft revisions; if you send them, treat them as informational and do not build client logic on top of them.
A trap: nginx's limit_req module returns 503 by default. Leave it that way and clients will read rate limiting as a server failure, and your monitoring will page you for an outage that is not happening.
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
server {
location /api/ {
limit_req zone=api burst=20 nodelay;
limit_req_status 429;
}
}
Here burst plays the role of bucket capacity. Without nodelay, nginx queues and delays requests within the burst instead of serving them immediately, which shows up on the client side as unexplained latency.
The client is the other half of the job. A client that gets a 429 should honour Retry-After, and if the header is missing, use exponential backoff with random jitter. Without jitter, thousands of clients that hit the limit together retry together, and the same wave hits the limiter every time.
Layers, and when not to do it
One limit in one place is not enough. A coarse, IP based limit at the edge, in the proxy or CDN, stops bulk abuse before it reaches the application. Inside the application, identity and tenant aware limits with cost weighting provide fairness. They solve different problems.
There are also cases where rate limiting is the wrong tool. If what you want to protect is the capacity of a backend database, limit concurrency rather than request rate; ten fast requests a second are harmless, ten slow queries open at the same time are not. For internal traffic between your own services, use queues and backpressure instead of hard rejection; answering an internal caller with 429 usually just moves the problem one layer up as a retry storm. And a rate limit is not a substitute for capacity planning: if normal traffic is approaching the limit, it is the infrastructure that needs to grow, not the limit.
One last practical tip: run the limiter in log only mode for a while before enforcing it. Setting thresholds without seeing which clients would actually be blocked by real traffic is the fastest way to block your best customer on day one.