How should an API client retry a failed request without making the outage worse?
Rate limiting shares capacity fairly, and the algorithm decides how
A limiter protects a shared resource from any one caller, malicious or just retrying badly. The algorithm decides which bursts get through.
For this question · chapter 3 of 3
Tell the client how to back off
- The limiter protects you, but the other half of rate limiting keeps your callers well behaved. Client A gets a 200 response with X-RateLimit-Limit 100, X-RateLimit-Remaining 1 and X-RateLimit-Reset 30 seconds. A client that can see its budget can stay inside it: Client A knows it has one request left until the reset in 30 seconds. The draft standard's RateLimit headers carry the same three numbers.
- This refusal is a helpful error, because it tells each client when to come back. Both clients get a 429 Too Many Requests with Retry-After: 30, so each one knows to wait 30 seconds. The Retry-After value can be a number of seconds or an HTTP date. A 429 means this one caller is over its limit, but a 503 would mean the whole service is down.
- Every instant retry costs the server work and achieves nothing. Client A ignores Retry-After, the header that says when to try again, and sends five retries in two seconds. The server answers all five with a 429 right away, so its refused count reaches 7. Retrying instantly is how one slow minute turns into an outage.
- Client B shows the client half of rate limiting: read the header, wait, then succeed. Client B obeys Retry-After and waits the full 30 seconds. Then Client B sends just one request. That request gets a 200 and a fresh budget of 99. This loop is the half of rate limiting that interviewers ask about.
- Retry-After tells each client how long to wait, so clients that all obey the same value come back at the same moment. Fifty clients got Retry-After: 30, so all fifty retry in the same second. The server receives 50 requests in 1 s. Fifty polite clients with the same wait are a burst, and the limiter was built to refuse a burst.
- The fix is to make each client wait a different time. Jitter adds a random extra wait, so the server receives the same fifty retries spread over 30 seconds, from second 30 to second 60. Exponential backoff doubles every client's wait after each failure, but it doubles all the waits together. So without jitter, the clients all retry together again in every round.
The limiter protects you, but the other half of rate limiting keeps your callers well behaved. Client A gets a 200 response with X-RateLimit-Limit 100, X-RateLimit-Remaining 1 and X-RateLimit-Reset 30 seconds. A client that can see its budget can stay inside it: Client A knows it has one request left until the reset in 30 seconds. The draft standard's RateLimit headers carry the same three numbers.
I keep a token bucket of 20 in Redis, so all three instances share one count. The 20 is the burst I allow on purpose, and the refill of 2 per second is the rate. My 429 carries Retry-After, so a good client waits. My clients also back off exponentially with jitter, so they never retry at the same moment.
© LearnThatStack - diagrams may not be republished without permission.
Want a quick review of the fundamentals? See the API Design cheatsheet.
90 of 110 API Design answers are in Pro.
Full answers, code samples, and AI explanations that go simpler or deeper. Cancel anytime.
- Full answers + code
- AI explanations, simpler or deeper
- 1,000 AI credits / month
- Cancel anytime