System Design Concepts · API Design
Async APIs and webhooks: when the work takes longer than the request
Do not keep the connection open. Return 202 with a job the client can poll, or send a webhook to whoever needs to know and handle every delivery problem.
Holding the connection breaks
- A normal read request gets its answer, and the connection closes after 80 ms. The client holds the connection exactly as long as the server works, and the wait costs nothing because the work is quick.
- Now the client asks for a data export and keeps the connection open while the work runs. The job needs ninety seconds, so the client holds the connection for the full ninety seconds.
- At 60 seconds, the load balancer (the proxy in front of the server) closes the connection, so the client gets 504 Gateway Timeout. The load balancer owns this 60-second limit, not the server, so the handler cannot raise the limit.
- The 4.1 MB export still finishes, but the reply goes to a socket that no client is holding. Cutting the connection did not cancel the work, so the export is orphaned: done, but nobody collected it.
- The client retries, so the server starts a second export before anyone collects the first one. The server has now spent 109 seconds on a file nobody has, and the client has waited 19 seconds for the same export again.
202 returns a job ID
- The server answers the export request in 40 ms with 202 Accepted and a ticket: a job id and a status URL. Accepted means the server has the work, not that the work is finished, so the client holds this ticket instead of a file.
- The connection is already closed, and the job waits in the queue at its own address, /jobs/42. So the requested work is now a resource, and the client can ask about the job at any time.
- A worker takes the job from the queue and starts writing the export, which is ninety seconds of work. The API handles only the job's status, so no connection stays open while the export runs.
- The client polls once by sending GET /jobs/42 to ask for the job's status. The server replies in 38 ms that the job is running at 61 per cent, and a Retry-After: 5 header tells the client when to ask again. Each poll is tiny, so asking costs almost nothing.
- The job is done, and its status now includes a result URL, the file's own address. So the answer comes back in a second short request, and the API never holds a request open for 4.1 MB.
- A job can fail long after the 202 reply, and here a second export fails at 34 per cent. So the failure is a status the client fetches, not a broken connection, with an error code and a request id in the response body.
Webhooks invert the call
- The consumer, the server that receives your events, registers /hooks/orders as the URL to POST events to. Your API returns a signing secret, the only thing that lets the consumer tell your calls apart from anyone else's.
- A payment completes, so your API is now the client and sends a delivery to their URL. The delivery includes the event id evt_9f3 and an HMAC signature of the body, made with the shared secret.
- The consumer computes the signature again with the shared secret, then accepts the delivery and records the event id evt_9f3. A POST from anywhere else has no signature at all, so the consumer refuses it with 401. The URL is public and takes a POST from anyone, so the signature is the only proof that a delivery came from you.
- The consumer records the next delivery, but then their endpoint answers 500. So their outage becomes your retry queue's job: try again after 1 second, then 5, then 25, and set the delivery aside as a dead letter after six tries.
- The retry arrives, but the consumer's first write had already worked, so the consumer now has the same event id twice. At-least-once delivery sometimes delivers an event twice, whatever care your side takes.
- The consumer's only defence is a record of the event ids it has already seen. Here evt_a02 is already in that record, so the consumer drops the second copy and applies only 2 of the 3 deliveries that arrived.
- A retry delays order.created, so order.updated arrives first. On each event, the consumer ignores the payload and fetches /orders/88 for the current state (shipped), so the order of events no longer matters.
© LearnThatStack - diagrams may not be republished without permission.
Request and reply is the default pattern, because it is easy to write and easy to reason about. This read takes 80 ms, and then the connection closes. The client waits exactly as long as the server works. Here the work is quick enough that the wait costs nothing. Trouble starts only when the work stops being quick.
Most people hold the connection open the first time they have to run something slow. The export needs ninety seconds of work on the server. So the client keeps the connection open for all ninety seconds. Nothing has gone wrong yet. The client has waited exactly as long as the server worked on the export.
The timeout that ends this request comes from the load balancer, the proxy in front of your service. The load balancer closes the connection at 60 seconds, so the client receives 504 Gateway Timeout. You did not set that limit, and your handler cannot raise it. In most stacks, the handler never finds out that the timeout happened.
Closing the connection does not cancel the server's work. The client closed the connection at 60 seconds, but the server has now worked for 90 seconds. So the server finishes the full 4.1 MB export after the client is gone. The server then sends the reply to a closed socket, so no client receives the file.
The server has spent 109 seconds on a file that nobody has received. The client retries, so the same work now runs twice. The retry doubles the cost and lowers the chance that either run finishes before the timeout. So instead of holding the connection open, hand back a ticket: a job the client can poll.
The server frees the connection after 40 ms, even though the export still needs ninety seconds. The server replies with 202 Accepted, a job id and a status URL. A 202 means the server has accepted the work, not that the work is finished. The job id and status URL give the client a way to check on the job later, not the file.
The job has its own URL, so 202 is a normal REST answer. The client does not have to keep the connection open, so the connection closes. The job waits in the queue at /jobs/42. The work is now a resource, so the client can send a GET request to /jobs/42 at any time.
A worker has started the job, and the job is now 18 per cent done. The worker spends the ninety seconds on the export, and the API only reports the job's status. No connection stays open while the export runs, so you can deploy the API servers and lose nothing.
Polling has a bad reputation because some clients ask every 200 ms. Here the client asks once, with GET /jobs/42. The server answers in 38 ms and says the job is running, 61 per cent done. Each poll is tiny, and the Retry-After: 5 header lets the server decide when the client asks again.
No connection stayed open at any point, but the client still got the answer with a second short request. The finished job returns a result URL, so the file has its own address. The API never holds a request open for a 4.1 MB file. The client fetches the file whenever it wants to.
A job can fail long after the server returns 202, so there is no response left to put the error in. A second export fails 34 per cent of the way through. So the client fetches the job's status and reads the failure there, with an error code and a request id in the body. People who draw only the happy path, where nothing goes wrong, leave this failed state out.
Returning 202 with a job to poll is for your own long-running work. With a webhook, you become the caller, because someone else needs to know that something happened. Their server, the consumer, gives you a URL to POST to, /hooks/orders, and gets a signing secret back from you. The signing secret is the only thing that lets the consumer tell your calls apart from anyone else's.
A webhook brings back every problem that the job resource (the job the client polls) let you avoid, but with the roles reversed. A payment succeeds, so your API becomes the client and sends a delivery to the consumer's URL. The delivery includes the event id evt_9f3 and an HMAC signature of the body. This time, the consumer has to handle those problems.
Anyone can send a POST to a public webhook URL, so only the signature proves that a delivery came from you. The consumer computes the same signature with the shared secret, so the consumer accepts and records the event evt_9f3. A forged POST from somewhere else has no signature at all, so the consumer returns 401. Compare the received and recomputed signatures in constant time: the check takes the same time whether the two signatures match or not.
Their server records this delivery, then fails with a 500. So their outage becomes work for your retry queue, which tries again after 1 second, then 5, then 25. Exponential backoff multiplies the wait after each failure, so your retries do not cause a second outage on their server. After six attempts, the delivery becomes a dead letter: your queue sets it aside and stops retrying it.
At-least-once delivery is the guarantee you can actually make. It means the consumer sometimes gets the same event twice. Here the retry arrives after the consumer's first write already worked, so the consumer now has the same event id twice. No amount of care on your side removes the risk of a duplicate.
An idempotent consumer applies each event only once, no matter how many times the event arrives. The consumer's only protection is the set of event ids it has already seen, in practice one table with a unique column. The table already holds evt_a02, so the consumer drops the second copy and applies only 2 of the 3 deliveries. Say the term idempotent consumer before the interviewer does.
Treat the event as a doorbell: a signal that something changed, not the new data itself. One retry makes order.created arrive after order.updated. So the consumer ignores the data in both payloads and fetches /orders/88, which returns shipped. The current state is the same whichever event arrives first. So fetching the current state means the order the events arrive in no longer matters.
For your own long work, return 202 with a job resource and a Retry-After header, because the server sets the poll interval. Give a failed job its own status and error body. Send a webhook when someone else needs to know, or stream when they need each event as it happens. Sign every delivery and retry with exponential backoff into a dead letter queue. At-least-once delivery sends some events twice, so make the consumer idempotent on the event id.
In interviews · 6 questions
Related Questions
- 01
- 02