LearnThatStack Ace your next interview

System Design Concepts · API Design

Async APIs and webhooks: when the work takes longer than the request

Do not keep the connection open. Return 202 with a job the client can poll, or send a webhook to whoever needs to know and handle every delivery problem.

asyncHolding the connection breaksa timeout you cannot setclient held19sserver work109sload balancertimeout 60 sClientretryingAPI serverPOST /data-exportwork on the server1 wasted, 1 runningexport #1orphanedexport #219 s of 90 s
asyncHolding the connection breaksclient held19sserver work109sload balancertimeout 60 sClientretryingAPI serverPOST /data-exportwork on the server1 wasted, 1 runningexport #1orphanedexport #219 s

Holding the connection breaks

  1. A normal read request gets its answer, and the connection closes after 80 ms. The client holds the connection exactly as long as the server works, and the wait costs nothing because the work is quick.
  2. Now the client asks for a data export and keeps the connection open while the work runs. The job needs ninety seconds, so the client holds the connection for the full ninety seconds.
  3. At 60 seconds, the load balancer (the proxy in front of the server) closes the connection, so the client gets 504 Gateway Timeout. The load balancer owns this 60-second limit, not the server, so the handler cannot raise the limit.
  4. The 4.1 MB export still finishes, but the reply goes to a socket that no client is holding. Cutting the connection did not cancel the work, so the export is orphaned: done, but nobody collected it.
  5. The client retries, so the server starts a second export before anyone collects the first one. The server has now spent 109 seconds on a file nobody has, and the client has waited 19 seconds for the same export again.

© LearnThatStack - diagrams may not be republished without permission.

Request and reply is the default pattern, because it is easy to write and easy to reason about. This read takes 80 ms, and then the connection closes. The client waits exactly as long as the server works. Here the work is quick enough that the wait costs nothing. Trouble starts only when the work stops being quick.

Most people hold the connection open the first time they have to run something slow. The export needs ninety seconds of work on the server. So the client keeps the connection open for all ninety seconds. Nothing has gone wrong yet. The client has waited exactly as long as the server worked on the export.

The timeout that ends this request comes from the load balancer, the proxy in front of your service. The load balancer closes the connection at 60 seconds, so the client receives 504 Gateway Timeout. You did not set that limit, and your handler cannot raise it. In most stacks, the handler never finds out that the timeout happened.

Closing the connection does not cancel the server's work. The client closed the connection at 60 seconds, but the server has now worked for 90 seconds. So the server finishes the full 4.1 MB export after the client is gone. The server then sends the reply to a closed socket, so no client receives the file.

The server has spent 109 seconds on a file that nobody has received. The client retries, so the same work now runs twice. The retry doubles the cost and lowers the chance that either run finishes before the timeout. So instead of holding the connection open, hand back a ticket: a job the client can poll.

The server frees the connection after 40 ms, even though the export still needs ninety seconds. The server replies with 202 Accepted, a job id and a status URL. A 202 means the server has accepted the work, not that the work is finished. The job id and status URL give the client a way to check on the job later, not the file.

The job has its own URL, so 202 is a normal REST answer. The client does not have to keep the connection open, so the connection closes. The job waits in the queue at /jobs/42. The work is now a resource, so the client can send a GET request to /jobs/42 at any time.

A worker has started the job, and the job is now 18 per cent done. The worker spends the ninety seconds on the export, and the API only reports the job's status. No connection stays open while the export runs, so you can deploy the API servers and lose nothing.

Polling has a bad reputation because some clients ask every 200 ms. Here the client asks once, with GET /jobs/42. The server answers in 38 ms and says the job is running, 61 per cent done. Each poll is tiny, and the Retry-After: 5 header lets the server decide when the client asks again.

No connection stayed open at any point, but the client still got the answer with a second short request. The finished job returns a result URL, so the file has its own address. The API never holds a request open for a 4.1 MB file. The client fetches the file whenever it wants to.

A job can fail long after the server returns 202, so there is no response left to put the error in. A second export fails 34 per cent of the way through. So the client fetches the job's status and reads the failure there, with an error code and a request id in the body. People who draw only the happy path, where nothing goes wrong, leave this failed state out.

Returning 202 with a job to poll is for your own long-running work. With a webhook, you become the caller, because someone else needs to know that something happened. Their server, the consumer, gives you a URL to POST to, /hooks/orders, and gets a signing secret back from you. The signing secret is the only thing that lets the consumer tell your calls apart from anyone else's.

A webhook brings back every problem that the job resource (the job the client polls) let you avoid, but with the roles reversed. A payment succeeds, so your API becomes the client and sends a delivery to the consumer's URL. The delivery includes the event id evt_9f3 and an HMAC signature of the body. This time, the consumer has to handle those problems.

Anyone can send a POST to a public webhook URL, so only the signature proves that a delivery came from you. The consumer computes the same signature with the shared secret, so the consumer accepts and records the event evt_9f3. A forged POST from somewhere else has no signature at all, so the consumer returns 401. Compare the received and recomputed signatures in constant time: the check takes the same time whether the two signatures match or not.

Their server records this delivery, then fails with a 500. So their outage becomes work for your retry queue, which tries again after 1 second, then 5, then 25. Exponential backoff multiplies the wait after each failure, so your retries do not cause a second outage on their server. After six attempts, the delivery becomes a dead letter: your queue sets it aside and stops retrying it.

At-least-once delivery is the guarantee you can actually make. It means the consumer sometimes gets the same event twice. Here the retry arrives after the consumer's first write already worked, so the consumer now has the same event id twice. No amount of care on your side removes the risk of a duplicate.

An idempotent consumer applies each event only once, no matter how many times the event arrives. The consumer's only protection is the set of event ids it has already seen, in practice one table with a unique column. The table already holds evt_a02, so the consumer drops the second copy and applies only 2 of the 3 deliveries. Say the term idempotent consumer before the interviewer does.

Treat the event as a doorbell: a signal that something changed, not the new data itself. One retry makes order.created arrive after order.updated. So the consumer ignores the data in both payloads and fetches /orders/88, which returns shipped. The current state is the same whichever event arrives first. So fetching the current state means the order the events arrive in no longer matters.

For your own long work, return 202 with a job resource and a Retry-After header, because the server sets the poll interval. Give a failed job its own status and error body. Send a webhook when someone else needs to know, or stream when they need each event as it happens. Sign every delivery and retry with exponential backoff into a dead letter queue. At-least-once delivery sends some events twice, so make the consumer idempotent on the event id.

In interviews · 6 questions

Related Questions

Also helps with

← All concepts