LearnThatStack Ace your next interview

System Design Concepts · API Design

Status codes and error design: how an API says NO

The status code tells machines whose fault the failure was and whether to retry. The error body tells a developer what to fix, and it is never an apology.

statusThe code routes the failurewhose fault, and whether to retryRequestany request2xx200 201 204, donedone4xx404 not there, 409 conflictyours5xx503, overloaded, Retry-Aftermineno retryretrythe twelve that matter200201204400401403404409422429500503
statusThe code routes the failureRequestany request2xx200 201 204, donedone4xx404 not there, 409 conflictyours5xx503, overloaded, Retry-Aftermineno retryretry200201204400401403404409422429500503

The code routes the failure

  1. The first digit of the status code sorts every request into three groups: 2xx done, 4xx the client's fault, 5xx the server's fault. Each step marks the code that its request received on the list of twelve status codes below.
  2. POST /orders sends a body the server cannot parse, so the server answers 400, a 4xx code that says the fault is the client's. Retrying will not help, because the same request will fail the same way again, so the client must change the request, not just send it again.
  3. GET /orders times out at the database. The fault is the server's, so the server answers 500, a 5xx code. The client can retry, because the same request may work in a moment. A second call finds the server overloaded, so the server returns 503, which may carry Retry-After. Both 500 and 503 are marked on the list of status codes.
  4. The server answers the 101st GET /orders in one minute with 429. Code 429 is a 4xx code, so the fault is the client's. But code 429 is the one 4xx where retrying helps. The client may still retry, but only after a wait. The header Retry-After: 30 tells the client when to try again. The list of codes seen so far now includes 429.
  5. GET /orders/7 arrives with no token, so the server does not know who is asking and answers 401. The fix is credentials, not a retry, so the retry gate tells the client to log in first. Now 401 is marked on the list of codes seen so far.
  6. A DELETE /orders/7 request arrives with user 124's token. The server knows exactly who is asking, but it still refuses, so the reply is 403. A retry will not change the answer. 403 joins the codes seen so far.
  7. Two more POST /orders requests get 4xx codes, so the fault is the client's. One request has a body the server cannot read, so the reply is 400 again. The server can read the other body, but qty -1 breaks a rule, so the reply is 422. Now 422 is a second client-fault code, next to 400.
  8. Five codes remain unused: 200, 201 and 204 for success, and 404 when a thing is missing, or the server will not confirm it exists. The fifth is 409, for a conflict with the current state, so the twelve codes are complete and cover almost every API.

© LearnThatStack - diagrams may not be republished without permission.

A request usually gets back one of three kinds of status code: 2xx is done, 4xx is your problem, 5xx is my problem. The first digit tells the caller whose problem this is. The same digit says whether retrying can help.

The server cannot parse this POST body, so the server returns 4xx. Retrying will not help, because 4xx means the same request will fail the same way again. So the client must change the request, not send it again. Code that retries a 400 wastes everyone's time.

A database timeout comes back as a 5xx, and this time a retry can help. A 5xx means the server failed, so the same request may work in a moment. A second call finds the server overloaded, so the server returns a 503, often with a Retry-After header that says how long to wait. Both codes permit a retry.

The 101st call within one minute gets 429, in the 4xx family. But this is the one 4xx where a retry helps, after a wait. The Retry-After header tells the client exactly how long to wait. The client should honour that header.

A request with no token gets 401. The server does not know who is asking, so it asks for a login first. A 401 is a question about identity, and the fix is credentials, not a retry. A retry with no credentials gets the same 401 again.

Now the same request carries user 124's token, and the server returns 403. The server knows exactly who is asking and still refuses. In interviews, 401 and 403 get confused more than any other pair. A 401 means I do not know who you are. A 403 means I know who you are, and the answer is still no. Retrying changes nothing.

Two more POSTs get codes in the 4xx range, the client-error codes. The server cannot read the first body, so the server returns 400 again. The server can read the second body, but that body breaks a rule, for example a qty of minus one. A readable body that breaks a rule gets 422.

Codes 200, 201 and 204 all mean success. Code 404 means the thing is not there, or the API will not confirm that the thing exists. Code 409 means the request conflicts with the current state. Those twelve codes cover almost every API you will design.

The status code is for machines. The body is for the person who has to fix the problem. Here the code is 422, which is correct, so the client takes the error path. But the body only says something went wrong. So the message helps nobody: not the client, not the user, and not the API owner (you).

The error body contains a code, VALIDATION_ERROR, and the client matches that exact string. The code exists for branching, so the wording never changes and the code is never a sentence. Treat the code like an enum you have published. Clients depend on the code from the day it ships.

The message is written for a person: email and password are invalid. No client logic should depend on the message, so you may reword the message at any time. The message is for a developer reading a log. The error code is for a program.

The details list has two entries, email and password. Each entry names its field, so the form can show the error where the user is looking. A validation error with no field name does not say where the problem is. So the user has to search the form for the field that is wrong.

The error body carries a request id. The log line records the same id next to the error message. Support asks for that id, because the id links a complaint in a ticket to a line in your logs. You search by timestamp when the id is missing.

The error body should never include a stack trace. Internal details such as file paths, SQL, and driver messages are not transparency. Those details show an attacker what runs behind the API. RFC 9457, called problem details, is the standard to cite if you want one.

A wrong status code makes the response body useless, because machines stop reading after the code. Let's start with the good version. The server times out on its database, so it responds with a 500. Three machines read that code and did the right thing: the monitor paged on-call, the cache stored nothing, and the client scheduled a retry.

Now the same failure comes back with status 200 instead of 500, and a body that says ok: false, error: db timeout. Someone wanted the client to handle the failure gracefully, so the true error information moved into the body. The failure is identical. But the status code now means success.

The monitor receives a 200 status code and correctly counts it as success. So the monitor's error count drops to zero and the dashboard reads healthy. Meanwhile the server's database timeouts rise to 2, and the outage is still happening. Nobody is paged, because monitoring only looks at status codes, and no alert rule reads JSON bodies.

The next 200 reaches the cache, which stores it for 60 seconds and serves it to the next caller. A cache never stores a 5xx. But a 200 on a GET with a max-age is exactly what a cache is for. The failure now lasts longer than the outage, because the cache followed the max-age that came with that 200.

The 200 status reaches the client, so the client's error handler never runs. Client libraries choose the success path or the error path from the status, before they read the body. So the client treats ok: false as ordinary data and shows an empty screen. The retry that would have fixed the failure never runs.

The server returns 500 again, and the body keeps its detail. All three machines recover at once: the error rate rises and triggers an alert, the cache stores nothing, the client retries. The fix was the status code. The body was never the problem.

The status code is the contract that tells machines whose fault the failure was and whether to retry. The body is the courtesy that tells a developer what to fix. A 200 with an error inside it claims a success that never happened, but three systems believe that 200. So never hide an error inside a 200.

In interviews · 3 questions

Related Questions

← All concepts