System Design Concepts · API Design
Status codes and error design: how an API says NO
The status code tells machines whose fault the failure was and whether to retry. The error body tells a developer what to fix, and it is never an apology.
The code routes the failure
- The first digit of the status code sorts every request into three groups: 2xx done, 4xx the client's fault, 5xx the server's fault. Each step marks the code that its request received on the list of twelve status codes below.
- POST /orders sends a body the server cannot parse, so the server answers 400, a 4xx code that says the fault is the client's. Retrying will not help, because the same request will fail the same way again, so the client must change the request, not just send it again.
- GET /orders times out at the database. The fault is the server's, so the server answers 500, a 5xx code. The client can retry, because the same request may work in a moment. A second call finds the server overloaded, so the server returns 503, which may carry Retry-After. Both 500 and 503 are marked on the list of status codes.
- The server answers the 101st GET /orders in one minute with 429. Code 429 is a 4xx code, so the fault is the client's. But code 429 is the one 4xx where retrying helps. The client may still retry, but only after a wait. The header Retry-After: 30 tells the client when to try again. The list of codes seen so far now includes 429.
- GET /orders/7 arrives with no token, so the server does not know who is asking and answers 401. The fix is credentials, not a retry, so the retry gate tells the client to log in first. Now 401 is marked on the list of codes seen so far.
- A DELETE /orders/7 request arrives with user 124's token. The server knows exactly who is asking, but it still refuses, so the reply is 403. A retry will not change the answer. 403 joins the codes seen so far.
- Two more POST /orders requests get 4xx codes, so the fault is the client's. One request has a body the server cannot read, so the reply is 400 again. The server can read the other body, but qty -1 breaks a rule, so the reply is 422. Now 422 is a second client-fault code, next to 400.
- Five codes remain unused: 200, 201 and 204 for success, and 404 when a thing is missing, or the server will not confirm it exists. The fifth is 409, for a conflict with the current state, so the twelve codes are complete and cover almost every API.
Error parts go to their fields
- The status code 422 is right, but the error body says only that something went wrong. The client handles the reply as an error, so the client can show only one banner for the whole form.
- The error body now has a code, VALIDATION_ERROR, and the second case in the client's switch matches that fixed string. The client branches on the code, so the code's wording never changes.
- The message is in plain words for a person to read: email and password are invalid. The message goes only to the log line. The message never reaches the switch where the client branches, so a change to the message's wording breaks nothing.
- The details field lists the two inputs that failed: email and password. Each entry names its own input field, so the client no longer needs one grey banner over the whole form.
- The server adds a request id to the error body and writes the same id in its log line, next to the message. The request id is the one thing support will ask for.
- The server does not add a stack trace as a fifth field of the error body. Internal details in an error body show an attacker how the server works, so sending them is not transparency.
A 200 with an error is a lie
- The server's database call times out, so the server answers GET /users/123 with a 500. That 500 reaches three machines: the monitor counts one more error and pages on-call, the cache stores nothing, and the client will retry in 2 seconds.
- The status code changes from 500 to 200. The body becomes {"ok": false, "error": "db timeout"}. The failure is the same, but the response now has a status code that means success.
- The 200 reaches the monitor. The monitor's error count drops to zero, so the dashboard reads healthy. The server's db timeouts rise to 2, but the monitor pages nobody.
- The next 200 reaches the cache, and the cache keeps a copy, because a 200 on a GET with max-age=60 is exactly what a cache stores. So the cache serves that copy to the next caller, and the cache still serves the failure after the outage ends.
- The 200 reaches the client, and the client's error handler never runs, because client libraries check the status before they read the body. So the client treats the error body as data, and the retry that would have fixed the failure never happens.
- The status goes back to 500, and the body keeps its detail. That one status line reaches all three machines again and fixes all three: the error rate rises, the cache stores nothing, and the client retries.
© LearnThatStack - diagrams may not be republished without permission.
A request usually gets back one of three kinds of status code: 2xx is done, 4xx is your problem, 5xx is my problem. The first digit tells the caller whose problem this is. The same digit says whether retrying can help.
The server cannot parse this POST body, so the server returns 4xx. Retrying will not help, because 4xx means the same request will fail the same way again. So the client must change the request, not send it again. Code that retries a 400 wastes everyone's time.
A database timeout comes back as a 5xx, and this time a retry can help. A 5xx means the server failed, so the same request may work in a moment. A second call finds the server overloaded, so the server returns a 503, often with a Retry-After header that says how long to wait. Both codes permit a retry.
The 101st call within one minute gets 429, in the 4xx family. But this is the one 4xx where a retry helps, after a wait. The Retry-After header tells the client exactly how long to wait. The client should honour that header.
A request with no token gets 401. The server does not know who is asking, so it asks for a login first. A 401 is a question about identity, and the fix is credentials, not a retry. A retry with no credentials gets the same 401 again.
Now the same request carries user 124's token, and the server returns 403. The server knows exactly who is asking and still refuses. In interviews, 401 and 403 get confused more than any other pair. A 401 means I do not know who you are. A 403 means I know who you are, and the answer is still no. Retrying changes nothing.
Two more POSTs get codes in the 4xx range, the client-error codes. The server cannot read the first body, so the server returns 400 again. The server can read the second body, but that body breaks a rule, for example a qty of minus one. A readable body that breaks a rule gets 422.
Codes 200, 201 and 204 all mean success. Code 404 means the thing is not there, or the API will not confirm that the thing exists. Code 409 means the request conflicts with the current state. Those twelve codes cover almost every API you will design.
The status code is for machines. The body is for the person who has to fix the problem. Here the code is 422, which is correct, so the client takes the error path. But the body only says something went wrong. So the message helps nobody: not the client, not the user, and not the API owner (you).
The error body contains a code, VALIDATION_ERROR, and the client matches that exact string. The code exists for branching, so the wording never changes and the code is never a sentence. Treat the code like an enum you have published. Clients depend on the code from the day it ships.
The message is written for a person: email and password are invalid. No client logic should depend on the message, so you may reword the message at any time. The message is for a developer reading a log. The error code is for a program.
The details list has two entries, email and password. Each entry names its field, so the form can show the error where the user is looking. A validation error with no field name does not say where the problem is. So the user has to search the form for the field that is wrong.
The error body carries a request id. The log line records the same id next to the error message. Support asks for that id, because the id links a complaint in a ticket to a line in your logs. You search by timestamp when the id is missing.
The error body should never include a stack trace. Internal details such as file paths, SQL, and driver messages are not transparency. Those details show an attacker what runs behind the API. RFC 9457, called problem details, is the standard to cite if you want one.
A wrong status code makes the response body useless, because machines stop reading after the code. Let's start with the good version. The server times out on its database, so it responds with a 500. Three machines read that code and did the right thing: the monitor paged on-call, the cache stored nothing, and the client scheduled a retry.
Now the same failure comes back with status 200 instead of 500, and a body that says ok: false, error: db timeout. Someone wanted the client to handle the failure gracefully, so the true error information moved into the body. The failure is identical. But the status code now means success.
The monitor receives a 200 status code and correctly counts it as success. So the monitor's error count drops to zero and the dashboard reads healthy. Meanwhile the server's database timeouts rise to 2, and the outage is still happening. Nobody is paged, because monitoring only looks at status codes, and no alert rule reads JSON bodies.
The next 200 reaches the cache, which stores it for 60 seconds and serves it to the next caller. A cache never stores a 5xx. But a 200 on a GET with a max-age is exactly what a cache is for. The failure now lasts longer than the outage, because the cache followed the max-age that came with that 200.
The 200 status reaches the client, so the client's error handler never runs. Client libraries choose the success path or the error path from the status, before they read the body. So the client treats ok: false as ordinary data and shows an empty screen. The retry that would have fixed the failure never runs.
The server returns 500 again, and the body keeps its detail. All three machines recover at once: the error rate rises and triggers an alert, the cache stores nothing, the client retries. The fix was the status code. The body was never the problem.
The status code is the contract that tells machines whose fault the failure was and whether to retry. The body is the courtesy that tells a developer what to fix. A 200 with an error inside it claims a success that never happened, but three systems believe that 200. So never hide an error inside a 200.
In interviews · 3 questions
Related Questions
- 01
- 02
- 03