System Design Concepts · API Design
Testing APIs: three different questions
Three suites, three questions: behaviour, compatibility, capacity. The contract test is the one only an API needs.
Each layer catches a different bug
- The four layers have fewer tests going up: 1,200 unit tests, 180 integration tests, 24 contract tests and 8 end-to-end journeys. Higher layers need more of the system running before a test means anything, but production, above the deploy line, is not the team's layer.
- A unit test catches formatMoney(0.1 + 0.2) returning "0.30000000000000004" in 0.3 seconds, on the developer's own machine, before the commit exists. Nothing else has to be running, because the bug is arithmetic inside one function, the only kind of bug this layer can see.
- The request GET /orders?status=open takes 6.4 seconds over 2 million rows, because the status column has no index. A unit test with a fake repository never touches an index, but an integration test, with the service and a real database running, finds the problem in 40 seconds.
- A commit renames total_cents to amount_cents and updates the unit and integration tests to match, so both layers still pass. The contract layer checks that consumers still get the fields they read, so it catches the rename in 12 seconds. This layer is high in the test pyramid because it has few tests, not because those tests are slow.
- Only end to end tests catch the login redirect that drops the session cookie, because the whole path is wrong, not any single service. These tests take nine minutes in a browser and fail for reasons unrelated to the code, so this layer should stay small.
- Without the contract layer, the same rename ships, so checkout-web shows 0.00 on every order until a customer reports the bug at 02:14, three days later. Skipping the layer costs three days instead of 12 seconds in the commit's own build, and somebody else finds the bug.
The contract test catches it first
- The consumer, checkout-web, calls GET /orders/8842 on orders-api, the provider in another team's repo, and renders the reply. The provider passes all 212 of its own tests, but its repo never names its consumer, so nothing links the two passing builds.
- The consumer's own test runs against a stub (a fake provider) and records what checkout-web actually read: a 200 and the three keys id, total_cents and status. The endpoint returns seventeen keys, but checkout-web has never touched the other fields and does not depend on them.
- The consumer publishes that recording to a broker as contract version 4, one file with the request, the expected status code and three required body keys. Both builds can fetch the contract from the broker between the two repos, so contract testing works across two teams.
- The provider's build now fetches the contract on every commit and replays it against the real orders-api, beside the team's 212 existing tests. The reply has all three keys, so the contract check passes in 12 seconds.
- A commit in orders-api, the provider, renames total_cents to amount_cents and also updates the provider's tests, so all 212 still pass. Nothing in the orders-api repo shows that anyone reads the old name, so the commit looks clean from inside that repo.
- All 212 provider tests still pass, but the contract check fails because it expects total_cents and gets amount_cents. So the consumer's own recorded usage fails the provider's build before deploy, and no request from checkout-web ever breaks.
- Schema validation is the cheaper option, and it catches any response that stops matching the OpenAPI spec in the provider's own repo. But the same commit also renamed the field in the spec, so the check misses this rename.
Four load shapes, four questions
- A load test holds the expected 800 requests per second for 30 minutes to answer one question: does the service work on a normal day? The p95 latency (the time 95% of requests finish within) stays at 180 ms, under a 300 ms budget, with 0.02% errors.
- A stress test raises the traffic until something breaks, and latency stays under the 300 ms budget up to 3,000 requests per second. Above that rate, latency goes over the budget and never drops back, reaching 4,900 ms at the top of the traffic ramp. So 3,000 requests per second is the ceiling, and the first thing to break is the database pool at 200 connections.
- A spike test raises the traffic from 800 to 5,000 requests per second in ten seconds, then holds it for thirty seconds. The p95 latency reaches 4,100 ms, and 11% of requests fail. But the test measures the 90 seconds the service takes to get back under the latency budget, not the peak.
- An endurance test holds the same normal traffic for eight hours, but p95 latency still rises from 180 ms to 940 ms. The memory heap grows in a straight line from 410 MB to 3.1 GB, a climb that a ten-minute run cannot see at all.
- Every signal says this same load run passed: 1.4 million responses, all with status 200, and p95 at 180 ms. Only the schema assertion failed, because every response body had an empty items field and no total field, so all those 200s proved nothing.
© LearnThatStack - diagrams may not be republished without permission.
Each test layer below the deploy line (the point where code ships) exists to stop a bug before it reaches production. Production is above that line, and it is not a layer you own. The higher the layer, the fewer the tests. There are 1,200 unit tests and 180 integration tests, then 24 contracts and 8 end to end journeys. But each higher layer needs more of the system running before a test means anything.
Speed is the whole reason there are 1,200 unit tests. One of them catches formatMoney(0.1 + 0.2) returning 0.30000000000000004 in 0.3 seconds, on your own machine. The bug is arithmetic inside one function, so the test needs nothing else running. You run these tests before the commit exists and barely notice the wait.
The next bug appears only against a real database. The request GET /orders?status=open takes 6.4 seconds over 2 million rows, because the status column has no index. A unit test with a fake repository never touches an index, so no test below this layer could have seen this problem. The integration test takes 40 seconds because it needs the service and a real database running.
The commit that renamed total_cents to amount_cents also updated the tests in both layers below, so neither layer catches the rename. The contract layer checks that consumers still get what they read. This check catches the rename in 12 seconds, faster than the integration layer. So a layer is high in this pyramid because it has few tests, not because its tests are slow.
The login redirect drops the session cookie. This bug is in the path between services, not inside any one service, so no lower layer of tests can see the bug. End to end tests take nine minutes in a browser and fail for reasons that have nothing to do with your code. So there are 8 end to end tests, not 800.
Skipping a test layer costs you, because the bugs that layer would catch reach production. Now delete the contract test layer and run the same rename again. The rename ships, so checkout-web shows 0.00 on every order until a customer reports the bug at 02:14, three days later. Your own build would have caught the bug in 12 seconds.
Of the four test layers, only an API needs the contract layer. Two teams each have their own repo and a passing build. The consumer, checkout-web, calls GET /orders/8842 and renders the response. The provider, orders-api, passes all 212 of its own tests, but nothing in its repo names a single consumer.
Consumer driven means the specification is what a consumer uses, not everything the API happens to offer. Here the consumer's own test runs against a stub, which is a fake provider. The test records only what checkout-web read: id, total_cents, status and a 200 response. So the specification covers three keys out of the seventeen this endpoint returns.
Both builds fetch the contract from a broker, so neither team has to ask the other for anything. The broker, a versioned file store between the two repos, is what makes contract testing work across teams. The consumer publishes its recording to the broker as version 4. The recording holds the request, the expected status code, and the three keys checkout-web needs in the body.
The provider's build now checks every commit against what checkout-web actually reads. The build fetches the contract and replays it against the real orders-api, with orders-api's own dependencies stubbed. All three keys are there, so the check passes in 12 seconds, alongside the 212 tests the team already had. Pact is the tool most people mean when they say contract testing.
No test in the provider's own suite can ever fail on this rename. The same people write the suite and the code, at the same time. So the commit that renames total_cents to amount_cents also updates the tests. All 212 provider tests still pass, so as far as the provider can tell, the commit is clean.
The provider's build fails on the contract, which is the consumer's own recorded usage. Another team wrote that contract in another repo. The contract check expected total_cents but got amount_cents, so the build stops before deploy. Nothing shipped, so checkout-web is still reading total_cents.
The cheaper option checks every response against your OpenAPI spec, with no broker and nobody to coordinate with. This check catches a response that stops matching the spec. But the check misses this rename, because the same commit also renamed the field in the spec. The spec is yours to change, but the contract belongs to the consumer.
Two suites have now proven behaviour and compatibility. Capacity is a different question. You ask it with the shape of the traffic. A load test keeps the traffic at the level you actually expect: 800 requests per second for 30 minutes. The p95 latency (the time within which 95% of requests finish) stays at 180 ms, under a 300 ms budget. So this run tells you about a normal day and nothing else.
A stress test gives two numbers, and the second one matters more. The first is the ceiling, 3,000 requests per second, where latency goes over the budget as the traffic keeps rising. Latency then stays over the budget and reaches 4,900 ms at the top of the ramp. The second number is the first thing to break, the database pool at 200 connections.
A spike test measures recovery, not the peak. Traffic rises from 800 to 5,000 requests per second in ten seconds and stays there for thirty seconds. Under that load, the p95 reaches 4,100 ms and 11% of requests fail. But the measurement is the 90 seconds the service takes to get back under the budget after the spike ends.
This endurance run finds a leak, and no ten minute run can show a leak's shape. The test sends the same ordinary traffic for eight hours. But p95 latency still rises from 180 ms to 940 ms, and the heap grows in a straight line from 410 MB to 3.1 GB. People skip this test, but it explains why a service needs a restart at 3am.
A test that checks only the status code proves almost nothing. Every signal says the load run passed: 1.4 million responses, all of them 200, and p95 at 180 ms. Only the schema assertion failed, because every response body had an empty items list and no total field. So assert an error rate, the body's shape, and a percentile instead of an average, because an average hides every slow response.
Say which question each suite answers. Functional tests prove behaviour. Contract tests prove compatibility. Performance tests prove capacity. Then name the contract test as the API-specific one. The reason: only the contract test fails your build on a change your own suite has already approved.
In interviews · 4 questions
Related Questions
- 01
- 02
- 03