LearnThatStack Ace your next interview

System Design Concepts · API Design

Testing APIs: three different questions

Three suites, three questions: behaviour, compatibility, capacity. The contract test is the one only an API needs.

testingEach layer catches a different bugand the layer below cannottime to know3 daysfound bya customerproductionfound at 02:14, three days after the deploydeployabove this line, somebody else finds itunit1,200 testsneeds nothing else. Just the functionintegration180 testsneeds the service and a real databasecontractdeletednobody is asking the consumer's questionend to end8 journeysneeds everything, in a browserdoes this function do what it saysformatMoney(0.1 + 0.2) returns "0.30000000000000004"does this endpoint work against the real storeGET /orders?status=open takes 6.4 s over 2 M rowsdoes this still give consumers what they readnothing runs here nowdoes the whole thing work for a personthe login redirect drops the session cookiecheckout-web shows 0.00 on every ordertotal_cents renamed to amount_centsthe layer you skip is the one that charges youthe same rename. 12 seconds in your own build, or three days and somebody else's evening
testingEach layer catches a different bugtime to know3 daysfound bya customerproductioncheckout-web shows 0.00, 02:14deployunit1,200 testsneeds nothing else runningformatMoney(0.1+0.2) = 0.30000000000000004integration180 testsneeds the service and a databaseGET /orders?status=open 6.4 s on 2 M rowscontractdeletednobody is asking the consumer's questionend to end8 journeysneeds everything, in a browserthe login redirect drops the session cookietotal_cents renamed to amount_centsthe layer you skip charges you12 s in your build, or 3 days and a customer

Each layer catches a different bug

  1. The four layers have fewer tests going up: 1,200 unit tests, 180 integration tests, 24 contract tests and 8 end-to-end journeys. Higher layers need more of the system running before a test means anything, but production, above the deploy line, is not the team's layer.
  2. A unit test catches formatMoney(0.1 + 0.2) returning "0.30000000000000004" in 0.3 seconds, on the developer's own machine, before the commit exists. Nothing else has to be running, because the bug is arithmetic inside one function, the only kind of bug this layer can see.
  3. The request GET /orders?status=open takes 6.4 seconds over 2 million rows, because the status column has no index. A unit test with a fake repository never touches an index, but an integration test, with the service and a real database running, finds the problem in 40 seconds.
  4. A commit renames total_cents to amount_cents and updates the unit and integration tests to match, so both layers still pass. The contract layer checks that consumers still get the fields they read, so it catches the rename in 12 seconds. This layer is high in the test pyramid because it has few tests, not because those tests are slow.
  5. Only end to end tests catch the login redirect that drops the session cookie, because the whole path is wrong, not any single service. These tests take nine minutes in a browser and fail for reasons unrelated to the code, so this layer should stay small.
  6. Without the contract layer, the same rename ships, so checkout-web shows 0.00 on every order until a customer reports the bug at 02:14, three days later. Skipping the layer costs three days instead of 12 seconds in the commit's own build, and somebody else finds the bug.

© LearnThatStack - diagrams may not be republished without permission.

Each test layer below the deploy line (the point where code ships) exists to stop a bug before it reaches production. Production is above that line, and it is not a layer you own. The higher the layer, the fewer the tests. There are 1,200 unit tests and 180 integration tests, then 24 contracts and 8 end to end journeys. But each higher layer needs more of the system running before a test means anything.

Speed is the whole reason there are 1,200 unit tests. One of them catches formatMoney(0.1 + 0.2) returning 0.30000000000000004 in 0.3 seconds, on your own machine. The bug is arithmetic inside one function, so the test needs nothing else running. You run these tests before the commit exists and barely notice the wait.

The next bug appears only against a real database. The request GET /orders?status=open takes 6.4 seconds over 2 million rows, because the status column has no index. A unit test with a fake repository never touches an index, so no test below this layer could have seen this problem. The integration test takes 40 seconds because it needs the service and a real database running.

The commit that renamed total_cents to amount_cents also updated the tests in both layers below, so neither layer catches the rename. The contract layer checks that consumers still get what they read. This check catches the rename in 12 seconds, faster than the integration layer. So a layer is high in this pyramid because it has few tests, not because its tests are slow.

The login redirect drops the session cookie. This bug is in the path between services, not inside any one service, so no lower layer of tests can see the bug. End to end tests take nine minutes in a browser and fail for reasons that have nothing to do with your code. So there are 8 end to end tests, not 800.

Skipping a test layer costs you, because the bugs that layer would catch reach production. Now delete the contract test layer and run the same rename again. The rename ships, so checkout-web shows 0.00 on every order until a customer reports the bug at 02:14, three days later. Your own build would have caught the bug in 12 seconds.

Of the four test layers, only an API needs the contract layer. Two teams each have their own repo and a passing build. The consumer, checkout-web, calls GET /orders/8842 and renders the response. The provider, orders-api, passes all 212 of its own tests, but nothing in its repo names a single consumer.

Consumer driven means the specification is what a consumer uses, not everything the API happens to offer. Here the consumer's own test runs against a stub, which is a fake provider. The test records only what checkout-web read: id, total_cents, status and a 200 response. So the specification covers three keys out of the seventeen this endpoint returns.

Both builds fetch the contract from a broker, so neither team has to ask the other for anything. The broker, a versioned file store between the two repos, is what makes contract testing work across teams. The consumer publishes its recording to the broker as version 4. The recording holds the request, the expected status code, and the three keys checkout-web needs in the body.

The provider's build now checks every commit against what checkout-web actually reads. The build fetches the contract and replays it against the real orders-api, with orders-api's own dependencies stubbed. All three keys are there, so the check passes in 12 seconds, alongside the 212 tests the team already had. Pact is the tool most people mean when they say contract testing.

No test in the provider's own suite can ever fail on this rename. The same people write the suite and the code, at the same time. So the commit that renames total_cents to amount_cents also updates the tests. All 212 provider tests still pass, so as far as the provider can tell, the commit is clean.

The provider's build fails on the contract, which is the consumer's own recorded usage. Another team wrote that contract in another repo. The contract check expected total_cents but got amount_cents, so the build stops before deploy. Nothing shipped, so checkout-web is still reading total_cents.

The cheaper option checks every response against your OpenAPI spec, with no broker and nobody to coordinate with. This check catches a response that stops matching the spec. But the check misses this rename, because the same commit also renamed the field in the spec. The spec is yours to change, but the contract belongs to the consumer.

Two suites have now proven behaviour and compatibility. Capacity is a different question. You ask it with the shape of the traffic. A load test keeps the traffic at the level you actually expect: 800 requests per second for 30 minutes. The p95 latency (the time within which 95% of requests finish) stays at 180 ms, under a 300 ms budget. So this run tells you about a normal day and nothing else.

A stress test gives two numbers, and the second one matters more. The first is the ceiling, 3,000 requests per second, where latency goes over the budget as the traffic keeps rising. Latency then stays over the budget and reaches 4,900 ms at the top of the ramp. The second number is the first thing to break, the database pool at 200 connections.

A spike test measures recovery, not the peak. Traffic rises from 800 to 5,000 requests per second in ten seconds and stays there for thirty seconds. Under that load, the p95 reaches 4,100 ms and 11% of requests fail. But the measurement is the 90 seconds the service takes to get back under the budget after the spike ends.

This endurance run finds a leak, and no ten minute run can show a leak's shape. The test sends the same ordinary traffic for eight hours. But p95 latency still rises from 180 ms to 940 ms, and the heap grows in a straight line from 410 MB to 3.1 GB. People skip this test, but it explains why a service needs a restart at 3am.

A test that checks only the status code proves almost nothing. Every signal says the load run passed: 1.4 million responses, all of them 200, and p95 at 180 ms. Only the schema assertion failed, because every response body had an empty items list and no total field. So assert an error rate, the body's shape, and a percentile instead of an average, because an average hides every slow response.

Say which question each suite answers. Functional tests prove behaviour. Contract tests prove compatibility. Performance tests prove capacity. Then name the contract test as the API-specific one. The reason: only the contract test fails your build on a change your own suite has already approved.

In interviews · 4 questions

Related Questions

Also helps with

← All concepts