Testing Strategies — Pyramid, Contracts, Property-Based & Flakes
Choose the cheapest test that would have caught the bug. Pyramid/Trophy allocate effort; contracts protect boundaries; property/fuzz find edge cases mocks miss; doubles need discipline; flakes destroy trust.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
A story needs tests before it merges
Prefer
Cheapest catching layer
Start at unit or property. Climb to integration, contract, or a thin end-to-end smoke only when the cheaper layer cannot express the failure.
- Pure rules fail in milliseconds and point at one function.
- Boundaries get a contract in CI, not a shared staging circus.
- A handful of journeys stay as smoke. Everything else moves down.
Alternative
Ice-cream cone
Most of the suite is brittle UI automation because QA was asked to own all testing.
- CI takes hours, so people stop running it locally.
- Random failures train the team to ignore red builds.
- Happy-path journeys still miss parsers, SQL, and schema drift.
Climb only when you must
The diamond in the source doc is the same rule, drawn as a ladder so the diagram stays one column.
- 1
Name the failure
Logic, wiring, a service boundary, a user journey, or nondeterminism. - 2
Stay cheap if you can
No I/O and one example: a unit test. An invariant over many inputs: property or fuzz. - 3
Climb for collaborators
Same process or a real database: integration. Another service's shape: a contract. - 4
Add smoke last
A critical journey gets a thin end-to-end check. If the cheaper layer already fails the bug, stop.
Overview
Testing is a product decision about feedback latency, change velocity, and risk tolerance. It is not a checkbox labeled done.
Every layer has a cost curve. Unit tests are cheap to write and run. End-to-end tests are expensive to write, slow to run, and brittle when the UI churns. Coverage that never fails is theater. Coverage that fails for the wrong reason, a flake or an over-mock, trains the team to ignore red CI.
Senior interviews probe the tradeoff. When does a contract beat a full staging end-to-end run? When does a property test beat ten golden fixtures? When does a fake beat a mock?
Ask this before you add a test: what failure modes matter for this change, and which layer detects them with the least ongoing cost?
Pyramid, trophy, and the ice-cream cone
Test pyramid (Mike Cohn, popularized by Martin Fowler). A broad base of unit tests, fewer integration tests, and a thin top of end-to-end or UI tests. It optimizes for fast feedback and a low flake rate. The risk is over-mocking: a "unit" test that lies about real collaborators.
Testing trophy (Kent C. Dodds). The mass shifts toward static analysis plus integration, real modules talking to each other, with a thinner unit band and a still-small end-to-end tip. It optimizes for confidence that wiring works. The risk is an integration suite that grows slow when Testcontainers and a shared database are undisciplined.
Ice-cream cone, the anti-pattern. Inverted: many brittle end-to-end tests, few unit or integration tests. It shows up when QA owns "all testing" through UI automation. Symptoms are hours-long CI, random failures, and fear of refactoring.
| Shape | Strength | Weakness | Prefer when |
|---|---|---|---|
| Pyramid | Fast CI, clear ownership | Over-mock drift | Libraries, domain-heavy services |
| Trophy | Real wiring confidence | Heavier test infra | Frontend and API apps |
| Ice-cream | Looks like user journeys | Slow, flaky, expensive | Never as the default |
Both healthy shapes keep end-to-end thin. They disagree on where the mass sits: isolated units, or integration plus static checks. The ice-cream cone is not a third strategy. It is what you get by accident.
Layer mechanics, including the same discount rule tested three ways, are Test Pyramid and Testing Trophy.
Which layer catches this bug
- Unit. Pure logic: parsers, reducers, pricing rules, state machines. No I/O. Failures point at one function.
- Integration. A repository plus a real database, or Testcontainers. An HTTP handler plus an in-process router. A message consumer plus a broker double. This layer catches wiring and SQL mistakes.
- End-to-end. Critical user journeys only, such as signup, then pay, then receive. Treat them as smoke, not an exhaustive matrix.
- Contract. Consumer and provider agreements for an API or an event schema. This beats shared staging for microservice boundaries. Here that means consumer-driven contracts (Pact) or a schema contract, not change data capture. The other CDC is CDC.
- Property and fuzz. Invariants over a large input space: encoding, serialization, permission algebra. They complement example-based tests. They do not replace a readable example of the happy path.
- Doubles. Stubs, fakes, and spies when a real dependency is slow or nondeterministic. Prefer fakes over mocks for domain repositories.
- Flake control. Isolation, frozen clocks, and quarantine with an owner. Without that, every other layer decays.
Flow
- 1
1. Name the bug or risk
- next2. Pure logic, no I/O
- 2
2. Pure logic, no I/O
- next3. Many inputs: property or fuzz
- 3
3. Many inputs: property or fuzz
- next4. One function: unit test
- 4
4. One function: unit test
- next5. Cross-service: contract
- 5
5. Cross-service: contract
- next6. Same process: integration
- 6
6. Same process: integration
- next7. Critical journey: thin E2E
- 7
7. Critical journey: thin E2E
- next8. Else stop at cheaper layer
- 8
8. Else stop at cheaper layer
Lesson map
Testing Strategies — Pyramid, Contracts, Property-Based & Flakes
Choose the cheapest test that would have caught the bug. Pyramid/Trophy allocate effort; contracts protect boundaries; property/fuzz find edge cases mocks miss; doubles need discipline; flakes destroy trust.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Name the bug or risk"] b["2. Pure logic, no I/O"] c["3. Many inputs: property or fuzz"] d["4. One function: unit test"] a -->|1. Name the bug or risk| b b -->|2. Pure logic, no I/O| c c -->|3. Many inputs: property or fuzz| d
Read the chain as a climb, not as eight mandatory tests. Stop at the first box that can express the failure. A pricing rule that never touches I/O stops at unit or property. A breaking JSON field between two services you own stops at a contract. A checkout journey that depends on cookies and a real browser earns one smoke, and only that smoke.
Rule of thumb: start at unit or property. Climb only when a cheaper layer cannot express the failure.
What the rest of the cluster adds
- Test Pyramid and Testing Trophy — unit versus integration versus end-to-end, and where the mass should sit.
- Contract Testing — consumer-driven Pact versus OpenAPI, AsyncAPI, and Protobuf schemas, verified in CI.
- Property-Based and Fuzz Testing — generators, invariants, shrinking, Hypothesis, and fast-check.
- Test Doubles — dummy, stub, spy, mock, and fake, plus classicist versus mockist.
- Flaky Tests — time, randomness, shared state, async races, order dependence, and quarantine.
- Load, Chaos and Production Validation — shadow traffic, canaries, synthetics, chaos experiments, and game days above the pyramid.
Schema design itself stays in API Design and Protobuf schemas. This cluster only asks whether a contract test is the right net for a boundary. Load, capacity, and coordinated omission stay in Microbenchmarks vs Load Tests. Do not turn a strategy discussion into a profiler or a k6 script.
Heavy end-to-end, and the better split
A large end-to-end suite is closest to what a user does. It can catch CSS, auth cookies, and a CDN misconfiguration. It demos well for stakeholders.
The bill is slow feedback, minutes to hours, so developers stop running the suite locally. Timing, third parties, and shared environments raise the flake rate. Selectors churn and the suite becomes a maintenance job. Happy-path journeys also create false confidence: the story passed, and the parser still corrupts an empty list.
The better split is a few end-to-end smokes, strong integration, contracts at boundaries, and property tests for parsers and rules.
Choose the layer in code
The sandboxes encode the climb. They do not run a browser or a broker. They only name the cheapest layer for a risk you already described.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
How do you choose between a unit test and an integration test?
Answer
If the failure mode is pure logic, write a unit test. If it involves SQL, transactions, serializers, or middleware order, write an integration test with real or containerized collaborators. Do not mock the thing you are trying to trust. A mock of the database encodes your assumption about the query. The bug you fear is that the assumption is wrong.
Why can full unit coverage still ship production bugs?
Answer
Mocks encode assumptions about collaborators. Wrong assumptions stay green. Integration and contract layers catch the interface drift those mocks hide. A coverage percentage does not say whether the collaborator was real.
When do contract tests replace staging end-to-end?
Answer
When the risk is the request, response, or event shape between services you own. Contracts verify provider compatibility in CI without orchestrating a full environment. Staging still helps for infrastructure and multi-team chaos. It should not be the only boundary check. Consumer-driven contracts here are Pact-style agreements, not change data capture.
What is the ice-cream cone, and how do you fix it?
Answer
Too many UI end-to-end tests, too few fast tests. Push assertions down. Extract pure logic for unit and property tests. Add API-level integration. Keep end-to-end for the top journeys. Delete the redundant UI cases that repeat a rule the unit suite already owns.
How does the testing trophy differ from the pyramid?
Answer
The trophy invests more in integration, and in static types and lint, relative to isolated unit tests. The pyramid emphasizes a larger unit base. Both keep end-to-end thin. They disagree on where the mass sits, not on whether a thousand browser tests are a strategy.
A PM wants end-to-end coverage for every story. How do you respond?
Answer
Propose a risk matrix. Critical revenue paths get end-to-end. CRUD variants get API integration. Business rules get unit or property tests. Show the CI time and the flake cost of blanket end-to-end. The goal is a failing test that points at the bug, not a badge that every story has a browser script.
Where do flaky tests fit in the strategy?
Answer
Flakes are a tax on every layer. Without quarantine plus an owner who fixes the root cause, teams mute CI and lose the strategy. Retries that hide a logic bug are not a layer. They are how the suite stops meaning anything. Isolation and clocks are the next pages after doubles.
What is one metric you watch for test-strategy health?
Answer
Median time-to-signal on a pull request, the flake rate of the default suite, and the share of failures that map to real product bugs rather than infrastructure noise. A suite that is red for the wrong reason is worse than a smaller suite that developers trust.
Where does this cluster stop, relative to load tests?
Answer
This cluster chooses the layer that catches a functional bug at the lowest ongoing cost. Load tests ask whether a latency or capacity SLO holds under concurrency. That question lives on the performance load-test page. A k6 script is not a substitute for a contract, and a contract is not a capacity test.
What do you say in the first minute?
Answer
Choose the cheapest test that would have caught the bug. Pyramid and trophy both keep end-to-end thin and disagree on unit versus integration mass. Contracts protect boundaries in CI. Property and fuzz tests search the input space examples miss. Doubles need discipline. Flakes destroy trust faster than a missing test.
Pitfalls
- Treating coverage that never fails as proof the change is safe.
- Mocking the collaborator whose behavior is the actual risk.
- Using shared staging as the only check that two services still agree.
- Adding an end-to-end case for every CRUD variant of the same rule.
- Letting flakes and retries train the team to ignore red CI.
- Calling a load test a test strategy, or a test strategy a capacity plan.
A checkout change touches a discount function, a SQL update, a partner webhook payload, the pay button, and a flaky clock in the trial banner. Name the layer for each risk, and name the one layer you refuse to use as the only net.