Test Doubles — Mocks, Stubs, Fakes, Spies & When Not to Mock
Mock/stub/fake/spy/dummy (Meszaros). Over-mocking couples tests to implementation; prefer fakes/testcontainers for domain logic.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
The repository is the collaborator under test
Prefer
Fake, then assert state
An in-memory repository implements the port for real. The test saves, renames, and reads back through the same API.
- The domain service does not know the double is fake.
- A wrong query shape fails if you later run the same test on Testcontainers.
- You assert the email that was stored, not the call transcript.
Alternative
Mock every method
The test pre-declares calls and returns. It passes when the production SQL is wrong, as long as the mock was told to return the row.
- Call order becomes part of the spec.
- Refactors that preserve behavior still break the test.
- Coverage looks high because the mock never executes the branch you fear.
Pick the double in order
The source diagram is a decision tree. This column is the order to consider them so the picture stays one rectangle wide.
- 1
Verify the call itself
A spy or a mock when the interaction is the behavior, such as 'capture was invoked once.' - 2
Need real behavior
An in-memory fake for multi-step domain logic. Testcontainers when SQL semantics matter. - 3
Only a canned value
A stub returns data. A dummy fills an unused parameter. Neither records calls.
Overview
Test doubles stand in for real collaborators. Gerard Meszaros named five: dummy, stub, spy, mock, and fake. Used well, they keep unit tests fast. Used poorly, especially as a reflex to mock everything, they couple tests to implementation details and green-wash bugs that integration would catch.
Prefer fakes, such as an in-memory repository, for domain logic. Reach for Testcontainers when the collaborator is the risk: SQL, Redis semantics, a transaction. Know the classicist and mockist styles, and when each fits.
People say "mock" for all five. In an interview, the precise word is the point.
The five doubles
| Double | Role | What you usually assert |
|---|---|---|
| Dummy | Passed in, never used | Nothing |
| Stub | Returns canned data | State of the system under test |
| Spy | Records how it was called, and often stubs too | That save was called with this value |
| Mock | Expectations up front, fails on unexpected calls | Interaction: calls and order |
| Fake | A working lightweight implementation | State, through the fake's own API |
A dummy exists because a constructor demands a logger you will not touch. A stub returns a user so you can test a formatter. A spy lets the call happen and then you ask what was recorded. A mock fails the test if an unexpected method runs, which is powerful at a process boundary and harmful in the middle of a domain model. A fake is a real implementation with a cheaper engine: a dictionary instead of Postgres.
Flow
- 1
1. Collaborator needed
- next2. Verify calls: mock or spy
- 2
2. Verify calls: mock or spy
- next3. Behavior: in-memory fake
- 3
3. Behavior: in-memory fake
- next4. Return value only: stub
- 4
4. Return value only: stub
- next5. Unused argument: dummy
- 5
5. Unused argument: dummy
- next6. SQL semantics: real DB
- 6
6. SQL semantics: real DB
Lesson map
Test Doubles — Mocks, Stubs, Fakes, Spies & When Not to Mock
Mock/stub/fake/spy/dummy (Meszaros). Over-mocking couples tests to implementation; prefer fakes/testcontainers for domain logic.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Collaborator needed"] b["2. Verify calls: mock or spy"] c["3. Behavior: in-memory fake"] d["4. Return value only: stub"] a -->|1. Collaborator needed| b b -->|2. Verify calls: mock or spy| c c -->|3. Behavior: in-memory fake| d
Read the chain as a filter, not as "always use all six." If you need to verify a call, stop at a spy or a mock and watch for over-mocking. If you need realistic multi-step behavior, stop at a fake. If the risk is the database engine, replace the fake with Testcontainers and stop mocking cursor.execute.
Classicist and mockist
Classicist, sometimes called Detroit style. Use real objects and fakes. Assert the final state. The test does not care which private helper ran.
Mockist, sometimes called London style. Put mocks around the neighbors. Assert the messages that were sent.
The senior split is classicist for domain cores, and selective mocks at process boundaries. A payment capture should hit a sandbox or a fake gateway. It should not charge a real card from a unit test, and it should not mock your own pricing function. Hexagonal ports make this natural: the domain depends on a repository port, and the test supplies an in-memory adapter. Production supplies Postgres. The same port is what you point at Testcontainers when SQL is the question.
Over-mocking owned types hides the bug. A suite with 95 percent coverage and a mocked repository still has not asked what catches a wrong WHERE. Ask that in review.
A fake repository you can run
No Vitest and no network. The spy beside the fake shows the difference: the fake is queried for state, the spy is queried for calls. Production code would take the same port.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
A Vitest suite would wrap the same fake in describe and expect. vi.fn() is a spy. Use it at the edge, for example to assert a mailer was called, not to replace UserService method by method.
Interview Q&A
What is the difference between a mock and a stub?
Answer
A stub returns canned data so the system under test can run. You assert state afterward. A mock is pre-programmed with expectations and fails if the calls do not match, including calls you did not expect. If you only needed a return value, a mock is a stricter tool than the test required.
What is a fake?
Answer
A working lightweight implementation. An in-memory database, a dictionary-backed repository, or a fake payment gateway that applies the same rules without calling the network. You assert through its real API. It is not a recording of expected calls.
When is mocking harmful?
Answer
When you mock types you own, or you mock the database and then claim the query is tested. The test locks the current call sequence. A refactor that still saves the right row goes red. A wrong WHERE stays green because the mock never ran SQL. That is how high coverage ships the bug.
Classicist or mockist?
Answer
Classicist, with fakes, for the domain core. Assert final state. Mockist at I/O edges where the message you send is the behavior: an HTTP call to a provider, a message published to a bus. Most codebases want both, with the mockist layer kept thin.
How do you test a payment capture?
Answer
Use the vendor sandbox or a fake gateway that honors the same success and decline rules. Never charge a real card from the unit suite. Do not mock your own domain object that decides the amount. If the risk is the provider's schema, add a contract. If the risk is your amount calculation, that function should be real.
Dummy or stub?
Answer
A dummy is passed and never used. It fills a parameter list. A stub supplies values the system under test reads. If the test fails when you pass null, it was not a dummy.
The suite shows 95 percent coverage and the repository is mocked. What do you ask?
Answer
What fails if the WHERE clause is wrong? If the answer is "nothing in this suite," the coverage number is not evidence. Add a fake that stores rows, or Testcontainers, and assert the row you read back. Keep a mock only if the interaction itself is the requirement.
How do hexagonal ports change the choice?
Answer
The domain talks to a port. Tests plug in an in-memory adapter. Production plugs in Postgres. You do not mock inside the domain. When engine semantics matter, the same port points at Testcontainers. The design makes the fake the default double instead of a special case.
Spy or mock for 'was save called with this email'?
Answer
A spy records the call and lets you assert afterward. A mock demands the expectation before the call and fails on surprises. For one argument check, a spy is enough. A mock is justified when unexpected extra calls are themselves a bug, such as a double charge.
What do you say in the first minute?
Answer
Dummy, stub, spy, mock, fake. Over-mocking couples the test to call sequences and hides SQL bugs. Prefer an in-memory fake for domain logic. Use Testcontainers when the engine is the risk. Classicist in the core, mockist at the boundary.
Pitfalls
- Using "mock" for all five and then picking the strictest one by habit.
- Mocking the repository, the clock, and the domain service in the same test.
- Asserting call order on code you own and refactor often.
- Skipping Testcontainers and also skipping a fake, so nothing stores a row.
- Hitting a real payment API from the unit job.
- Treating a fake's happy path as proof of SQL dialect behavior.
You have a test that stubs repo.get and expects repo.save with a hard-coded email. Rewrite the assertion as state on an in-memory repository. Say which one behavior you would still spy: the outbound mailer, or the repository.