CQRS — Commands, Queries, Projections & Consistency
CQRS splits commands that change state from queries shaped for screens. In an event-driven system the read side is usually a projection, and lag is an SLO rather than a stuck write.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does CQRS separate?
Answer
Commands that change state from queries that answer screens. Each model is shaped for its job.
L2
Do you need two databases?
Answer
No. Two modules or two tables inside one service can be CQRS. Separate stores show up when scale or shape diverges.
L3
Does CQRS require event sourcing?
Answer
No. The write model can be an ordinary store. Events or CDC update the projection.
L4
How do you explain lag to a product manager?
Answer
The write path is the source of truth. The read path has a freshness SLO. Critical screens read the write model or wait for a version.
L5
How does a projector avoid applying an older event over a newer row?
Answer
Idempotent upserts keyed by business id, applied only when the event version is newer than the row.
L6
How do you rebuild a corrupted read model?
Answer
Replay from the event log or a CDC checkpoint. The version gate still rejects older facts during catch-up.
L7
Where does cache invalidation fit?
Answer
An event can signal a delete or fill a cache. The invalidation lesson owns TTL, event-driven deletes, and versioned keys.
Failure modes
Stale read after a successful command
The write committed and the projection has not applied that version yet.
Late event overwrites a newer row
The projector upserts without a version check, so an older fact wins.
Read model with no rebuild
A bug in the projector cannot be repaired because nobody can replay the log.
Misconceptions
CQRS means a microservice per model.
You can split command and query modules inside one service.
CQRS requires event sourcing.
Event sourcing is one way to feed projections. A normal write model plus events is enough.
Interviewer traps
Promising read-your-writes on the projection without a version wait.
Return the write model, or poll until the read model version catches the command.
Inventing a third cache strategy on the whiteboard.
Point at the cache-invalidation lesson and keep this answer on projections.
Design scenario
Same prompt for every reader.
Requirements
The status page must not claim the order is still new once version 7 is visible. Search can lag. A corrupted status table must be rebuildable.
Traffic / scale
About 5k commands per second at peak, and several times that in status reads.
Latency
Status reads that can tolerate lag stay on the projection. A client that just submitted waits until version 7 or times out to the write model.
Consistency
The write model is authoritative. Each projection applies an event only when its version is newer.
Availability
A crashed projector grows lag and alerts on freshness. Commands still commit.
Failure assumptions
- The projector can crash after the command has returned.
- Events can be redelivered.
- One read model can be rebuilt without stopping writes.
Constraints
- Do not serve the status page from an unversioned cache fill.
- Partition projection work by order id so one order stays ordered.
Prompt
Checkout returns 202 with version 7. The order-status page reads a projection. Email and search use other projections of the same event.
API
What does the command return so the client knows which version to wait for?
Data
What does the status row store besides the status string?
Architecture
Which component reads the outbox, and which screens are allowed to be stale?
The status page is not the ledger
Prefer
Write model for commands, projection for the screen
The command commits invariants and an outbox row. A worker upserts a query-shaped row. The screen reads that row and shows its version.
- Hot reads do not take locks on the write path.
- Search, status, and a customer timeline can each have a shape.
- Lag is visible. Freshness is an SLO.
Alternative
One CRUD model for every query
Transactions are simpler and read-your-writes is natural. The schema becomes a compromise, and the status page competes with checkout.
- Awkward queries pile indexes onto the write table.
- A read replica still shares that schema.
- Cache-aside alone does not give you a rebuild story.
Command succeeds, screen catches up
The 202 is not a lie. The projection is behind.
- 1
Commit on the write model
Validate the invariant. Persist the order and the outbox row together. - 2
Return the version
The client learns which version must appear before the status page is trustworthy. - 3
Project in order per entity
Apply the event only when its version is newer. Duplicates and late copies no-op. - 4
Serve the screen
If the row is behind, say so, poll, or fall back to the write model for this session. - 5
Alert on freshness
A dead projector is lag, not a failed command. Page on the SLO, then rebuild if the model is wrong.
Overview
CQRS (Command Query Responsibility Segregation) splits the write model from read models. Commands enforce invariants and emit facts. Queries answer screens from data shaped for that screen.
This page is the architecture layer: consistency dials, projection design, and the interview tradeoff. It is not a Kafka tutorial and not an event-sourcing course. Payload styles are the previous lesson. Broker partitions are Apache Kafka — Topics, Partitions, Brokers & Consumer Groups.
CQRS is not microservices. Two modules or two tables inside one service are enough. CQRS is not event sourcing. The write model can be an ordinary store. Events or CDC update the projection.
Commands and queries
| Side | Job | Typical store | Failure you should name |
|---|---|---|---|
| Command | Validate invariants, persist, emit | OLTP database or an aggregate store | Reject a bad command. Lose the event if publish is a second write |
| Query | Answer a screen quickly | SQL read model, document store, search, cache | Stale reads. Expensive rebuilds. No version gate |
The dual-write on the command side is the next lesson. Here, assume the event leaves through an outbox and ask what the reader is allowed to see.
Consistency dial
Decisions
- 1
1. User issues a command
- next2. Write model commits
- 2
2. Write model commits
- next3. Event leaves through the outbox
- 3
3. Event leaves through the outbox
- next4. Projection worker applies it
- 4
4. Projection worker applies it
- next5. Read model updates
- 5
5. Read model updates
- next6. Does this UI need read-your-writes
- ?
6. Does this UI need read-your-writes
- yes7. Read the write model or wait for the version
- no8. Read the eventually consistent view
- 7
7. Read the write model or wait for the version
- 8
8. Read the eventually consistent view
Lesson map
The read can lag version 7
Version 7 is committed. The projector has not applied it. The read model still says NEW.
Architecture. Write model Accepted. Projector Behind. Read model NEW
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB write["Write model Accepted"] projector["Projector Behind"] read["Read model NEW"] write -->|Version 7| projector projector -->|Upsert| read read -->|Read-your-writes| write projector -->|Lag| read
Three ways to get read-your-writes without pretending the projection is synchronous:
- After the command, return the representation from the write model.
- The client polls until the read model version is at least the command version.
- A short-lived session cache of recent writes, still keyed by that version.
Product copy should match the dial. "Submitted" from the command response is honest. "Status: new" from a lagging page needs a freshness story, not a bug ticket against the write path.
Projection design
- Idempotent upserts keyed by business id plus event version or offset.
- Ordered apply per entity. The partition key is the entity id. Cross-partition views are unordered on purpose. The ordering lesson owns that choice. Throughput math for Kafka keys stays on the Kafka topics lesson.
- A rebuild story. Replay from the start of the log, or from a snapshot plus catch-up. The version gate still drops older events.
- Fan-out. One event type can feed search, analytics, and a customer timeline. Isolate those workers so one bad projector does not freeze the others. Blast radius is the failure-modes lesson.
Where CQRS sits next to simpler reads
| Approach | Strength | Cost |
|---|---|---|
| Single CRUD model | Simple transactions and easy read-your-writes | Awkward queries and lock contention |
| Read replica of the write database | Familiar operations | Schema stays coupled. Limited denormalization |
| CQRS projections | Query-shaped data and isolated load | Lag, a second schema, and repair jobs |
| Cache-aside only | Fast hot keys | Invalidation is its own problem |
Cache-aside is not a rejected idea. It is a different lesson: Cache Invalidation — TTL vs Event-Driven vs Versioned Keys. Events can invalidate a key or fill a cache. Do not invent a third invalidation theory while designing the projection.
Lag the user can see
Sequence
- 1
User → Command API
1. Submit order
- 2
Command API → Write DB
2. Persist the order and the outbox row
- 3
Command API → User
3. Accepted at version 7
- 4
User → Read API
4. Get status
- 5
Read API → User
5. Still NEW because of lag
- 6
Projector → Write DB
6. Read outbox event version 7
- 7
Projector → Read API
7. Upsert status SUBMITTED
- 8
User → Read API
8. Get status again
- 9
Read API → User
9. SUBMITTED
- 10
Projector
A crashed projector grows lag. Alert on freshness.
Step 5 is not a failed submit. Step 3 already told the truth. The page owes the user a way to wait for version 7 or to read the write model.
Version-gated upsert (run this)
An older event returns false and leaves the row alone. That is the whole correctness rule for a projection that sees duplicates and delays.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Wait for the command version (run this)
The helper polls an in-memory read model. The first call gives up while the row is still at version 1. After the projector writes version 2, the same wait returns the row. Production sleeps with a budget. This sketch only counts attempts so it can run in the page.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Define CQRS.
Answer
Separate models for commands and queries. Each one is optimized for its job.
Must you use two databases?
Answer
No. Two tables or two modules can suffice. Separate stores appear when scale or the shape of the data diverges.
How do you explain lag to a product manager?
Answer
The write path is the source of truth. The read path has an SLO, for example a p99 freshness under 2 seconds. A screen that cannot tolerate that reads the write model or waits on the version.
How do you repair a corrupted read model?
Answer
Rebuild from the event log or a CDC checkpoint. Version gates stop older events from clobbering newer state during the replay.
Where does cache invalidation fit?
Answer
Events can signal a delete or populate a cache. Use Cache Invalidation — TTL vs Event-Driven vs Versioned Keys. Do not invent a third strategy in the middle of a CQRS answer.
What do you return from the command?
Answer
An acceptance and the version you just committed. The client uses that version to decide whether the status page has caught up.
Can one event feed many read models?
Answer
Yes. Search, analytics, and a timeline can each subscribe. Isolate their lag and their failures. A shared consumer that does all three is one blast radius.
Why partition the projector by order id?
Answer
So one order's events apply in order. Other orders stay parallel. Global order is not required for a status row. The ordering lesson is the full argument.
Pitfalls
Draw the command API, the write row, the outbox, one projector, and the status API. Write the version on the 202 and on the status row. Circle the only boxes a user may trust before those versions match.