Resolvers & N+1 — DataLoader, Batching & Query Planning
GraphQL walks the selection set and calls a resolver per field. A child fetch per parent row is the N+1 outage. DataLoader batches those keys inside a single request, and a join or a persisted plan beats the loader when the shape is fixed.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is a resolver?
Answer
The function the runtime calls for one field, with parent, args, context, and info.
L2
What is GraphQL N+1?
Answer
One query for the list, then one query per parent for a child field.
L3
What must a batch function return?
Answer
Values aligned by index with the input keys, including a null or an error for a missing key.
L4
Why is the loader per request?
Answer
A shared cache can leak another tenant's row and can serve stale data after a mutation.
L5
When is a JOIN better?
Answer
When the nested read is predictable and one query returns the tree cheaper than several batches.
L6
Does DataLoader fix authorization?
Answer
No. The batch query still filters by tenant and actor. Coalescing IO is not an auth check.
L7
How do you catch a forgotten loader?
Answer
Query-count budgets in CI or shadow traffic, plus a field span. Request latency alone hides the fan-out.
Failure modes
Forgotten loader on a new field
The schema looks fine. Production query count climbs with list size.
Loader cached across requests
User A receives user B's email, or a mutation's follow-up read is stale.
Unbounded IN list
A page of ten thousand parents becomes one giant batch. Cap maxBatch and chunk.
Cache not primed after a write
Later fields in the same request still see the pre-mutation row.
Misconceptions
DataLoader is a Redis cache.
It is a per-request batch and memo. TTL and invalidation live in the caching cluster, not here.
CDN caching removes N+1.
N+1 happens while building the response. The edge never sees the extra queries.
Returning a map from the batch function is enough.
Classic DataLoader wants the output list in key order. Align it yourself if the SQL result is not ordered.
Interviewer traps
Describing a global cache layer when asked about N+1.
Say request-scoped batch first. Mention Redis only as a different problem.
Claiming federation removes N+1.
Entity hops can recreate it across the network. Batch inside the subgraph and at the gateway.
Design scenario
Same prompt for every reader.
Requirements
One database round trip per entity type per request. No cross-request cache. Missing products become null, not thrown holes, unless the schema says otherwise.
Traffic / scale
Pages of 20 orders, a few lines each, a few thousand screens per minute.
Latency
The list path stays flat as the page size grows from 5 to 20.
Consistency
A mutation in the same request sees the new email if a later field reads that user.
Availability
A failed product batch fails those fields, not the whole order list, when the product field is nullable.
Failure assumptions
- A new reviewer field ships without a loader.
- Someone stores the loader on a global singleton.
Constraints
- Loaders are created in request context.
- Batch size is capped.
Prompt
An order list screen selects the buyer email and each line's product name. Today each child field queries once.
API
Which fields are the parent list and which fields call load?
Data
What is the IN query, and how do you align a missing id?
Architecture
Where is the loader constructed, and what would a join replace?
Overview
For each selected field the runtime calls a resolver. Lists then run child resolvers once per element. orders { user { email } } is one orders query plus N user queries if user is naive.
That is the production outage pattern. The schema did not change. The list got longer.
Flow
- 1
Client sends nested query
- nextRuntime loads parent list
- 2
Runtime loads parent list
- nextChild fields call load
- 3
Child fields call load
- nextLoader coalesces keys this tick
- 4
Loader coalesces keys this tick
- nextOne query with id IN
- 5
One query with id IN
- nextAlign values to key order
- 6
Align values to key order
- nextJSON selection set
- 7
JSON selection set
Lesson map
Resolvers & N+1 — DataLoader, Batching & Query Planning
GraphQL walks the selection set and calls a resolver per field. A child fetch per parent row is the N+1 outage. DataLoader batches those keys inside a single request, and a join or a persisted plan beats the loader when the shape is fixed.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Client sends nested query"] b["Runtime loads parent list"] c["Child fields call load"] d["Loader coalesces keys this tick"] a -->|Client sends nested query| b b -->|Runtime loads parent list| c c -->|Child fields call load| d
DataLoader contract
- The batch function accepts a list of keys and returns a list of values or errors aligned by index. A map is fine only if you re-order it to the keys.
- Scope the loader to one request. Never share the cache across users.
- The per-request memo is optional when a mutation must see a fresh row before the request ends. Clear or prime those keys.
- After a mutation returns an entity, prime the loader so later fields in the same request do not read the old row.
How to fetch the child
Prefer
Per-request batch loader
Resolvers stay field-shaped. Keys coalesce. The schema does not change.
- Works for sparse relations and many independent fields.
- Wrong max batch size still builds a huge IN list.
- Does not replace a join for a dense, known tree.
Alternative
Join or a stored plan
One query returns the tree the screen always asks for.
- Faster when the shape is stable.
- Couples the resolver to that query.
- Persisted operations make the plan honest because the selection set is known.
One request, one basket
- 1
Context is born
The HTTP handler creates loaders once and hangs them on context. - 2
Fields call load
Each child resolver enqueues a key and returns a promise. - 3
The tick flushes
Unique keys go to one batch function, chunked by maxBatch. - 4
Results align
Index i matches key i. Missing keys are null, not a shifted list. - 5
Request ends
The cache dies with the request. The next user gets a new loader.
Where the loader lives
context.loaders.users
context.loaders.products
context.actorResolvers must not construct loaders. A loader built inside a field has its own empty cache and misses the batch. In a federated graph each subgraph has its own context. Batch inside the subgraph. The gateway's entity batch is a second hop, covered in federation.
Beyond the loader
| Approach | When | Tradeoff |
|---|---|---|
| DataLoader | Sparse relations, many resolvers | Easy, still one query per type |
| SQL join | A known dense nested read | Faster, coupled to the shape |
| Persisted query plan | Mobile apps with fixed operations | Best latency, less ad-hoc freedom |
| Denormalized read model | A hot path | Write complexity elsewhere |
| Entity reference batch | Subgraphs | Extra network hop |
Observability
- Field latency histograms, not only the request timer.
- Database query count per operation name.
- Batch size distribution. A batch of one means the loader is not batching.
- Parent-to-child spans. A fan of identical child spans is the picture of N+1.
Sandbox: align by index
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
A page returns 20 orders. Each order has a user and a product. How many queries does the naive resolver issue? How many does a loader issue? What changes if product is a join instead?
Interview Q&A
What is GraphQL N+1?
Answer
Child field resolvers hit storage once per parent in a list. One orders query becomes one plus N user queries.
Why must DataLoader be per request?
Answer
A cache shared across requests risks stale rows and cross-tenant leaks. The memo is an optimization for one selection set, not a system of record.
Must batch results preserve order?
Answer
Yes for the classic API. Value i belongs to key i. If the database returns rows in another order, map them back before you resolve the promises.
Does DataLoader fix authorization?
Answer
No. Batching coalesces IO. The batch query and the field still enforce the actor and the tenant. See AuthZ.
When is a JOIN better than loaders?
Answer
When the nested read is predictable and one query is cheaper than several batched lookups. Loaders win when many independent fields each need a sparse relation.
Can a CDN fix N+1?
Answer
No. N+1 is server work before the response exists. An edge cache can hide a repeated identical response. It cannot merge the queries that built the first one.
How do mutations interact with loaders?
Answer
Clear or prime the mutated keys so later fields in the same request see the write. The next request gets a new loader anyway.
What do you graph in production?
Answer
Query count and batch size per operation name, plus field latency. A flat p50 with a climbing query count is the forgotten loader.