GraphQL — Schema Design, N+1, Federation & When Not to Use It
GraphQL is a query language plus a typed schema so clients can ask for the fields they need on one endpoint. This hub is the decision map for schema nullability, resolver N+1, federation, and when REST or gRPC is the better tool.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a GraphQL server return?
Answer
A data object shaped like the selection set, and an errors array when a field fails. Partial data is normal.
L2
When do you choose GraphQL over REST?
Answer
Several clients need different nested shapes, and the team can govern the schema, resolver cost, and query budgets.
L3
When is gRPC the better fit?
Answer
Service-to-service calls, protobuf contracts, and streaming inside the mesh. Browsers are not the first client.
L4
What is N+1 in a resolver?
Answer
A list returns N parents and a child field hits storage once per parent. Batch with a per-request loader or a join.
L5
Why is a non-null field dangerous?
Answer
A null result becomes an error and nulls the nearest nullable parent. Mark a field non-null only when it is guaranteed for the life of the API.
L6
How do you stop an expensive query?
Answer
Depth limits, a complexity score, alias caps, timeouts, rate limits, and persisted queries for first-party apps.
L7
Monolith schema, federation, or no GraphQL?
Answer
One team owns a monolith schema. Federation needs domain teams plus a gateway SLO. Skip GraphQL for CDN-heavy reads, bulk export, and simple partner CRUD.
Failure modes
GraphQL for a CDN-first public GET
POST queries skip naive edge caches. Cacheable resources stay on REST.
Non-null cascade
One missing field blanks a whole card because the parent was also non-null.
Unbounded public query
Depth, aliases, and list multipliers melt the database before AuthN failure would have mattered.
Federation as a fashion
A gateway and composition CI show up before a second team exists.
Misconceptions
GraphQL replaces REST and gRPC.
It is a product graph. REST still wins cacheable resources. gRPC still wins the mesh.
Mutations are automatically idempotent.
Send an idempotency key and dedupe the write the same way any other API would.
A root auth check covers nested fields.
Authorize the field or the row. A viewer check does not hide another user's email.
Interviewer traps
Redesigning REST path names or cursor pagination when the question was the graph.
Point at the API Design cluster, then stay on schema, resolvers, cost, and ownership.
Re-teaching gRPC streaming shapes.
Name unary versus stream in one line and send the listener to the gRPC cluster.
Design scenario
Same prompt for every reader.
Requirements
One round trip for the order screen. Payments stay on the existing RPC. Photos stay cacheable.
Traffic / scale
A few thousand screen loads per minute at peak.
Latency
The screen returns in one client round trip, under a few hundred milliseconds when the stores are healthy.
Consistency
Order status matches the payments service after a successful place-order mutation.
Availability
If recommendations are down, the order header and lines still render.
Failure assumptions
- A client sends a deeply nested query with repeated aliases.
- A child resolver fetches one user per order line.
Constraints
- Do not move payments off gRPC.
- Do not put image bytes on the graph.
Prompt
A mobile app and a web app need an order screen. Payments already speak gRPC. Product photos are CDN-cached.
API
Which root fields are queries, and which mutation places the order?
Data
Where do orders live, and how do you batch the user lookup?
Architecture
Monolith schema, thin BFF, or a federated gateway?
Where the product contract should live
Prefer
GraphQL as a product graph
Several clients select different subtrees of one schema. The team owns nullability, resolver batching, and query budgets.
- One round trip for a nested screen.
- Additive fields instead of a new public URL for every client tweak.
- The cost is field AuthZ, N+1, and weaker HTTP caching.
Alternative
REST resource or gRPC method
A stable cacheable resource stays on REST. An internal high-QPS typed call stays on gRPC. GraphQL is the wrong default for both.
- Public GET plus CDN still wins for media and simple partner CRUD.
- Protobuf streaming and deadlines stay in the gRPC cluster.
- A thin BFF can still sit in front of those leaves.
Decide before you publish a schema
Client shape first, ownership second, cost controls before traffic. Sibling pages hold the mechanics.
- 1
Name the clients
Different nested screens are the GraphQL case. One stable resource is REST. Service-to-service only is gRPC or internal REST. - 2
Own the schema
Non-null is a forever promise. Mutation inputs and payloads evolve by addition. Deprecate before you remove. - 3
Batch the resolvers
A list of parents times a child fetch is N+1. A per-request DataLoader or a join is the fix. - 4
Budget the query
Depth, complexity, aliases, timeouts, and field AuthZ. Persisted queries for first-party apps. - 5
Split only with a platform
Federation when domain teams and a gateway SLO both exist. Otherwise keep one schema or skip GraphQL.
Overview
Interviewers care less about "one endpoint" and more about when the graph is the wrong tool, how nullability breaks clients, and whether you can stop a query from melting a database.
Prefer GraphQL when several product clients need different nested shapes and a team can staff schema review, resolver performance, and abuse controls.
Prefer REST for cacheable public reads, simple partner CRUD, file upload, and bulk export. Path and versioning depth stays in API Design. Page tokens stay in pagination.
Prefer gRPC inside the mesh. Streaming and wire compatibility stay in gRPC and Protobuf. The decision matrix that already exists is gRPC versus REST and GraphQL. Do not rebuild it here.
Why the graph exists
REST makes the server pick the JSON shape. A mobile order screen then either downloads unused fields or makes a second call per line. GraphQL lets the client name the subtree. That power moves work onto the server: every selected field is a function call, and a list multiplies those calls.
Senior loops come back to five questions:
- Can the schema grow without breaking old clients?
- Do you see N+1 before production does?
- What stops a malicious selection set?
- When is a gateway worth the operational tax?
- When should you refuse GraphQL?
GraphQL versus REST versus gRPC
| Dimension | GraphQL | REST (HTTP/JSON) | gRPC (Protobuf) |
|---|---|---|---|
| Contract | Typed schema and operations | Resources and status codes | Service and message definitions |
| Who shapes the payload | The client selection set | The server, plus a few query params | The RPC response |
| HTTP caching | Hard. Persisted queries and a CDN plan | Natural GET, ETag, CDN | Usually internal |
| Streaming | Subscriptions over WebSocket or SSE | Separate SSE or WebSocket | Unary and streams are first-class |
| Change | Add fields, deprecate, then remove | URL or header versions are common | Field numbers and reserved ranges |
| AuthZ | Every field | Route or resource | Method or message |
| Best fit | BFFs and product graphs | Public CRUD and CDN reads | Mesh, low latency, typed streams |
| Refuse when | Bulk export, upload-only, tiny CRUD, CDN media | A screen needs many round trips | A browser is the only client and there is no gateway |
Rule of thumb: GraphQL shines as a product aggregation layer. It is a poor file pipe and a poor CDN.
Decisions
- 1
Product clients?
- NogRPC or internal REST
- YesDifferent shapes?
- 2
gRPC or internal REST
- ?
Different shapes?
- NoREST plus versioning
- YesOwn schema and perf?
- 4
REST plus versioning
- ?
Own schema and perf?
- NoREST plus versioning
- YesOne deployable graph?
- ?
One deployable graph?
- YesMonolith GraphQL
- NoRun a gateway?
- 7
Monolith GraphQL
- nextLimits and DataLoader
- ?
Run a gateway?
- YesFederation or stitch
- NoREST or gRPC plus BFF
- 9
Federation or stitch
- nextLimits and DataLoader
- 10
REST or gRPC plus BFF
- 11
Limits and DataLoader
- nextAuthZ and persisted queries
- 12
AuthZ and persisted queries
Lesson map
GraphQL — Schema Design, N+1, Federation & When Not to Use It
GraphQL is a query language plus a typed schema so clients can ask for the fields they need on one endpoint. This hub is the decision map for schema nullability, resolver N+1, federation, and when REST or gRPC is the better tool.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Product clients?"] b["gRPC or internal REST"] c["Different shapes?"] d["REST plus versioning"] a -->|No| b a -->|Yes| c c -->|No| d
Single column. Each decision has two exits. No subgraph.
When not to use GraphQL
- Bulk download or an analytics extract. The client wants pages of rows. A selection set adds overhead and no product value.
- CDN-first media. A REST GET with cache headers wins. A GraphQL POST misses a naive CDN.
- Ultra-simple public CRUD. One resource and a few fields. Partners read REST more easily.
- File upload pipelines. Multipart REST or a signed URL is simpler than GraphQL upload quirks.
- Mesh-scale fan-out. gRPC streams or a dedicated bus. GraphQL subscriptions fit a product UI, not every telemetry pipe. Realtime protocol choice: WebSockets and MQTT.
- No schema owners. An ungoverned graph becomes one unmaintainable mega-endpoint.
Cluster map
| Lesson | Focus | Then |
|---|---|---|
| Hub (this page) | When to use the graph | Schema design |
| Schema design | Types, nullability, inputs, evolution | Resolvers |
| Resolvers and N+1 | DataLoader, batching, query plans | Operations |
| Operations | Query, mutation, subscription versus REST and gRPC | AuthZ |
| AuthZ and abuse | Field checks, complexity, persisted queries | Federation |
| Federation | Gateway, stitching, ownership | Back to this hub |
Schema is the product contract
Treat the schema like a public API:
- Prefer non-null only when the business guarantees the value for the life of the field.
- Use input types on mutations. Do not hide writes on the query type.
- Ship additive fields first. Deprecate, then delete.
- Teach clients that
errorscan arrive next to partialdata.
Depth: Schema design.
type Query {
viewer: User
order(id: ID!): Order
}
type User {
id: ID!
email: String
orders(first: Int = 20): [Order!]!
}
type Order {
id: ID!
status: OrderStatus!
lines: [OrderLine!]!
}
enum OrderStatus {
PENDING
PAID
SHIPPED
CANCELLED
}
type Mutation {
placeOrder(input: PlaceOrderInput!): PlaceOrderPayload!
}
input PlaceOrderInput {
sku: String!
qty: Int!
idempotencyKey: String!
}viewer stays nullable so anonymous traffic is honest. order may be null when the id is missing. List items are non-null only when a row cannot be a hole. The idempotency key is a hook into idempotency keys, not a second lesson on dedupe storage.
Resolvers and N+1, in one pass
A naive child resolver runs once per parent row. Fix it with a batch function and a cache that live for one request. Measure field latency and query counts, not only the HTTP timer.
Depth: Resolvers and N+1.
AuthZ, cost, and federation, as pointers
- Authenticate at the edge. Authorize in the field or the loader. A root check does not cover nested email or notes.
- Put a depth limit and a cost score on public graphs. Persisted queries for first-party apps.
- A monolith schema is enough until domain teams are real. A gateway is a platform product with its own SLO.
Depth: AuthZ and abuse protection and federation.
Pros and cons
Pros
- Fewer round trips for nested product screens.
- Tooling from a typed schema and introspection.
- One BFF can sit over many REST or gRPC leaves.
Cons
- CDN and HTTP cache semantics are worse than GET.
- N+1 and abusive queries are the default without discipline.
- The AuthZ surface grows with every field.
- Upload, bulk export, and tiny partner APIs are often worse DX.
Sandbox: a tiny cost score
Real servers walk the AST. This sketch only shows the interview idea: list fields cost more than scalar fields, and a multiplier stands in for first.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Sandbox: batch keys once
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Name the clients. Mark which payload is GraphQL, which URL stays REST, and which hop stays gRPC. Point at the field that would N+1, and name the budget you would reject.
Interview Q&A
When would you choose GraphQL over REST?
Answer
When several clients need different nested shapes, and the team can own schema review, resolver performance, and query cost. A single cacheable resource stays on REST. See API Design for names and versions.
When is gRPC a better fit?
Answer
Internal service-to-service calls, protobuf contracts, and streaming. Not a browser-first public API without a gateway. Depth stays in gRPC and gRPC versus REST and GraphQL.
What is the N+1 problem in GraphQL?
Answer
A list resolver returns N parents. A child field resolver hits storage once per parent. Batch with DataLoader or a join-aware planner. The mechanic is the resolvers lesson.
Why is a non-null type dangerous?
Answer
If the resolver returns null, the runtime adds an error and nulls the nearest nullable parent. That can blank a whole card. Only mark a field non-null when the value is guaranteed for the life of the API.
How do you stop expensive queries?
Answer
Depth limits, a cost analysis, alias caps, rate limits, persisted queries or an allowlist, timeouts, and field-level AuthZ. See abuse protection.
Federation, stitching, or a monolith schema?
Answer
A monolith if one team owns the graph. Federation when domain teams own subgraphs and accept gateway operations. Stitching is flexible and harder to standardize. A small system should stay a BFF. See federation.
Are GraphQL mutations automatically idempotent?
Answer
No. Pass an idempotency key and dedupe on the server like any other write. The store and replay rules live in idempotency keys.
Name three reasons not to use GraphQL.
Answer
CDN-heavy public GETs, bulk exports, and a team that cannot staff schema governance plus abuse controls.