Federation, Gateways & Schema Stitching Tradeoffs
Federation, schema stitching, and a monolith schema are ownership models. A gateway is worth it when domain teams and a platform SLO both exist. It does not remove N+1, field AuthZ, or the need for a cost budget inside each subgraph.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What problem does federation solve?
Answer
Several teams contribute to one client graph without one repo owning every resolver.
L2
What is an entity key?
Answer
The fields that identify an object so another subgraph can reference it and the owner can resolve it.
L3
What does the gateway do at runtime?
Answer
It builds a query plan and fetches subgraphs in sequence or in parallel, then stitches the selection set.
L4
How is stitching different?
Answer
Stitching merges remote schemas with custom delegation. Federation standardizes entities and composition checks.
L5
Does federation remove N+1?
Answer
No. Entity resolution can be one hop per parent unless the gateway and the subgraph batch those references.
L6
What breaks composition?
Answer
Conflicting field types, missing keys, and shared enums that disagree.
L7
When is federation a bad answer?
Answer
Two services, no platform team, a latency budget that cannot pay for hops, or a graph that is still one domain.
Failure modes
Gateway without an owner
Router upgrades, schema CI, and the SLO have no team. Every domain waits on a hero.
Network N+1
Each order triggers an entity fetch to the users subgraph instead of a batched reference query.
Shared type with two writers
Two subgraphs define the same field with different types and composition fails, or worse, the rules are loose and clients see a clash.
Day-one federation
The platform cost arrives before a second team exists. A modular monolith would have shipped.
Misconceptions
Federation is a library choice, not an org chart.
It is an ownership model. The gateway is a product with an on-call.
Stitching and federation are the same diagram.
Both have a gateway. Federation adds a standard entity contract and composition. Stitching is custom glue.
The gateway authorizes every field.
Each subgraph still authorizes its fields and rows. The router is not a substitute for that.
Interviewer traps
Recommending federation for a two-service BFF.
Say monolith schema or a thin BFF. Name the gateway cost you are refusing.
Drawing protobuf streaming as the alternative without a pointer.
Internal RPCs can stay gRPC under any of these graphs. Link the gRPC hub and stop.
Design scenario
Same prompt for every reader.
Requirements
Clients keep a single endpoint if you split. Composition fails closed on type conflicts. Entity fetches are batched.
Traffic / scale
Order screens dominate. User fields are a small slice of each screen.
Latency
A second network hop must fit the screen budget, or the user fields stay in-process.
Consistency
A user reference resolves to the same id the order stored.
Availability
If the users subgraph is down, the order header still returns and the user field errors.
Failure assumptions
- Entity resolution runs once per order.
- Nobody owns router deploys.
Constraints
- Do not federate until a platform owner exists.
- Subgraphs keep their own field AuthZ and loaders.
Prompt
Orders and users are separate teams. The product client wants one graph. There is no platform group yet. Latency budget is one extra hop at most.
API
Which type is an entity, and which field is the key?
Data
How are user references batched back to the users subgraph?
Architecture
What stays a modular monolith until the platform team exists?
Overview
When one team cannot own the whole product graph, you are choosing an ownership model, not a dependency.
| Model | Who owns fields | Runtime | Prefer when | Avoid when |
|---|---|---|---|---|
| Monolith schema | One team, one deployable | Simple | Early product, one domain | Many orgs fight over one schema file |
| Federation | Domain teams own subgraphs | Gateway plus composition | Clear boundaries and a platform team | Nobody will run the router |
| Schema stitching | Whoever writes the merge | Gateway transforms | Odd legacy schemas, a gradual merge | You want one entity standard |
| Many BFFs, no shared graph | Each app team | Per-app GraphQL or REST | Strong app autonomy | Every app must share one graph |
Flow
- 1
Clients
- nextGateway router
- 2
Gateway router
- nextComposed supergraph
- 3
Composed supergraph
- nextUsers service
- nextOrders service
- 4
Users service
- 5
Orders service
- nextBatched entity refs
- 6
Batched entity refs
- nextUsers service
Lesson map
Federation, Gateways & Schema Stitching Tradeoffs
Federation, schema stitching, and a monolith schema are ownership models. A gateway is worth it when domain teams and a platform SLO both exist. It does not remove N+1, field AuthZ, or the need for a cost budget inside each subgraph.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Clients"] b["Gateway router"] c["Composed supergraph"] d["Users service"] a -->|Clients to Gateway router| b b -->|Gateway router to Composed supergraph| c c -->|Composed supergraph| d
The gateway plans a query using the composed schema. An entity lets Orders return a user reference. Users contributes email. Catalog can be a third subgraph in the same shape. The diagram keeps two so the column stays readable.
Federation mechanics
- Subgraphs publish a schema with entity directives. The common teaching words are key, requires, and provides. The exact spellings are spec details, not the interview.
- Composition runs in CI and fails on conflicts.
- The gateway builds a plan: some subgraph calls are sequential, some parallel.
- Entity resolution batches references. If it does not, you have N+1 across the network. The in-process fix is still DataLoader, on each side of the hop.
From subgraph publish to a client response
- 1
Subgraph publishes
A domain team ships its schema and resolvers. - 2
Composition checks
Shared types, keys, and enums must agree. - 3
Router loads the plan
The supergraph says which service owns which field. - 4
Entity batch
References to User go out as one batched lookup, not one call per order. - 5
Client sees one tree
Partial errors still apply if a subgraph fails.
What you gain
- Domain teams ship without a single resolver repo.
- Subgraphs deploy on their own cadence.
- Clients keep one product schema.
What you pay
- The gateway is on the critical path.
- Cross-subgraph fields add hops.
- Someone must govern shared types. Who owns User?
- Local development means a router plus several processes.
- Reading a query plan becomes a skill.
Standard entities or custom glue
Prefer
Federation when the contract is shared
Entity keys and composition give every team the same rules.
- CI can reject a conflicting field.
- Entity batching has a known shape.
- You accept a router and a platform on-call.
Alternative
Stitching when the schemas are odd
Custom delegation merges legacy graphs that do not share an entity model.
- Flexible for a gradual migration.
- Consistency is whatever the gateway code does.
- Harder to staff as the number of merges grows.
Neither option deletes field AuthZ or cost limits. Put those inside each subgraph. The router is not the only place a nested field can leak.
When not to federate
- Two services. A BFF monolith is cheaper.
- No platform team for router upgrades and composition.
- The latency budget cannot buy another hop.
- One domain. Use a modular monolith schema, packages per area, one deploy.
- The real split is already gRPC plus a thin edge. Keep that. See gRPC versus REST and GraphQL.
Start with a modular monolith. Extract a subgraph when team boundaries hurt more than the gateway costs. Do not federate on day one for a microservices aesthetic.
Sandbox: a reference and a conflict
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The Python sketch is a teaching stand-in for composition, not a gateway.
Pitfalls
Two teams, or one team with two packages? If you split, what is the entity key, who runs the router, and what happens to order history when the users subgraph times out?
Interview Q&A
What problem does federation solve?
Answer
Multiple teams contribute to one client-facing graph without one repository owning every resolver.
What is an entity key?
Answer
The fields that uniquely identify an object so other subgraphs can reference it and the owning subgraph can resolve it.
Stitching versus federation?
Answer
Stitching is a flexible merge with custom delegation. Federation is a standardized composition model built around entities and a router.
Does federation eliminate N+1?
Answer
No. It can add a network N+1. Batch entity resolution at the gateway and batch IO inside the subgraph. See resolvers.
Who runs the gateway?
Answer
A platform or API team with an SLO, schema CI, and an on-call. If that team does not exist, do not federate yet.
What is the alternative?
Answer
A modular monolith schema, a per-app BFF, or REST and gRPC with no unified graph. Internal RPC depth stays in gRPC.
What breaks composition?
Answer
Conflicting field types, missing keys, and shared enums that do not agree. CI should fail closed.
When is federation a weak interview answer?
Answer
When the system is small, or the answer cannot name gateway cost, composition, and entity batching. The decision map is the hub.