Studies
Topic, then cluster, then study. Recently added is the short list at the top.
Recently added
Show more- 6.Engineering Compliance into Systems - Data Classification, Audit Logs, Retention, JIT Access, Policy-as-Code & Continuous EvidenceBuilding compliance into the platform: data classification tags and lineage, key ownership, tamper-evident WORM audit logs, retention with legal hold, access reviews vs JIT access, policy-as-code gates, continuous evidence (how to evaluate Vanta, Drata and Secureframe), vendor risk, a breach notification runbook across GDPR 72h, HIPAA 60d, state laws, SEC 8-K and PCI, and compliance in CI/CD. Includes a runnable hash-chain audit log, retention engine and policy gate.
- 5.GDPR & CCPA Privacy Engineering - Lawful Bases, DSARs, Erasure in Practice, Data Transfers, DPIAs & ConsentGDPR vs CCPA/CPRA in practice: controller vs processor, the six lawful bases, DSAR workflows and deadlines, erasure in backups, event logs, search indexes, warehouses and processors (with crypto-shredding), international transfers (EU-US DPF, SCCs, TIAs), DPIAs, consent and Global Privacy Control, and the 2026 CPPA regulations. Includes a runnable crypto-shredding demo and a DSAR orchestrator with legal holds.
- 2.HIPAA & PHI - Covered Entities, Business Associates, BAAs, Security Rule Safeguards, Breach Notification & De-identificationHIPAA for builders: PHI vs health data that is not PHI, covered entities vs business associates, BAAs (including cloud and no-view providers), the Privacy Rule's minimum necessary standard, Security Rule safeguards and the status of the 2025 proposal, and breach notification (60 days, the 500 thresholds, the encryption safe harbor). Also covers Safe Harbor vs Expert Determination de-identification, tracking-technology guidance and enforcement cases, with a runnable de-identifier and PHI-safe logger.
- 4.PCI DSS v4.0.1 - Cardholder Data, CDE Scoping, Segmentation, Tokenization, SAQ Types & Payment Page ScriptsPCI DSS v4.0.1: cardholder data vs sensitive authentication data, CDE and connected-to scoping, segmentation and its testing, tokenization vs encryption, and how iframes, your own JS fields or a direct post lead to SAQ A, A-EP or D (including the January 2025 SAQ A change). Covers payment page script controls 6.4.3 and 11.6.1 and the future-dated requirements that are now mandatory, with a runnable token vault and scope calculator.
- 1.Security & Data-Protection Compliance - Why Frameworks Exist, the Shared Control Set & Choosing HIPAA, SOC 2, ISO 27001, PCI DSS, GDPR or CCPAHub: why compliance frameworks exist (law vs contract vs market), the shared control set every framework asks for, and HIPAA vs SOC 2 vs ISO 27001 vs PCI DSS vs GDPR/CCPA (plus HITRUST, FedRAMP 20x, NIST CSF) compared by trigger, assessor, artifact and cadence. Includes a runnable control-mapping matrix and framework triage in Python and TypeScript, and an engineering-guidance, not legal-advice note.
- 3.SOC 2 & ISO 27001 - Trust Services Criteria, Type I vs Type II, ISMS, Annex A, Statement of Applicability & Which to ChooseSOC 2 (an AICPA attestation against the Trust Services Criteria, Type I vs Type II, observation windows, CUECs and bridge letters) vs ISO/IEC 27001:2022 (a certifiable ISMS with clauses 4 to 10, 93 Annex A controls, the Statement of Applicability and a surveillance cycle). Covers which to choose by market and how engineers produce evidence, with a runnable Type II sampling simulation and SoA builder.
System design
Capacity, trade-offs, and request paths you can defend on a whiteboard.
Geospatial Indexing
6 studies- 1.Geospatial Indexing — Cells, Trees & Proximity at ScaleRide-hail, delivery, and maps live or die on nearby search. Index lat/lng with cells (geohash, S2, H3), adaptive trees (quadtree), or MBR trees (R-tree). Always fan out neighbors and filter true distance.
- 2.Geohash — Prefix Locality, Encoding, Precision & Edge CasesGeohash turns lat/lng into a base32 string (or interleaved int) whose shared prefixes mean spatial locality — until a cell border. Encode cheaply, pick precision from radius, fan out neighbors, then haversine-filter.
- 3.Quadtrees & Space Partitioning — Adaptive Cells, Density & UpdatesQuadtrees split a 2D region into four children when a capacity threshold is hit. Depth follows density — downtown deep, desert shallow. Cheap in-process; awkward as a durable cross-service shard key for moving objects.
- 4.R-Trees & R-star — MBRs, Bulk Load & Range/KNN QueriesR-trees index rectangles via nested minimum bounding rectangles. Databases (PostGIS GiST) use them for range, intersect, and kNN. Bulk-load static POIs; do not page-split a moving fleet — keep cells for membership.
Push Notifications
6 studies- 1.Push Notifications Architecture — APNs, FCM & Event-Driven AlertsMobile push is a distributed system: token registries, APNs/FCM provider APIs, priority and collapse, rich-content extensions, abuse-safe critical alerts, and a funnel from send to engagement. This hub maps the cluster. In-app realtime and durable buses are cousins — cross-link, do not re-teach.
- 2.iOS Notification Service Extension — Rich Content, Mute Rules & Time BudgetsNSE runs briefly before iOS presents a remote notification. You get a short wall-clock budget to mutate title/body/attachments or apply mute rules. Timeout or crash fail-open: the user still sees the original payload.
- 3.FCM & Multi-Platform Fan-out — Tokens, Topics, Collapse Keys & PriorityFCM is a common fan-out plane for Android (and often a pass-through toward APNs/web). Interviews test token lifecycle, topics vs device multicast, collapse keys, priority, and how your service abstracts APNs vs FCM behind one outbound pipeline.
- 4.Critical & Time-Sensitive Alerts — Interruption Levels, Channels & Abuse ControlsCritical and time-sensitive paths bypass Focus/DND in limited, entitlement-gated ways. Interviews look for product judgment: which events deserve interrupt, how Android channels and iOS interruption levels differ, and how you stop abuse that would get the app removed or all notifications disabled.
Rate limiting
6 studies- 1.Rate Limiting: Token Bucket, Leaky Bucket & Sliding WindowFixed-window 2× burst; sliding O(1) counter; token vs leaky bucket; Redis+Lua; 429/Retry-After.
- 2.Token Bucket vs Leaky Bucket vs Sliding WindowThree classic limiters: token bucket (burst + sustained rate), leaky bucket (smooth drain), sliding window (fairer than fixed windows). Interviews want tradeoffs, not just names.
- 3.Redis + Lua Atomic Rate LimitersDistributed limiters need atomic read-modify-write. Redis + Lua (EVAL/EVALSHA) runs check+debit in one script so concurrent replicas cannot both undercount. Prefer hash tags for Cluster slot affinity; keep scripts short.
- 4.Distributed Rate Limits Across GatewaysA limit of 100 req/s per key means nothing if each of N gateway pods enforces it on its own: the real limit becomes N x 100 and changes every time the fleet autoscales. Accurate limits need one shared counter (Redis with an atomic Lua script); fast limits need a local check that costs no network hop. Production systems layer them: the edge drops obvious abuse, a local per-pod bucket rejects clear overage for free, and a global Redis bucket enforces the real quota. When Redis is slow or down, degrade to the local share and alert, instead of either blocking everything or admitting everything.
APIs
HTTP semantics, retries, idempotency, GraphQL tradeoffs, and contract design.
API Design
6 studies- 1.API Design — Naming, Paths, Routing & ContractsInterviewers rarely ask you to design REST. They ask whether /getUser is wrong, how deep nesting should go, when to version, which status codes clients can trust, and how you paginate without breaking caches. This hub maps naming, nesting, methods/status, versioning, and list contracts.
- 2.Resource Naming & URI Design — Plurals, Case, IDs & CollectionsAlmost every API design interview starts with a path sketch. Interviewers check plural nouns, verbs smuggled into URLs, kebab vs snake consistency, opaque IDs, and whether filters live in the query string. We design the nouns idempotent retries operate on — without re-teaching Idempotency-Key.
- 3.Path Nesting, Routing & Actions — Depth Limits & Sub-resourcesNesting feels natural until clients need cross-org search, gateways rewrite paths, or authz must check every segment. Cap depth around 2–3 resource pairs, keep a canonical item URL, and model non-CRUD verbs as :customAction or /actions — not controller soup.
- 4.HTTP Methods & Status Codes for Resource APIsInterviewers expect CRUD mapped to methods and status codes clients can automate on. Always-200 with {success:false} is a classic fail. This lesson covers GET/POST/PUT/PATCH/DELETE, 201 Location, Prefer, If-Match, and the 4xx grid — pointing POST retries and 429 at the Idempotency and rate-limit docs.
GraphQL
6 studies- 1.GraphQL — Schema Design, N+1, Federation & When Not to Use ItGraphQL is a query language plus a typed schema so clients can ask for the fields they need on one endpoint. This hub is the decision map for schema nullability, resolver N+1, federation, and when REST or gRPC is the better tool.
- 2.Schema Design — Types, Nullability, Inputs & EvolutionA GraphQL schema is a versioned product contract. Nullability, input types, and additive evolution decide whether the next ship breaks clients or only adds a field they can ignore.
- 3.Resolvers & N+1 — DataLoader, Batching & Query PlanningGraphQL walks the selection set and calls a resolver per field. A child fetch per parent row is the N+1 outage. DataLoader batches those keys inside a single request, and a join or a persisted plan beats the loader when the shape is fixed.
- 4.Queries, Mutations & Subscriptions vs REST/gRPCGraphQL operations are Query for a read, Mutation for a write, and Subscription for a product event stream. Map them to REST methods and gRPC shapes, then leave CDN reads, idempotent POST retries, and mesh streams in their own clusters.
gRPC & Protobuf
6 studies- 1.gRPC & Protobuf — Contracts, Streaming & API EvolutiongRPC is HTTP/2 + Protocol Buffers + generated stubs: a contract-first RPC stack for service-to-service calls. Interviewers care less about binary speed and more about contracts, streaming shapes, evolution rules, and when gRPC is the wrong tool (browsers without grpc-web, public REST ecosystems). This hub maps the cluster: schemas, unary vs streaming, deadlines/errors, load balancing/channels, and the REST/GraphQL decision matrix.
- 2.Protobuf Schemas — Fields, Wire Types & CompatibilityProtobuf's wire identity is the field number, not the field name. Types, optional / repeated / map, proto3 defaults, reserved, and wire types determine whether two binaries can talk forever. Interviewers probe: can you add a field safely? Why never reuse numbers? What goes wrong with JSON mapping?
- 3.Unary vs Streaming RPCs — Client, Server, Bidi & BackpressuregRPC RPCs are not only request/response. Four shapes — unary, server-streaming, client-streaming, bidirectional — change memory, latency, and failure modes. Senior interviews ask when each fits, how HTTP/2 flow control provides backpressure, and which anti-patterns (unbounded server streams, ignoring cancel) blow up production.
- 4.gRPC Deadlines, Cancellation & Error ModelgRPC calls carry a deadline (absolute time), not only a per-hop timeout. Cancellation propagates so servers stop work the client no longer needs. Errors are status codes + optional details + trailers, not HTTP status alone. Interviewers compare this to HTTP codes and ask which statuses are retryable.
Idempotency
6 studies- 1.API Idempotency KeysStripe-style Idempotency-Key for safe retries — fingerprint, unique (account,key), in_progress/completed, response cache, ~24h TTL.
- 2.At-Least-Once vs Exactly-Once DeliveryAt-most-once loses; at-least-once duplicates; exactly-once needs transactional producer+consumer+sink; effectively-once is idempotent consumers + dedup.
- 3.Request Fingerprinting for Idempotent APIsCanonical JSON (sorted keys, stable numbers) hashed with method+path; include/exclude map; Stripe-style 422 on mismatch; schema-evolution with null defaults.
- 4.Retry Storms, Backoff & JitterNaive retries amplify outages. Combine exponential backoff, full jitter, Retry-After, circuit breakers, client budgets, and idempotency keys.
Caching
Latency wins, stampede control, and invalidation that does not lie.
Redis cache
8 studies- 1.Redis Cache-Aside, Invalidation & Stampede PreventionCache-aside + delete-on-write; TTL jitter; single-flight SET NX PX + Lua unlock; XFetch; negative caching.
- 2.Write-Through vs Write-Behind CachingCache-aside invalidates after DB commit. Write-through warms cache on the write path. Write-behind acks from cache/queue and flushes later — faster writes, durability risk.
- 3.Cache TTL Design & JitterTTL is a staleness bound. Identical TTLs create expiry cliffs. Add jitter; use soft TTL / XFetch on hot keys; size TTL from business tolerance, not habit.
- 4.Single-Flight Locking for Cache FillsCoalesce concurrent misses with SET NX PX. Lock TTL vs load p99. Token + Lua unlock. Waiters poll the data key. Combine with in-process singleflight.
Databases
Indexes, isolation, storage engines, shard and partition keys, and zero-downtime migrations you can ship without a maintenance window.
- 1.Database Replication — Leader-Follower, Multi-Leader & Leaderless QuorumsReplication keeps the same data on several machines for durability, read scale, or multi-region latency. The three shapes are leader-follower, multi-leader, and leaderless. Who may accept a write and when the client hears committed set lag, RPO, failover risk, and whether you must resolve conflicts.
- 2.Sync vs Async vs Semi-Sync Replication — Durability, RPO & Physical vs Logical LogsAsync, semi-sync, quorum, and fully sync commit modes trade latency for RPO. Postgres synchronous_commit levels and physical versus logical logs (plus replication slots) decide what a failover can lose and what CDC can see.
- 3.Replication Lag & Session Guarantees — Read-Your-Writes, Monotonic Reads & Consistent PrefixAsync replicas introduce read-your-writes, monotonic-read, and consistent-prefix anomalies. Fix them with LSN/GTID tokens, sticky replicas, and bounded-staleness routing — measure lag correctly first.
- 4.Failover & Split Brain — Detection, Promotion, Fencing & Lost WritesFailover is detect, elect, pick the most caught-up replica, fence the old leader, promote, repoint, and rewind. Split brain without fencing loses or duplicates writes. Patroni, managed HA, and GitHub 2018 make the trade-offs concrete.
Database storage engines
6 studies- 1.Database Storage Engines — WAL, B-Trees & LSM TreesStorage engine = on-disk layout + recovery; B-Tree vs LSM families share a WAL durability spine; pick by write:read, latency SLO, and vacuum/compaction ops cost.
- 2.Write-Ahead Log — Durability, Checkpoints & Crash RecoveryForce log before data (ARIES-style); LSN, checkpoints, REDO/UNDO; sync vs group commit vs async; torn pages / doublewrite.
- 3.B-Tree Internals — Pages, Splits & Buffer PoolPages, fanout, leaf/internal, splits; buffer pool caches dirty pages async; WAL still required; Postgres/InnoDB.
- 4.LSM Trees — MemTable, SSTables & CompactionMemTable → immutable → flush SSTable; leveled vs size-tiered compaction; blooms; RocksDB/Cassandra model.
DB Sharding & Partitioning
6 studies- 1.Database Sharding & Partitioning — Keys, Hotspots & RebalancingVertical scaling and a single primary eventually hit CPU, IOPS, storage, or write-throughput walls. Sharding splits data across nodes for independent write capacity. Table partitioning splits one table inside an engine for prune and maintenance. This hub is the map for when each is enough, how a key avoids hotspots, why scatter-gather explodes cost, and how online resharding works.
- 2.Partition Strategies — Range, Hash, List & CompositeBefore you invent an application shard layer, master in-engine partitioning: range, hash, list, and composite. The same vocabulary maps onto Vitess, Citus, and TiDB. Pick the strategy from the access pattern. A wrong one creates hotspots, immovable partitions, or queries that never prune.
- 3.Shard Keys — Hotspots, Fan-out Queries & CardinalityThe shard key is the product decision you cannot easily undo. Evaluate cardinality, skew, stability, and whether the hot transaction stays on one shard. Low-cardinality columns, monotonic time, and celebrity keys become hotspots. Queries that omit the key become fan-out.
- 4.Cross-Shard Queries — Scatter-Gather, Aggregations & AvoidanceThe day you shard, any query without the shard key becomes a scatter-gather. Latency is the slowest shard plus the merge, and every such query multiplies load by N. Merge sums and counts, never averages of averages. Keep scatter off the OLTP hot path.
Elasticsearch & OpenSearch
6 studies- 1.Elasticsearch & OpenSearch — Inverted Indexes, Relevance & OpsInterview hub: inverted indexes, BM25, shards, ILM, and Elasticsearch versus OpenSearch versus Solr. Use a search cluster for full-text, filters, and aggregations, and keep multi-row transactions in an OLTP store.
- 2.Inverted Index, Analyzers, Tokenization & MappingsAn inverted index maps each term to a posting list. Analyzers turn raw strings into those terms, and mappings decide which analyzer and field type apply. A wrong mapping silently drops recall or explodes the cluster.
- 3.Query DSL, Relevance Scoring (TF-IDF/BM25) & Filters vs QueriesQuery DSL composes bool, match, term, and range clauses. Relevance defaults to BM25. Filters answer yes or no and cache well. Queries compute the score that ranks hits.
- 4.Sharding, Replicas, Routing & Cluster HealthAn index is split into primary shards with zero or more replicas. Writes route by a key, searches scatter and gather, and cluster health is green, yellow, or red based on whether those shards are allocated.
MVCC & isolation
6 studies- 1.MVCC, Snapshot Isolation & Write Skewxmin/xmax visibility; SI vs SSI; write skew; SERIALIZABLE + 40001 retry; FOR UPDATE; VACUUM/bloat.
- 2.Isolation Levels Deep DiveIsolation levels are the contract for what concurrent transactions may see. ANSI labels hide engine gaps: Postgres RR is SI, MySQL RR uses gap locks, and write skew sits outside the ANSI phenomena list.
- 3.SSI vs Snapshot IsolationSnapshot Isolation stops lost updates on the same row but allows write skew. Postgres SERIALIZABLE (SSI) tracks rw-conflicts and aborts one txn with SQLSTATE 40001; retry the whole transaction.
- 4.Row Locks: SELECT FOR UPDATESELECT FOR UPDATE materializes the read set under SI so concurrent writers wait instead of write-skewing. Know NOWAIT, SKIP LOCKED, lock duration, and deadlock order.
- 1.Vector Indexes — Semantic Search, ANN, and When Exact kNN WinsSemantic search turns text, images, or events into vectors and answers what is closest to the query vector. Exact kNN scans every vector: perfect recall, zero build cost, trivial filters, and cost linear in N. ANN indexes give up a little recall for sub-linear queries. This hub is the decision map across HNSW, IVF, PQ, LSH, ScaNN, and DiskANN, including when brute force is the right answer.
- 2.HNSW — Layers, Greedy Search, M and efHNSW is a stack of proximity graphs. Layer 0 holds every vector, and each higher layer keeps an exponentially smaller sample. A query greedily walks the sparse top layers, then runs a beam of width efSearch on layer 0. M, efConstruction, and efSearch are the three knobs that set memory, graph quality, and the recall versus latency dial.
- 3.IVF, PQ, and ScaNN — Partition, Quantize, Re-rankIVF partitions space with k-means and scans only the nprobe cells nearest the query. PQ compresses each vector into a few bytes of codebook ids and scores them with table lookups. Together, with an exact re-rank, that is the classic recipe from hundreds of millions to billions of vectors. ScaNN spends its quantization error where inner-product rank actually moves.
- 4.Filtered and Hybrid Search — Metadata, BM25, and Broken RecallReal queries are nearest vectors for this tenant, in stock, that this user may see, and often matching an exact error code. Filters fight ANN structures: post-filtering returns too few hits, a matching-only walk strands, and a filter-aware walk gets expensive when the predicate is selective. Hybrid search adds a BM25 leg and fuses ranks so one score scale cannot silently win.
Zero-Downtime DB Migrations
6 studies- 1.Zero-Downtime Database Migrations — Expand/Contract, Dual Write & Online DDLZero-downtime migrations change schema and data paths without taking the app offline. Expand first and contract only after every reader and writer is gone. Online DDL bounds lock time. Dual-write, an idempotent backfill, and a feature-flag rollback cover moves to a new table or store.
- 2.Expand/Contract Pattern — Additive Schema Changes Without DowntimeExpand/contract keeps old and new code running on one schema: add the new shape, migrate traffic and data while both exist, and drop the old shape only after every reader and writer is gone. A rename is a new column plus dual-write, not one RENAME in the same deploy.
- 3.Online DDL — Locks, CONCURRENTLY, gh-ost & pt-oscOnline DDL applies schema changes while the database serves traffic, with locks held only for brief metadata steps. Postgres has CREATE INDEX CONCURRENTLY and a leftover INVALID index on failure. MySQL has INSTANT, INPLACE, and COPY. Hot rewrites move to gh-ost (binlog) or pt-osc (triggers).
- 4.Dual Write / Dual Read — Migrating Data Paths SafelyWhen rows must live in a new table or store, dual-write keeps both paths updated and dual-read flips traffic. The order is write-both/read-old, shadow compare, read-new, then write-new-only. Partial failure, ordering, and idempotency dominate. An outbox makes the second write durable intent.
Networking
TCP, TLS, load balancers, and why the p99 lives in the handshake.
CDN & edge cache
6 studies- 1.CDNs, Cache Hierarchy & Origin ShieldingEdge→mid-tier→shield→origin; s-maxage/SWR; cache keys; purge races; request collapsing.
- 2.Cache Keys & VaryThe cache key is the identity of a variant. Cookie in Vary or raw query strings explode cardinality; normalize host/path and put only true content axes in the key.
- 3.SWR & stale-if-errormax-age is hard freshness; stale-while-revalidate hides revalidation; stale-if-error keeps serving through origin 5xx. Together they are a latency and availability tool, not a correctness tool.
- 4.Purge & Generation TokensInstant purge races with shields; soft purge plus generation tokens / surrogate keys make invalidation deterministic. Double-purge or bump the token when the shield can refill stale.
- 1.HTTP — Versions, TLS & Connection LifecycleHTTP is a versioned application protocol on TCP or QUIC, almost always under TLS. This hub maps HTTP/1.1, HTTP/2, and HTTP/3, plus where TLS ends and how pools, timeouts, and retries fail.
- 2.HTTP/1.1 — Keep-Alive, Pipelining & Head-of-Line BlockingHTTP/1.1 made persistent connections the default so TCP and TLS are amortized. Pipelining failed in practice because responses stay in order, which is application head-of-line blocking.
- 3.HTTP/2 — Multiplexing, Streams, HPACK & Push TradeoffsHTTP/2 is a binary framed protocol. Many streams share one TCP and TLS connection, HPACK compresses headers, and server push is a dead end in browsers. TCP loss still stalls every stream.
- 4.HTTP/3 & QUIC — UDP, Migration, 0-RTT & HOL FixesHTTP/3 runs on QUIC over UDP, with TLS 1.3 inside the transport. Loss stays on one stream, a connection can survive an IP change, and 0-RTT early data can be replayed.
Load balancing
6 studies- 1.Load Balancing — L4 vs L7, Algorithms & Health ChecksWhat an LB owns (VIP, backends, health, drain); L4 vs L7 matrix; algorithm zoo; health/drain; global LB; sticky vs stateless — comparative tradeoffs for senior interviews.
- 2.L4 vs L7 Proxies — Connection Routing, TLS Termination & Protocol AwarenessTCP/UDP passthrough vs HTTP/gRPC routing; TLS terminate vs passthrough; SNI; HTTP/2 multiplexing implications; when L4 is enough; gRPC/WebSocket affinity needs.
- 3.Balancing Algorithms — Round-Robin, Least-Conn, Maglev & Power-of-Two ChoicesRR, WRR, least-conn, least-time, Maglev, P2C + bounded load; skew/hot-key behavior; when each wins.
- 4.Health Checks, Slow Start & Connection DrainingActive vs passive health; HTTP vs TCP checks; thresholds; slow start after recover; connection draining / deregistration delay; fail-open vs fail-closed.
Private Networking
6 studies- 1.Private Networking — VPC, NAT, SSM, Tunnels & IngressPrivate networking decides who can reach what, by which path, and under whose control. This hub maps VPC addressing, NAT egress, SSM access, tunnels, and north-south ingress. Sidecar/mesh and CDN stay in their own clusters.
- 2.VPC Fundamentals — Subnets, Route Tables, Security Groups & NACLsPublic versus private is a route, not a checkbox. A public subnet sends 0.0.0.0/0 to an internet gateway. Security groups are stateful ENI allow lists. NACLs are stateless subnet filters. Plan non-overlapping CIDRs before you peer.
- 3.NAT Gateways & Egress — Private Subnets, NAT vs NAT Instance, Egress ControlA private subnet has no direct internet-gateway route. NAT lets IPv4 workloads start outbound flows. It is not a firewall and not inbound publishing. Prefer endpoints when the destination is a supported AWS API, and one NAT gateway per AZ when you pay for resilience.
- 4.Secure Access Without SSH — AWS SSM Session Manager, Bastions & AlternativesSession Manager gives a shell and supported port forwarding without inbound SSH or a public IP. The agent dials Systems Manager outbound. IAM authorizes the operator. A private subnet still needs NAT or interface endpoints, and logging is not automatic.
Sidecar & Service Mesh
6 studies- 1.Sidecar Pattern & Service Mesh — Out-of-Process Proxies, Architecture & TradeoffsA sidecar is an out-of-process data plane next to the app: the workload speaks plain HTTP/gRPC, the proxy owns identity, retries, telemetry, and traffic policy. This hub maps mechanics, proxy choices, control-plane discovery, cloud equivalents, and when not to mesh.
- 2.Sidecar Mechanics — Transparent Proxy, iptables/eBPF Capture & Container LifecycleA sidecar only works if app traffic actually hits it. Transparent capture (redirect without app config) and pod lifecycle (init, ready, drain) are where interviews get concrete. Envoy, nginx, and Linkerd-proxy all need a capture story and a start/stop story.
- 3.Data-Plane Proxies — Envoy, nginx, Linkerd-proxy & What They OwnThe data plane is the process that touches every request. Envoy, nginx, and Linkerd-proxy play the same role with different L7 depth, memory, and ops. Interviews want what lives in the proxy vs the app — not an Envoy product tutorial.
- 4.Control Plane & Discovery — xDS-style Config, Endpoints & ConvergenceSidecars are useless with stale config. The control plane discovers endpoints, compiles routes and listeners, and pushes them until the fleet converges. Envoy xDS is the best-known API shape — the ideas apply to any mesh.
Concurrency
Locks, isolation, and the bugs that only show up in production.
Atomics, Memory Ordering & Lock-Free
6 studies- 1.Atomics, Memory Ordering & Lock-Free — Races, Barriers & When Locks WinInterview map for process-local concurrency: data races versus race conditions, atomics and CAS, memory orders, lock-free structures and ABA, and when a mutex still wins.
- 2.Data Races & Happens-Before — UB, Tools & Mental ModelsA data race is conflicting access with no happens-before edge. C++ and Rust call that undefined behavior. This page is the definition, the edges, and the tools that catch it.
- 3.Atomic Operations & CAS Loops — Progress GuaranteesAtomics are indivisible read-modify-write on one word. CAS loops build lock-free algorithms and accidental spinlocks. This page is the operations, the retry shape, and the progress words.
- 4.Memory Ordering — Relaxed, Acquire/Release, SeqCstAtomicity stops a torn word. Memory order controls which other accesses may move around that word. Relaxed, acquire/release, and seq_cst are the ladder, and a publish flag is the pattern to know.
Concurrency
6 studies- 1.Mutexes, Condition Variables, Deadlocks & Happens-BeforeMutex HB; CV while-loop (Mesa); Coffman + lock ordering; atomics ≠ compound atomicity; locks vs lock-free.
- 2.Mutex vs RWLockA mutex gives one thread exclusive access. An RWLock (shared mutex) lets many readers in at once or one writer. Reach for an RWLock only when reads far outnumber writes and each read section does real work (I/O, parsing, a long scan), so readers actually overlap. For short critical sections a plain mutex is usually faster: every RWLock reader still does an atomic update on the shared reader count, and that cache line ping-pongs between cores. Default to a mutex and measure before switching.
- 3.Condition Variables: Mesa vs HoareA condition variable lets a thread sleep until a predicate on shared state becomes true, always paired with the mutex that guards that state. Under Hoare semantics, signal hands the mutex straight to the woken waiter, so the predicate is guaranteed true when it runs. Under Mesa semantics (pthreads, Java, C++, Go, Rust, Python), signal is only a hint: the signaler keeps running, the waiter re-acquires the mutex later, and by then the predicate may be false again. So always wait in a while loop, never an if.
- 4.Deadlock Prevention & AvoidanceDeadlock needs Coffman’s four — mutual exclusion, hold-and-wait, no preemption, circular wait. Production systems break circular wait with lock ordering, avoid hold-and-wait with try-lock + backoff, or redesign with channels / actors so resources are not multi-locked.
Distributed locks
6 studies- 1.Distributed Locks — Correctness, Leases & Fencing TokensDistributed locks need mutual exclusion, crash safety via leases, and fencing so stale holders cannot corrupt shared state. Prefer CAS/queues when you do not need a multi-step exclusive critical section.
- 2.Redis SET NX EX vs Redlock — Single-Instance Safety & the Kleppmann DebateAtomic SET NX EX is best-effort on one Redis; Redlock’s majority story is contested (Kleppmann). Use single Redis + fencing for soft work; etcd/ZK/DB CAS for correctness-critical paths.
- 3.ZooKeeper / etcd Lock Recipes — Ephemeral Nodes, Sessions & WatchesConsensus locks use ephemeral/lease-bound keys, sequential recipes, and watches; session timeout is the fence boundary. Prefer Curator / etcd concurrency APIs over hand-rolled nodes.
- 4.Fencing Tokens — Stopping Stale Lock Holders After PartitionA monotonic fencing token must be checked by storage; without it, TTL/lease locks are unsafe after GC pauses or partitions.
Observability
SLIs, traces, and the dashboard you would actually page on.
Observability triad
5 studies- 1.Observability Triad — Metrics, Logs & Distributed TracingMetrics, logs, and traces answer different questions on one incident; correlate with exemplars and trace_id.
- 2.SLIs, SLOs & Error BudgetsSLI is the experience ratio; SLO is the target; error budget is 1 minus SLO. Choose nines and journey vs request SLIs deliberately.
- 3.Metric Cardinality & Prometheus Label DesignA series is metric times labels. Unbounded labels explode memory — bound labels; use histograms and recording rules.
- 4.Trace Context Propagation (W3C) & Sampling StrategiesW3C traceparent and tracestate keep traces intact. Head vs tail sampling trades cost against capturing rare failures.
Resilience Patterns
6 studies- 1.Resilience Patterns — Circuit Breakers, Bulkheads & Load SheddingWhen systems overload, naive retries amplify traffic, connection pools saturate, and one sick dependency can take the whole site down. Resilience patterns—circuit breakers, bulkheads, load shedding, and deadline budgets—give you deliberate ways to stop calling sick dependencies, isolate blast radius, drop work early, and bound wait time. This hub maps the cluster and shows when to use each pattern versus the alternatives.
- 2.Circuit Breaker Mechanics — Closed, Open, Half-Open, Thresholds & ProbesA circuit breaker wraps calls to a dependency and tracks recent failures (and often slow calls). In Closed it passes traffic; after a threshold it Opens and fail-fasts; after a reset timeout it Half-Opens with a limited probe. Tuned well, it stops retry storms against a sick service; tuned poorly, it flaps or never recovers.
- 3.Bulkheads — Thread Pools, Connection Pools, Queues & Failure DomainsBulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
- 4.Load Shedding & Admission Control — Drop Early, Protect the CoreLoad shedding deliberately refuses or degrades work when you are the bottleneck—before queues explode and every request times out. Admission control at the edge/gateway drops low-priority or expensive work first, returns 503 + Retry-After, and protects critical paths (checkout over recommendations). Unlike rate limiting (fairness/quotas), shedding is survival under overload.
SLOs & tracing
6 studies- 1.SLOs, Error Budgets & Distributed TracingSLI/SLO/SLA; error budget & burn rates; multi-window burn alerts; OTel spans + W3C Trace Context; histograms not averages.
- 2.SLI Design PatternsAn SLI is good/total of something users feel. Availability, latency (threshold ratio), freshness, quality, and journey SLIs each fail differently if you count retries or averages.
- 3.Multi-Window Burn AlertsPage on burn rate times window, not error rate ≥ SLO for 10 minutes. Multi-window alerts catch fast burns and slow leaks without paging on a single failed request.
- 4.Histogram vs Average Latency SLIsAverages hide the tail. Latency SLIs are histogram threshold ratios (good = faster than T); exemplars attach traces to the buckets users actually felt.
AI / ML
Retrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
Causal ML & Double Machine Learning
6 studies- 1.Causal ML & Double Machine Learning — From Association to EffectHub: correlation is not causation. Defend ATE, ATT, and CATE first, then identification, then estimators. Predictive ML is the wrong tool for treatment decisions; this cluster maps graphs, DML, CATE forests, causal time series, and health ethics.
- 2.Causal Graphs & Identification — DAGs, Confounders & BackdoorDAGs make identification inspectable before any DML fit. Distinguish confounders, colliders, and mediators; apply backdoor and positivity; name unmeasured confounding and confounding by indication.
- 3.Double Machine Learning — Nuisance Models, Orthogonalization & Cross-FittingChernozhukov DML: residual-on-residual / orthogonal scores, flexible ML nuisances, and K-fold cross-fitting so a low-dimensional ATE stays valid. DML does not invent identification — it beats naive ML-on-treatment when the backdoor set is already right.
- 4.Heterogeneous Treatment Effects — Causal Forests & Personalized MedicineCATE vs ATE: who benefits, not only whether the average helps. Causal forests and metalearners with honest splitting; personalized treatment only under overlap, multiplicity control, and validation — without overclaiming precision medicine.
Choosing ML Algorithms
6 studies- 1.Choosing ML Algorithms — Problem Shape to Model FamilySenior interviewers do not want an algorithm encyclopedia. They want you to map problem shape to model family under constraints: tabular vs sequence, interpretability vs accuracy, latency, data size, and drift. Start with a baseline; ship the simplest model that meets the metric and the ops budget.
- 2.Decision Trees — Splits, Interpretability & When They FailA tree greedily partitions feature space into axis-aligned regions. That makes it readable — and brittle. Defend when a single tree is enough (debug, rules, small data, compliance) versus when it underfits or overfits, and why ensembles exist.
- 3.Random Forests & Bagging — Variance Reduction & Feature ImportanceA senior answer for “why random forests?” is variance reduction via averaging decorrelated trees — not “it is an ensemble.” Explain bootstrap aggregating, feature bagging, OOB error, and why impurity-based importance lies when features are correlated.
- 4.Gradient Boosted Trees — XGBoost / LightGBM / CatBoost Tradeoffs“Just use XGBoost” is not a senior answer. Explain sequential residual fitting, the learning-rate × n_estimators tradeoff, early stopping, and honest differences among XGBoost, LightGBM, and CatBoost — then defend when GBM beats RF, and when calibration or latency make RF or linear wiser.
LLM Inference & Distributed Training
6 studies- 1.LLM Inference Engines & Distributed Training — SGLang, TensorRT-LLM, ONNX & PyTorch ParallelismSenior SWE/ML interviews ask you to pick a serving stack and a parallelism strategy under latency SLOs, cost per token, hardware mix, and team maturity. Mis-picking SGLang vs TensorRT-LLM vs ONNX Runtime vs raw PyTorch — or confusing DDP with FSDP — produces symptoms that look like “the model is slow” but are architectural. This hub maps the cluster.
- 2.SGLang — RadixAttention, Continuous Batching & Structured GenerationMost LLM serving pain is the runtime, not the model. Agents, chat, RAG, and multi-turn loops resend long shared prefixes. SGLang’s bet is prefix reuse via RadixAttention, plus continuous batching and a constrained-decoding API. This lesson is the runtime layer — it does not re-teach JSON Schema or CFGs.
- 3.TensorRT-LLM — Engine Build, In-Flight Batching & QuantizationTensorRT-LLM compiles a model into a serialized TensorRT engine bound to one GPU SKU, one parallelism layout, and one quantization recipe. That buys NVIDIA latency and lock-in. This lesson covers convert → build → run, in-flight batching, FP8/INT4/INT8, multi-GPU topology, and the rebuild tax.
- 4.ONNX & ONNX Runtime — Export, Graph Optimizers & Execution ProvidersONNX is the interchange format; ONNX Runtime executes and rewrites the graph. Export cleanly, pick the right opset, and order Execution Providers — or you get numerical drift, unsupported-op failures, or a GPU that sits idle while nodes fall back to CPU. This lesson is the portable path, not an LLM-server lesson.
LLM Inference in Production
5 studies- 1.LLM Inference in Production - Quantization, Prompt Caching, Multi-LoRA Serving & Capacity PlanningHub: fleet-level serving levers beyond the existing runtime pages; which lever (quantization, prefix caching, multi-LoRA, batching/capacity) moves TTFT, TPOT, memory or $/1M; runnable lever simulation + symptom-to-lever picker; triage chart.
- 2.LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offsRTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
- 3.Prefix & Prompt Caching Across Requests - Cache Keys, Routing, TTL & Hit-Rate EconomicsKV block hashing (vLLM APC) vs radix tree (SGLang); what goes into cache keys (adapter, tokenizer, images, salts); prefix-aware routing across replicas; provider prompt caching write premiums, read discounts, TTL; runnable fleet hit-rate sim + economics.
- 4.Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & MemoryMerged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
LLM Inference Runtime
6 studies- 1.LLM Inference Runtime — Prefill, KV, Batching & ParallelismSenior interviews fail when they name a serving product before the bottleneck. This hub is the mechanism ladder under the live LLM Inference Engines cluster: prefill versus decode, KV bytes, PagedAttention, continuous batching, speculative decode, then serving tensor parallel, expert parallel, and prefill/decode disaggregation.
- 2.Prefill vs Decode & KV Cache MechanicsTime-to-first-token and tokens per second split because prefill and decode are different kernels. Prefill burns FLOPs on the prompt. Decode streams the KV cache from memory for one new token. This lesson is the byte math and the autoregressive loop those metrics sit on.
- 3.PagedAttention & Continuous BatchingPagedAttention and continuous batching are the default pair behind modern open serving. Block tables make KV virtual memory. An iteration-level scheduler admits and finishes requests without waiting for a static batch. RadixAttention is a different sharing layer and stays in the SGLang study.
- 4.Speculative DecodingWhen decode is memory-bandwidth bound, speculative decoding spends spare FLOPs to cut sequential target steps. A draft proposes several tokens and the target verifies them in one forward. Acceptance rate decides whether you win or waste the step.
LLM Post-Training
6 studies- 1.LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & EvalsHub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
- 2.Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss MaskingSFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
- 3.Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward HackingBradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
- 4.Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPODPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
Production AI Agents
6 studies- 1.Production AI Agents — Architecture, Scaling, Caching & ReliabilityA production AI agent is a policy loop: model, tools, memory, and guardrails, with durable state, cost and latency SLOs, and cache layers. It is not a chat box glued to one completion. This hub maps architecture, memory, scale-out, caching, and reliability, and only cross-links the RAG, structured-outputs, and inference hubs.
- 2.Agent Architecture — Planner/Executor, Tool Registry & Control LoopsProduction agents need an explicit control architecture: who plans, who executes, how tools are registered and validated, and where humans gate irreversible actions. Compare ReAct, plan-then-act, and a graph, and point at structured outputs for argument shape without re-teaching constrained decoding.
- 3.Agent Memory — Working Context, Long-term Store, Compaction & SessionsAgent memory is a budgeted systems problem. Working context fits in the window. Episodes and semantic facts live outside it. Compaction must keep goals and commitments and must not leak PII across sessions. Retrieval of documents is a different hub.
- 4.Scaling Agents — Concurrency, Queues, Fan-out, Rate Limits & Cost BudgetsScaling agents is a distributed-systems problem: horizontal workers, job queues, per-tenant quotas, backpressure, and idempotent tool side effects. Every parallel tool call is a budgeted fork-join. Fan-out multiplies capability and cost together.
RAG & Vector Databases
6 studies- 1.RAG & Vector Databases — Retrieval, Embeddings & GroundingRAG grounds an LLM on your documents at query time: embed, retrieve, assemble context, and generate with citations. This hub is the decision map for when retrieval beats fine-tuning or long context, and how to judge grounding.
- 2.Embeddings & Similarity — Dense Vectors, Metrics & Chunking BasicsEmbeddings map text to dense vectors so similar meaning lands nearby. The usable part is the metric, the model, and the chunk size. A strong embedder with a 4k-token mush window still retrieves mush.
- 3.Vector Indexes — HNSW, IVF & Product Quantization TradeoffsBrute-force top-k dies at millions of vectors. HNSW, IVF, and product quantization trade a measured amount of recall for latency and memory. Pick them the way you pick a B-tree versus a hash index: with a recall curve and a p99.
- 4.Hybrid Retrieval — BM25 + Vectors, Reranking & Metadata FiltersProduction RAG rarely ships vectors alone. Hybrid retrieval fuses BM25 with dense similarity, reranks a shortlist, and applies metadata filters for tenant, ACL, and time before the model ever sees a chunk.
Structured outputs
6 studies- 1.Structured Outputs / Constrained DecodingStrict JSON Schema vs JSON mode; CFG token masking; required fields + additionalProperties:false; Pydantic/Zod; refusals/truncation still break validity.
- 2.JSON Schema Strictness for Structured OutputsStrict structured outputs need a closed schema: root object, every property required, additionalProperties false, provider-supported subset, schema_version, null for absence.
- 3.Tool/Function Calling vs Structured OutputsTools execute side effects and fetch live data; structured outputs constrain a final JSON contract; agent loops mix both. Pick by side effects, latency, and security — not by habit.
- 4.Grammars & CFG Constrained DecodingCompile a CFG/GBNF to an automaton, mask illegal next tokens, keep a parse stack. Same idea as JSON Schema SO, but you can constrain SQL, arithmetic, or custom DSLs.
System One & Jev
6 studies- 1.System One Models & Jev — Fast Structured Decisions for SoftwareChat models optimize for preferred strings. Software needs typed, probabilistic decisions it can branch on without parsing prose. System One Models (TypeSafe, announced 2026-09-15) take unstructured state in and emit typed probabilistic decisions. Jev is the first public System One model — this hub maps the cluster.
- 2.Choice, Score & Noul — State, Parallel Questions & Typed AnswersSystem One usefulness lives in three primitives. Mis-picking the type is a common interview fail: using a Noul when you need ordered levels, or a free-form LLM parse when a closed Choice would do. This lesson teaches state, question IDs, criteria, parallel evaluation, and answer shapes — with sandbox mocks and no API keys.
- 3.Confidence-Gated Routing, Composite Scoring & Workflow DecompositionTyped answers alone do not make automation safe. Confidence (Choice/Score) and noul magnitude (yes/no) are the second axis: what vs whether to act. This lesson teaches stake-scaled thresholds, composite scores owned in code, and speculative fan-out — without re-teaching Structured Outputs validation loops.
- 4.Generators, yield & yield* — Composing Streams and Decision WorkflowsTwo composition styles collide in AI backends: sequential token streams via generators/yield, and one-shot parallel decision round-trips (System One). Interviews expect you to implement async generators for LLM chunks and to compose typed decision steps without confusing them with token streaming.
SQL
Query plans, joins, and the index the interviewer hopes you mention.
SQL Analytics
6 studies- 1.SQL Analytics — Window Functions, CTEs & Set-Based ThinkingSenior interviews fail when SQL is written as a loop over rows. Window functions, CTEs, and set-based thinking turn Top-N, running totals, and hierarchies into declarative plans. This hub maps the cluster. Indexes and EXPLAIN stay in the existing SQL Indexes pages.
- 2.Window Functions — PARTITION BY, ORDER BY & FramesOVER () computes an aggregate without collapsing rows. Interviews love frames because the default is surprising. Master PARTITION BY, ORDER BY, and ROWS vs RANGE before ranking and running totals.
- 3.Ranking & Top-N — ROW_NUMBER, RANK, DENSE_RANK, NTILETop-N per group is a senior interview staple. The wrong ranking function silently changes product behavior on ties. Compare ROW_NUMBER, RANK, DENSE_RANK, NTILE, and Postgres DISTINCT ON.
- 4.Running Totals, LAG/LEAD & Gaps-and-IslandsTime-series analytics in SQL is cumulative sums, period-over-period deltas, and session islands. These are window problems first. Recursive CTEs are usually overkill for a linear sequence.
SQL indexes
6 studies- 1.Indexes, Cardinality & EXPLAIN PlansB-tree vs hash; selectivity; composite/covering; Seq/Index/Bitmap; Nested/Hash/Merge; EXPLAIN ANALYZE pitfalls.
- 2.B-tree vs Hash vs GIN vs GiSTB-tree is the default for equality + range + ORDER BY. Hash is equality-only (niche). GIN shines for arrays/JSONB/full-text (inverted). GiST powers geometry, ranges, and trigram similarity. Pick the access method by operator class, not vibes.
- 3.Composite & Covering IndexesComposite indexes follow leftmost prefix rules — (a,b,c) serves a, a,b, a,b,c, not b alone. Covering indexes (INCLUDE / index-only scans) store enough columns to avoid heap fetches. Column order is equality first, then range/sort.
- 4.Selectivity, Cardinality & StatisticsSelectivity is the fraction of rows matching a predicate. Cardinality is the estimated row count after filters/joins. Planners use statistics (ANALYZE) — n_distinct, MCV, histograms — to choose Seq Scan vs Index Scan. Bad stats mean bad plans.
SQL Query Execution & Optimizer
6 studies- 1.How a SQL Query Actually Executes - Parse, Rewrite, Plan, Execute, Volcano vs Vectorized & Reading Plan TreesInterview hub: parse -> analyze -> rewrite -> plan -> execute; Volcano iterator vs vectorized vs compiled execution (runnable model counting next() calls, LIMIT pipelining); real PostgreSQL 17 run where one query shape gets index+nested-loop vs seq-scan+hash-join plans depending on the constant; reading plans top-down (control) vs bottom-up (data), inclusive vs exclusive time and q-error (runnable); PG vs MySQL vs SQLite vs DuckDB planner contrasts.
- 2.Join Algorithms - Nested Loop vs Index Nested Loop vs Hash Join vs Merge JoinNested loop vs index nested loop vs hash join vs merge join: runnable work-count model, textbook I/O cost formulas and Postgres-style batch math (runnable), real PG 17 plans forcing each algorithm (incl. Memoize) and a hash join spilling to Batches: 8 under small work_mem; when each wins table, the planned-10-got-1M nested loop failure, MySQL (no merge join, hash join 8.0.18+), SQLite automatic indexes, DuckDB range joins.
- 3.Cost-Based Optimizer & Join Ordering - Cost Constants, Dynamic Programming, Left-Deep vs Bushy, join_collapse_limit & GEQOPostgreSQL cost constants verified exactly against pg_class (Seq Scan = relpages + reltuples x cpu_tuple_cost); random_page_cost 4 vs 1.1 flipping seq vs index scan in real PG 17; runnable index-vs-seq crossover model and physical correlation; runnable Selinger DP join enumeration with left-deep vs bushy and the n!/Catalan explosion; join_collapse_limit/from_collapse_limit/GEQO with a real written-order plan; CBO vs rule-based vs Cascades vs adaptive vs learned.
- 4.Cardinality Misestimates & Plan Regressions - Correlated Columns, Extended Statistics, Generic Plans & the Slow-Overnight RunbookWhy estimates go wrong: independence assumption with correlated and anti-correlated columns (runnable q-errors and multiplicative error growth, Leis et al. VLDB 2015); real PG 17 CREATE STATISTICS dependencies/ndistinct fixing estimates; stale stats and out-of-range values; custom vs generic plans for skewed prepared statements (real plans + runnable heuristic sim), parameter sniffing contrast; auto_explain/pg_stat_statements; step-labeled slow-overnight runbook; pg_hint_plan and plan-pinning trade-offs.
Messaging
Kafka, event-driven architecture, outbox and CQRS, WebSockets, MQTT, and queues you can defend in interviews.
Event-Driven Architecture
6 studies- 1.Event-Driven Architecture — Sync vs Events, Patterns & TradeoffsEvent-driven architecture publishes facts that already happened. This hub maps sync versus events, notification versus carried state versus sourcing, and where CQRS, the outbox, ordering, and failure modes sit.
- 2.Event Types — Notification, Event-Carried State Transfer & Event SourcingNotification, event-carried state transfer, and event sourcing are different contracts. Payload richness decides coupling, races, and whether the log is the source of truth.
- 3.CQRS — Commands, Queries, Projections & ConsistencyCQRS splits commands that change state from queries shaped for screens. In an event-driven system the read side is usually a projection, and lag is an SLO rather than a stuck write.
- 4.Transactional Outbox, Inbox & Consumer IdempotencyA dual-write commits the database and publishes as two steps, so one can succeed alone. The outbox makes the intent to publish part of the local commit. The inbox makes at-least-once delivery safe.
Kafka messaging
5 studies- 1.Apache Kafka — Topics, Partitions, Brokers & Consumer GroupsAppend-only partitioned logs; consumer group assigns each partition to at most one member; durability=ISR; scale=partitions; order=within partition.
- 2.Partition Keys — Ordering Guarantees vs Parallel Throughputhash(key)%N sticky order; null keys sticky/RR; repartition breaks affinity; hot keys/sticky partitioner.
- 3.Delivery Semantics — At-Least-Once, At-Most-Once & Exactly-OnceAMO vs ALO vs EOS (idempotent producer+transactions); external side effects still need idempotency keys.
- 4.Consumer Rebalancing, Lag & BackpressureCooperative sticky vs eager; lag as offset gap; pause/resume; scale consumers vs partitions.
WebSockets & MQTT
6 studies- 1.WebSockets & MQTT — Real-Time Protocols for Senior InterviewsInterviewers ask how you push live updates without burning sockets, batteries, or ops. WebSockets give full-duplex browser streams; MQTT gives topic-based pub/sub for devices. SSE, long-polling, and Kafka-style logs fill different niches — this hub maps the cluster.
- 2.WebSocket Protocol — Handshake, Frames, Ping/Pong & Close CodesRFC 6455 upgrade is HTTP 101 plus Accept = base64(SHA1(key + GUID)). Frames carry opcodes and masking; ping/pong is the application keepalive. Close 1001 drains deploys; 1006 is an abnormal drop with no close frame.
- 3.MQTT Essentials — Topics, QoS 0/1/2, Retained Messages & SessionsMQTT is broker-centric pub/sub: hierarchical topics, hop-scoped QoS 0/1/2, retained last-value, and clean vs persistent sessions. Effective QoS is the min of publish and subscribe. QoS 2 is not Kafka exactly-once.
- 4.WebSocket vs SSE vs Long-Polling vs MQTT — When to Choose WhatChoose a realtime channel from directionality, client environment, proxies, battery, topic routing, and durability. Pick the simplest protocol that fits; bridge to Kafka only when replay is a product requirement.
Data engineering
Pipelines, sketches, approximate aggregations, object storage, and stream processing with event time, windows, and exactly-once sinks.
Approximate Aggregations
6 studies- 1.Approximate Aggregations — Sketches for Quantiles, Cardinality & Merge PipelinesExact p99 over billions of events does not fit in memory, and averaging shard percentiles is wrong. Interviewers expect mergeable sketches: fixed-size summaries you update online, serialize, and combine leaves to regions to global.
- 2.KLL Quantile Sketches — Error Bounds, k Parameter & Merge SemanticsKLL (Karnin–Lang–Liberty) is a mergeable quantile sketch with a provable rank-error bound. Interviewers want the k parameter, what ε means, and how merges preserve guarantees — not vendor trivia.
- 3.T-Digest — Centroid Compression, Tail Accuracy & Heuristic LimitsT-Digest stores mergeable centroids with a compression parameter that spends accuracy on the tails. Interviewers want how centroids work, why p99 looks good, and where the heuristic breaks versus KLL’s theorems.
- 4.Sketch Merge Pipelines — Incremental, Hierarchical & Cross-Shard AggregationSketches only pay off when the pipeline merges them correctly: incrementally on a stream, hierarchically across regions, and across shards without double-counting. Interviewers want leaf-to-region-to-global and the retry failure mode.
CDC & Debezium
5 studies- 1.Change Data Capture — WAL Tailing, Debezium & Event PipelinesDual-write splits one business fact across a database commit and a later publish. CDC reads the database change log so the commit is the event. This hub maps log versus poll, Debezium, the existing outbox lesson, exactly-once effects, and failure modes.
- 2.WAL Tailing vs Query-Based CDC — Log vs Poll TradeoffsPoll CDC reads a watermark column and misses hard deletes. Log CDC reads commit order from WAL or binlog and keeps a replication slot until the consumer confirms. Pick the log when you can operate it.
- 3.Debezium & Kafka Connect — Snapshots, Offsets, Schema History & HeartbeatsDebezium snapshots a consistent read, then streams from a stored position. Connect offsets are the resume token. Schema history decodes DDL. Heartbeats advance a quiet slot so WAL can be released. Domain events still go through the outbox lesson.
- 4.Exactly-Once CDC Pipelines — Idempotent Consumers, Keys & At-Least-Once RealityCDC capture is at-least-once. Exactly-once effects come from a stable key plus an idempotent sink: an LSN guard for projections, an inbox for side effects. Kafka transactions do not make an email or a charge exactly-once.
Object Storage
6 studies- 1.Object Storage — Buckets, Keys, Consistency & ScaleObject storage is the default durable store for lakes, backups, media, and model artifacts: a flat bucket plus key over HTTP, not a POSIX mount. This hub maps block versus file versus object, then consistency, multipart, lifecycle, and the IAM threat model. Metadata partition keys and change streams stay on the sharding and CDC hubs.
- 2.Data Model — Buckets, Objects, Keys, Versioning & MetadataThe object data model is a bucket, bytes plus system metadata, a UTF-8 key, optional user metadata, and optional versions. Interviews fail when prefixes are treated as directories, ETags and version ids are ignored, or one hot prefix throttles PUTs and listings.
- 3.Consistency — Read-after-write, Listing & Conditional WritesModern S3 is strongly consistent for new objects, overwrites, deletes, and listings in-region. Interviews still expect the pre-2020 failure mode, what If-Match buys you, and how replication, CDNs, and client caches put staleness back in front of a strong API.
- 4.Multipart Uploads, Parallelism & ThroughputLarge objects should not ride one fragile HTTP PUT. Multipart splits bytes into parts, uploads them in parallel, and publishes one object at complete. Interviews expect the state machine, part size, checksums, and abort hygiene so incomplete uploads do not bill forever.
- 1.Stream Processing — Event Time, Windows, State & Exactly-OnceStream processing is continuous computation over unbounded data: you keep results fresh as events arrive instead of recomputing everything on a schedule. The hard part is not reading from Kafka fast. It is answering four questions when data arrives late, out of order, and forever: what you compute, where in event time, when you emit, and how refinements relate. A production engine must also keep large keyed state consistent across crashes and rescaling, and get results into sinks exactly once.
- 2.Windowing — Tumbling, Hopping, Session & Global Windows with TriggersA window turns an infinite stream into finite chunks you can aggregate. The window type is a product decision, not a tuning knob: tumbling windows answer per minute, hopping windows answer over the last five minutes updated every minute, session windows answer per visit, and a global window answers ever, until you say so. Triggers decide when a window emits, and the accumulation mode decides what each emission means downstream. Getting these wrong produces numbers that look plausible and are silently wrong.
- 3.Event Time & Watermarks — Allowed Lateness, Side Outputs & Idle SourcesIn event-time processing the engine never knows a window is complete. A phone that was in airplane mode can upload yesterday's clicks right now. A watermark is the engine's declaration that it believes no more events older than W will arrive. It drives on-time window firing, timers, and state cleanup. Too aggressive and you drop or mis-count late data. Too conservative and every result is delayed. Allowed lateness keeps window state around a bit longer for corrections, a side output captures anything later than that, and idleness handling stops one silent partition from freezing the whole job.
- 4.Stateful Streaming — Keyed State, RocksDB, Checkpoints & RescalingStateful streaming means the operator remembers things between events: running counts, open windows, join buffers, dedupe sets, feature rows. That state can be terabytes. It must survive crashes consistently with the source offsets, and it must be redistributable when you change parallelism. Flink's reference design is keyed state partitioned into key groups, stored in RocksDB, snapshotted by barrier checkpoints into object storage, and restorable through savepoints. Spark and Kafka Streams solve the same problem with different trade-offs.
Workflow Orchestration
6 studies- 1.Workflow Orchestration - Scheduled DAGs vs Durable Execution (Airflow, Temporal & Friends)Concept hub: two families of orchestrators, scheduled batch DAGs (Airflow, Dagster, Prefect, Argo) vs durable execution (Temporal, Cadence, Step Functions, Durable Functions, Restate, Inngest); the shared kernel (durable state, scheduler, queue, workers, retries, heartbeats/leases, at-least-once units so idempotency matters); comparison tables; runnable task-level vs step-journal crash demo and lease/heartbeat kernel; decision chart.
- 2.DAG Fundamentals & Airflow Architecture - Topological Scheduling, the Scheduler Loop, Executors, Task States & PoolsWhy DAGs and what topological order buys (waves, critical path, cycle detection); Airflow 3.x architecture (Dag processor, Dag bundles, scheduler, API server and Task Execution API, triggerer, metadata DB); task-instance states and trigger rules; pools and concurrency limits; executors compared (Local, Celery, Kubernetes, Edge, multiple executors); HA schedulers via row locks; runnable mini scheduler and wave/mapping simulation; executor decision chart.
- 3.Writing Correct Pipelines - Data Intervals, Catchup & Backfill, Idempotent Tasks, Deferrable Sensors & AssetsLogical date and data intervals, catchup and backfill in Airflow 3 (catchup off by default, logical_date None for asset/API runs), idempotent tasks with partition overwrite (runnable sqlite demo: append vs delete+insert), DST and time zones (runnable America/Chicago demo), XCom limits, sensors vs deferrable operators vs assets and AssetWatcher, dynamic task mapping, Deadline Alerts replacing SLAs, failure modes; Dagster/Prefect comparisons; waiting decision chart.
- 4.Durable Execution & the Temporal Model - Workflows vs Activities, Event History, Replay & DeterminismWorkflows vs activities, event history and replay, determinism rules (SDK time/random, no I/O in workflow code) with a runnable replay engine showing a wall-clock non-determinism bug and the fix; timers, signals, queries and updates (runnable race demo); task queues and workers; retry policy defaults and the four activity timeouts; what exactly-once means (and does not) with idempotency keys; history and payload limits; workflow-vs-activity decision chart.
Distributed systems
Raft consensus, replication, consistent hashing, saga-style distributed transactions, two-phase commit, and conflict-free replicated data types.
Consistent hashing
6 studies- 1.Consistent Hashing: Rings, Virtual Nodes & Replica PlacementModulo remaps ~all keys on membership change; consistent hashing remaps ~K/N via a clockwise hash ring. Vnodes fix skew and fan out failures; RF walks collect distinct physical nodes (topology-aware).
- 2.Rendezvous Hashing (HRW): Highest Random WeightHRW scores hash(key, node) and picks the max. No ring to maintain; membership change remaps about 1/N; lookup is O(N) unless approximated. Weights fold into the score.
- 3.Jump Consistent Hash: Dense Buckets, Almost No MemoryJump hash maps a key onto 0..N-1 with almost no memory and ~K/N movement. Buckets must be a dense integer range — no arbitrary node ids, weights, or AZ walks.
- 4.Vnode Rebalancing & Membership: Stream ~K/N Without Split-BrainAdding a node only steals ~1/N of keys, but streaming those keys still needs throttling, versioned membership, and hinted handoff so clients and replicas do not split-brain.
CRDTs
6 studies- 1.CRDTs — Conflict-Free Types, Convergence & When Consensus WinsInterview map for conflict-free replicated data types: a join that converges, the state versus op versus delta split, and the invariants that still belong on Raft or one ledger.
- 2.Counters & Registers — G-Counter, PN-Counter, LWW & MV-RegisterG-Counter and PN-Counter merge per-replica counts with a component-wise max. Last-writer-wins drops a concurrent value. The multi-value register keeps it.
- 3.Sets & Maps — G-Set, 2P-Set, OR-Set, OR-MapG-Set only adds. 2P-Set removes forever. An observed-remove set tags each add with a dot so a later add can win. An OR-Map nests a CRDT under each key.
- 4.Sequences & Collaborative Text — RGA, LSEQ, Yjs/AutomergeSequence CRDTs give each insert a stable identity so two people typing at the same place converge. RGA, LSEQ, Yjs, and Automerge are that idea with different identifiers. A last-writer-wins string is not.
Durable Objects
6 studies- 1.Cloudflare Durable Objects - Single-Instance Actors, Edge State & When to Use ThemHub: a Durable Object is a single-threaded actor with its own SQLite, one live instance per ID worldwide; routing with getByName/idFromName/newUniqueId, stubs and RPC; DO vs KV, D1, R2, Queues, Redis, Postgres row locks, Orleans and Akka with a decision chart; runnable routing simulation; what happens if you pick the alternative.
- 2.Durable Objects Storage - SQLite vs Legacy KV, Transactions, Write Coalescing & Point-in-Time RecoveryStorage: SQLite backend vs legacy KV, SQL API, sync KV API, transactionSync, write coalescing, in-memory cache, allowUnconfirmed/noCache, PITR bookmarks with ctx.abort, limits quoted from docs; output-gate and PITR simulations; storage decision chart.
- 3.Durable Objects Concurrency - Single Thread, Input & Output Gates, blockConcurrencyWhile & the Races That RemainConcurrency: single thread plus input and output gates, where interleaving still happens (fetch, timers, other objects), blockConcurrencyWhile cost, the external-call oversell race reproduced under wrangler dev (reserved 6 of 3) and fixed with claim-first and version check.
- 4.Durable Objects Real-Time - WebSocket Hibernation, Alarms, Chat, Presence & Collaborative EditingReal-time: WebSocket Hibernation API (acceptWebSocket, tags, attachments, auto-response), real chat room with presence and history, alarms, collaborative editing via CRDT or server ordering, fan-out batching and backpressure, hibernation cost math ($20.65 vs $420.65 docs example), transport decision chart.
Raft consensus
5 studies- 1.Raft Consensus — Leader Election, Log Replication & SafetySingle leader, append-only log, commit after majority; terms/roles/heartbeats/log matching vs Multi-Paxos; used in etcd/Consul/TiKV/K8s metadata.
- 2.Quorums & Majority — Why 2f+1, Read Quorums & Stale ReadsN=2f+1 tolerates f crashes; majority intersection prevents conflicting commits; linearizable reads need ReadIndex/lease not blind follower reads.
- 3.Leader Election Deep Dive — Timeouts, Randomized Election, Split VotesHeartbeat silence → new-term election via RequestVote; exclusive votes + up-to-date log; randomized timeouts; pre-vote reduces disruption.
- 4.Log Replication & Commit Index — Matching, Conflict Resolution & SafetyAppendEntries + prevLog match; nextIndex backoff; truncate divergent suffixes; commitIndex on majority current-term matchIndex (Figure 8); safety sketch.
Sagas & Distributed Transactions
6 studies- 1.Sagas & Distributed Transactions — Orchestration, Choreography & CompensationsInterview hub on sagas versus 2PC, orchestration versus choreography, compensations, deadlines, and reconciliation for multi-service business transactions.
- 2.Two-Phase Commit vs Sagas — Why 2PC Breaks at ScaleWhy 2PC blocks at scale, how that compares with sagas, and when a shared-database ACID transaction still wins.
- 3.Orchestration vs Choreography — Central Coordinator vs Event DanceCentral coordinator versus an event dance for sagas, with decision guidance and pointers to outbox, inbox, and deadlines.
- 4.Compensating Transactions — Idempotent Undo & Semantic RollbackIdempotent semantic undo for sagas: reverse-order compensations, void versus refund, and reversing ledger entries.
Time, Clocks & Ordering
6 studies- 1.Time, Clocks & Ordering in Distributed Systems - Physical Clocks, Lamport, Vector Clocks, HLC & TrueTimeInterview hub: why no machine knows the real time; wall vs monotonic; the ladder from NTP wall clocks to Lamport, vector clocks, HLC and TrueTime with a decision flow; runnable LWW-on-skewed-clocks data loss vs Lamport vs vector clocks; comparison table incl. timestamp oracles; Cloudflare 2017 leap second, Spanner, CockroachDB, Dynamo, Snowflake/UUIDv7 anchors.
- 2.Physical Clocks - NTP/PTP, Drift & Skew, Wall vs Monotonic Time & Leap SecondsHow clocks are kept in sync: oscillator drift (ppm), NTP four-timestamp offset/delay math and the delay/2 error bound (runnable), slew vs step, NTP vs PTP vs cloud time (ClockBound); wall vs monotonic APIs per language; runnable lease bug under an NTP step; leap seconds step vs smear; failure catalog (LWW, leases, TTL/JWT, negative durations).
- 3.Lamport Clocks - Happens-Before, Logical Timestamps & Total OrderHappens-before precisely; Lamport clock rules and total order with (ts, pid) tie-break (runnable); runnable proof that the clock condition holds but L(a)<L(b) does not imply causality; Lamport vs wall vs vector vs HLC vs consensus log index; Raft terms and fencing tokens as logical clocks; pitfalls (tie-break, persistence).
- 4.Vector Clocks vs Version Vectors - Detecting Concurrent Writes, Siblings & Dotted Version VectorsVector clock rules and four-way compare; runnable Dynamo-style sibling store with context and merge; vector clocks vs version vectors vs client-id vclocks vs dotted version vectors; runnable LWW vs per-server VV (lost write) vs DVV (siblings); size growth and pruning; LWW vs siblings vs CRDTs vs consensus; Dynamo, Riak 2.0, Cassandra.
- 1.Two-Phase Commit — Protocol, Coordinator & ParticipantsInterview hub: 2PC coordinator, votes, forced logs, blocking vs 3PC, Paxos Commit, and sagas.
- 2.Prepare & Commit - Votes, Logging & the DecisionPrepare and commit votes, force-before-send logging, and the decision record.
- 3.Failures & Recovery - Blocking, Coordinator Crash & Participant CrashCoordinator and participant crashes, blocking, termination rules, and heuristic decisions.
- 4.Optimizations - Presumed Abort, Presumed Commit & Read-Only VotesPresumed abort, presumed commit, and read-only votes without shrinking the uncertainty window.
Security
AuthN, AuthZ, OAuth/OIDC, OWASP web attacks (injection, XSS, CSRF, SSRF, CORS), secrets/KMS, tokens, and service identity you can defend in interviews.
Authorization
6 studies- 1.Authorization — RBAC, ABAC, ReBAC & Policy EnginesHub decision tree for AuthZ models and PEP/PDP placement; AuthN left to OAuth & OIDC cluster.
- 2.RBAC — Roles, Permissions & Role ExplosionRoles, permissions, and how role explosion pushes teams toward ABAC/ReBAC.
- 3.ABAC — Attributes, Policies, PDP & PEPAttribute-based policies with PDP/PEP separation for env and resource attributes.
- 4.ReBAC & Zanzibar — Relationship Tuples & ConsistencyRelationship-based auth with Zanzibar-style tuples and consistency tradeoffs for sharing graphs.
OAuth & OIDC
7 studies- 1.OAuth 2.1 & OIDC — Authorization Code + PKCEOAuth 2.1 makes Authorization Code + PKCE the default for public clients and retires implicit and password grants. Authentication (who) and authorization (what) are different questions: OIDC identity is the next lesson, JWT is only a token format, and the decision matrix picks session, BFF, PKCE, client credentials, device code, or mTLS. Interviews expect the redirect sequence, S256, exact redirect URIs, and state versus nonce.
- 2.OpenID Connect — ID Tokens, UserInfo, Discovery & NonceOpenID Connect is the identity layer on OAuth 2. The client asks for scope openid, validates an ID token aimed at itself (iss, aud, exp, nonce), and may call UserInfo. Discovery publishes the endpoints. Never send the ID token to your API as a bearer.
- 3.JWT vs Opaque Tokens — Validation, JWKS & RevocationJWTs validate locally via JWKS (iss/aud/exp/kid) and scale reads; opaque tokens make revocation trivial via introspection or a store. Choosing wrong means day-long stolen JWTs or introspection bottlenecks at the edge. Pair short AT TTL with refresh rotation and never skip audience checks.
- 4.Refresh Token Rotation & Reuse DetectionRefresh token rotation issues a new RT on every refresh and invalidates the old one; reuse detection treats a replayed ancestor as theft and revokes the whole token family. This is OAuth 2.1 / BCP guidance for public clients and pairs with short-lived access tokens and BFF storage.
Secrets & KMS
6 studies- 1.Secrets & KMS — Envelope Encryption, Rotation & Blast RadiusCredentials, API keys, and data-encryption keys are high-blast-radius assets. This hub is the senior-SWE decision map for where secrets live, how KMS wraps data keys (envelope encryption), how you rotate without downtime, and how apps fetch secrets at runtime — without re-teaching OAuth/OIDC token flows or RBAC/ABAC policy engines. Focus: Vault, cloud Secrets Manager, and Kubernetes Secrets tradeoffs, the DEK/KEK/CMK hierarchy, dual-read rotation, injection paths, and leakage threat models.
- 2.Envelope Encryption — DEK, KEK & CMK HierarchyEnvelope encryption separates bulk data crypto (fast local AES with a DEK) from key protection (a KMS-held CMK or KEK wraps the DEK). This lesson is the DEK, KEK, and CMK hierarchy, the encrypt and decrypt paths, AAD, key policies, and rotation by re-wrap versus re-encrypt.
- 3.Secret Stores Compared — Vault, Cloud Secrets Manager & Kubernetes SecretsPicking a secret store is an architecture choice: dynamic leases in Vault, a managed cloud Secrets Manager, or Kubernetes-native Secrets plus CSI. This lesson compares trust boundaries, auth to the store, encryption at rest, HA, and anti-patterns — without rehashing OAuth grant types or full RBAC engines.
- 4.Credential Rotation — Dual-Read, Overlap Windows & Break-GlassRotation without downtime needs a dual-read overlap, version stages, and a rehearsed break-glass path. This lesson covers overlap windows, consumer lag, database password patterns, API key versioning, and what fails when only half the fleet has the new secret.
- 1.Security & Data-Protection Compliance - Why Frameworks Exist, the Shared Control Set & Choosing HIPAA, SOC 2, ISO 27001, PCI DSS, GDPR or CCPAHub: why compliance frameworks exist (law vs contract vs market), the shared control set every framework asks for, and HIPAA vs SOC 2 vs ISO 27001 vs PCI DSS vs GDPR/CCPA (plus HITRUST, FedRAMP 20x, NIST CSF) compared by trigger, assessor, artifact and cadence. Includes a runnable control-mapping matrix and framework triage in Python and TypeScript, and an engineering-guidance, not legal-advice note.
- 2.HIPAA & PHI - Covered Entities, Business Associates, BAAs, Security Rule Safeguards, Breach Notification & De-identificationHIPAA for builders: PHI vs health data that is not PHI, covered entities vs business associates, BAAs (including cloud and no-view providers), the Privacy Rule's minimum necessary standard, Security Rule safeguards and the status of the 2025 proposal, and breach notification (60 days, the 500 thresholds, the encryption safe harbor). Also covers Safe Harbor vs Expert Determination de-identification, tracking-technology guidance and enforcement cases, with a runnable de-identifier and PHI-safe logger.
- 3.SOC 2 & ISO 27001 - Trust Services Criteria, Type I vs Type II, ISMS, Annex A, Statement of Applicability & Which to ChooseSOC 2 (an AICPA attestation against the Trust Services Criteria, Type I vs Type II, observation windows, CUECs and bridge letters) vs ISO/IEC 27001:2022 (a certifiable ISMS with clauses 4 to 10, 93 Annex A controls, the Statement of Applicability and a surveillance cycle). Covers which to choose by market and how engineers produce evidence, with a runnable Type II sampling simulation and SoA builder.
- 4.PCI DSS v4.0.1 - Cardholder Data, CDE Scoping, Segmentation, Tokenization, SAQ Types & Payment Page ScriptsPCI DSS v4.0.1: cardholder data vs sensitive authentication data, CDE and connected-to scoping, segmentation and its testing, tokenization vs encryption, and how iframes, your own JS fields or a direct post lead to SAQ A, A-EP or D (including the January 2025 SAQ A change). Covers payment page script controls 6.4.3 and 11.6.1 and the future-dated requirements that are now mandatory, with a runnable token vault and scope calculator.
Web Application Security
6 studies- 1.Web Application Security - OWASP Top 10, Injection, XSS, CSRF & SSRFAlmost every web bug on the OWASP Top 10 is untrusted data parsed as code or as authority at a sink. Fix that sink with an API that keeps code and data apart, then add a second platform layer for the day the first control is skipped.
- 2.Injection Attacks - SQLi, Command & Template Injection, Parameterized QueriesInjection is string-built commands: SQL, a shell, a template, LDAP, or a NoSQL query. The attacker's bytes close your literal and open theirs. Bind values, pass argv, and allowlist identifiers. Escaping is the fallback you will get wrong.
- 3.Cross-Site Scripting (XSS) - Contextual Encoding, Strict CSP Nonces & Trusted TypesXSS is attacker JavaScript running in your origin. Reflected, stored, and DOM XSS are three delivery routes. Contextual encoding is the fix. A strict nonce CSP and Trusted Types are the layer that still blocks the script when someone uses an escape hatch.
- 4.CSRF - SameSite Cookies, Anti-CSRF Tokens & Fetch MetadataCSRF works because the browser attaches cookies by itself. A page on another site can submit a form to yours, and the session rides along. SameSite, a CSRF token, and Origin or Sec-Fetch-Site each prove the request came from your pages. Bearer headers set by your own script are a different trade.
DevOps
Kubernetes workloads, CI/CD pipelines, artifact digests, supply-chain controls, Git rebase, merge, and recovery, and Terraform state, modules, and safe change you can defend in interviews.
Build Systems & Monorepos
6 studies- 1.Build Systems & Monorepos - Graphs, Caching, Hermeticity & Reproducible EnvironmentsConcept hub: a build is a function over a dependency graph; the four properties (correctness, incrementality, caching, hermeticity) and how they stack; task runner vs build system vs package/environment manager; what Turborepo, Nx, Bazel and Nix share and where they differ; runnable mini build function and graph-granularity demo; decision chart.
- 2.Dependency Graphs & Incremental Builds - Task DAGs, Topological Scheduling, Content Hashing & Affected DetectionPackage vs task vs action graphs; topological waves and the critical path; mtime vs content hash vs verifying and constructive traces; early cutoff; runnable Python mini build engine (topo sort, three rebuilders, cycle error) and TS monorepo task engine with affected detection; static vs dynamic dependencies and the Build Systems a la Carte grid.
- 3.Content-Addressed Caching - Cache Keys, Remote Cache, Remote Execution & PoisoningWhat goes into a cache key; undeclared env var wrong-cache-hit demo and the strict-env fix (runnable); local vs remote cache vs remote execution (REAPI CAS + action cache); constructive vs deep traces; cache economics and the critical-path floor (runnable); poisoning table incl. untrusted writers; caching decision chart.
- 4.Hermeticity & Reproducible Builds - Sandboxes, Pinned Toolchains, the Nix Store & Bit-for-Bit OutputPinned vs hermetic vs deterministic vs reproducible; sources of non-reproducibility; SOURCE_DATE_EPOCH archive demo (runnable); the Nix model (derivations, store paths, closures, substituters, input- vs content-addressed) with a store-path demo (runnable); lockfiles vs Docker vs Bazel sandboxes vs Nix/Guix; SBOMs and provenance.
CI/CD Pipelines
6 studies- 1.CI/CD Pipelines — Stages, Artifacts, Caching & Supply ChainInterview hub on CI/CD pipeline design: stages/gates, immutable artifacts, CI speed, branching/previews, and supply-chain controls (OIDC/SBOM/signing).
- 2.Pipeline Anatomy — Stages, Gates, Environments & PromotionStages, soft/hard gates, environments, and build-once promotion anatomy with GitHub Actions sketch + promotion helpers.
- 3.Artifacts & Registries — Digests, Provenance & ImmutabilityDigests vs tags, SLSA/provenance, registry immutability, and digest-handoff patterns for staging→prod.
- 4.CI Performance — Caching, Parallelism & Flaky JobsCI caching layers, parallelism/sharding, flake economics, and metrics that actually drive feedback time.
- 1.Feature Flags — Targeting, Experimentation & Kill SwitchesFeature flags (feature toggles) decouple deploy from release: you ship dark code safely, then turn behavior on for segments, percentages, or experiments, and turn it off instantly when metrics or incidents demand it. This hub teaches release, experiment, ops, and permission toggles, sticky bucketing, targeting, A/B guardrails, progressive delivery, kill switches, and hygiene so flags do not become permanent debt.
- 2.Flag Types & Evaluation — Boolean, Multivariate, Percentage & Sticky BucketsFlag types and evaluation mechanics decide whether your rollout is trustworthy. Booleans gate on or off. Multivariate flags return named variants. Percentage rollouts need a sticky hash so the same subject stays in the same bucket. This lesson implements those sketches, contrasts hash choices, and shows why a fresh random draw breaks experiments and UX.
- 3.Targeting & Context — Attributes, Segments, Rules & PrecedenceTargeting answers who gets a variant: attributes on the evaluation context, reusable segments, and ordered rules with a clear precedence. Bad targeting causes support chaos and biased experiments. This lesson shows a minimal rules engine, how segments are composed, and how exposure differs from authorization.
- 4.Experimentation & A/B — Exposure, Metrics, Guardrails & PeekingAn experiment is a sticky multivariate flag plus exposure logging, a success metric, and guardrail metrics, run with statistical discipline. This lesson covers assignment versus exposure, sample ratio mismatch, peeking, a one-sentence view of CUPED, and when a dedicated experiment platform earns its place next to the flag.
Git
6 studies- 1.Git — Everyday Commands, Rebase vs Merge & Safe HistoryGit is a content-addressed DAG of commits plus movable refs and a staging index. This hub is the decision matrix for rebase, merge, and squash, and the map to objects, history surgery, recovery, and pull-request hygiene.
- 2.Git Objects, Refs, Index & Working Tree — Mental ModelEverything in Git is an immutable content-addressed object or a mutable ref. The index is the staging area between the working tree and the next commit. Rebase, reset, and reflog stop feeling like magic once those three trees are obvious.
- 3.Rebase vs Merge vs Squash — When, Why & TradeoffsMerge records parallel work with a join commit. Rebase replays commits onto a new base and mints new ids. Squash collapses a branch into one commit. The wrong choice on a shared branch hurts the team. The wrong choice on a private branch mostly wastes review time.
- 4.History Surgery — Amend, Interactive Rebase, Fixup & AutosquashHistory surgery cleans private commits before review. Amend fixes the tip. Interactive rebase reorders, squashes, rewords, or edits. Fixup and autosquash fold review follow-ups into the commit they belong to. Every one of these tools mints new commit ids.
- 1.Infrastructure as Code — Terraform State, Modules & Safe ChangeInfrastructure as code turns console clicks into reviewable change with a known blast radius. This hub is the Terraform decision map for remote state, modules, saved plans, drift, and policy, with a light contrast to Pulumi, CloudFormation, and Crossplane.
- 2.State, Backends, Locking & WorkspacesTerraform state maps configuration addresses to real resource IDs. This page covers remote backends, locking, encryption, and why prod usually gets its own root instead of a workspace.
- 3.Resources, Providers & the Dependency GraphProviders turn HCL into API calls, and resources are the nodes in the graph. This page covers aliases, implicit edges, lifecycle, and why for_each beats count when a set changes shape.
- 4.Modules, Composition & VersioningA Terraform module is a versioned interface, not a dump of the whole account. This page covers composition, pins, registries, and when a root should stay flat.
Kubernetes Internals
6 studies- 1.Kubernetes Internals - Declarative Reconciliation from kubectl apply to Running PodHub: Kubernetes as a database of intentions plus independent watch-driven loops; imperative vs declarative reconciliation; the full kubectl apply -> running pod path (authn/authz, admission, etcd, Deployment/ReplicaSet controllers, scheduler, kubelet, CRI/CNI/CSI); component contracts, bottlenecks and a layer-picking decision chart. Goes one layer below the existing Workloads series.
- 2.Kubernetes Control Plane - API Server Request Path, etcd, Watches, Informers, Optimistic Concurrency & CRDsControl plane: API server request pipeline (authn, API Priority and Fairness, RBAC, mutating/validating admission, schema validation), etcd limits and compaction, resourceVersion, watches, 410 Gone and relist, 409 optimistic concurrency, shared informers and workqueues, CRDs and operators vs aggregated APIs.
- 3.Scheduler, Controllers & the Kubelet - Filter and Score, Taints, Preemption, Reconcile Loops, Garbage Collection & Pod SyncScheduler, controllers and kubelet: scheduling framework (PreFilter, Filter, Score, Reserve, Permit, Bind), taints/tolerations, affinity and topology spread, priority and preemption; level-triggered reconcile loops, owner references, finalizers and garbage collection; kubelet pod sync, node-pressure eviction and a Pending/ContainerCreating triage chart.
- 4.Kubernetes Storage - Volumes, PV, PVC, StorageClass, CSI, StatefulSets & Failure ModesStorage: volume types, PV/PVC/StorageClass, static vs dynamic provisioning, Immediate vs WaitForFirstConsumer, provision -> bind -> attach -> mount via CSI, access modes (RWO/RWOP/RWX), reclaim policies, StatefulSets with volumeClaimTemplates, expansion and snapshots, zone pinning, Multi-Attach and node-loss failure timelines.
Kubernetes Workloads
6 studies- 1.Kubernetes Workloads — Deployments, Probes, Resources & Progressive DeliveryControllers reconcile desired vs actual. Deployments own ReplicaSets that own Pods. Probes gate traffic and restarts. Requests/limits set QoS and scheduling. Rollouts need PDBs and a progressive strategy. This hub maps the cluster; Ingress/Gateway stays in the Private Networking pages.
- 2.Pods, ReplicaSets & Deployments — Desired State & ControllersThe Deployment controller owns ReplicaSets that own Pods. Labels and selectors wire the tree. maxUnavailable/maxSurge and revision history define how a rollout moves. Prefer Deployments over hand-managing ReplicaSets.
- 3.Probes & Pod Lifecycle — Liveness, Readiness, Startup & PreStopStartup probes buy slow boots. Readiness gates Service traffic. Liveness restarts stuck processes — never point it at a dependent database or you restart-storm yourself. PreStop plus terminationGracePeriodSeconds is how you drain cleanly.
- 4.Requests, Limits & QoS — CPU Throttling, Memory OOM & SchedulingRequests drive scheduling; limits cap cgroups. Guaranteed / Burstable / BestEffort decide who dies under pressure. CPU over limit throttles; memory over limit OOMs. Overcommit without quotas is how noisy neighbors win.
Performance
Profiling, flame graphs, latency budgets, and load tests you can defend in interviews.
Performance Engineering
6 studies- 1.Performance Engineering — Profiling, Flame Graphs, Latency Budgets & Hot PathsMeasure→hot-path→fix→re-measure loop; latency vs throughput vs capacity; when to profile CPU vs memory vs wait/IO.
- 2.CPU & Wall-Clock Profiling — Sampling vs InstrumentationSampling vs instrumentation; on-CPU vs wall-clock; GIL/event-loop traps that make “CPU” misleading.
- 3.Flame Graphs & Stack Collapse — Reading Hot PathsWidth = sample share, not one slow call; how to read plateaus and avoid common misreads.
- 4.Memory Profiling — Allocations, Leaks & GC PressureAllocation rate vs retained heap; leaks vs caches; GC pressure on p99.
Testing
Pyramid, contracts, property-based tests, doubles, and flakes you can defend in interviews.
Testing Strategies
7 studies- 1.Testing Strategies — Pyramid, Contracts, Property-Based & FlakesChoose the cheapest test that would have caught the bug. Pyramid/Trophy allocate effort; contracts protect boundaries; property/fuzz find edge cases mocks miss; doubles need discipline; flakes destroy trust.
- 2.Test Pyramid & Testing Trophy — Unit vs Integration vs E2EUnit = fast/narrow; integration = real collaborators; E2E = journeys but slow/flaky. Trophy shifts effort toward integration. Avoid ice-cream cone (mostly E2E).
- 3.Contract Testing — Consumer-Driven Contracts vs Schema ContractsConsumer-driven (Pact) vs provider schema (OpenAPI/AsyncAPI/Protobuf) vs shared-env e2e. Contracts catch breaking API changes in CI without staging.
- 4.Property-Based & Fuzz Testing — Generators, Shrinking & InvariantsGenerators + invariants + shrinking find edge cases example tests miss. Hypothesis (Python) and fast-check (TS); fuzz parsers/protocols.
DSA & Algorithms
Interview pattern map — arrays to DP/backtracking — with Big-O, templates, and LeetCode drills.
DSA Advanced & Company Favorites
16 studies- 1.DSA Advanced & Company Favorites - Tries, Greedy, Strings, Design & MoreAdvanced company-favorite DSA patterns beyond the fundamentals track: tries, bits, greedy, strings, DP advanced, design, concurrency, and mock drills.
- 2.Tries / Prefix TreesPrefix trees for autocomplete, Word Search II, and dictionary prefix queries with O(L) ops.
- 3.Bit Manipulation & BitmasksBit tricks, XOR family, subset bitmasks, and when FAANG asks bit puzzles.
- 4.Greedy Algorithms - Proof Sketches & Exchange ArgumentsWhen greedy works: exchange arguments, stay-ahead proofs, and classic interview greeds.
DSA Interview Patterns
21 studies- 1.DSA for Interviews — Pattern Map, Complexity & How to DrillInterview DSA pattern map, Big-O cheat sheet, how to pick a pattern, and drill pacing for 20 core patterns.
- 2.Arrays & Two Pointers — Opposite Ends, Same Direction & PartitionOpposite-end and same-direction two pointers for sorted arrays, partitions, and O(n) pair scans.
- 3.Sliding Window — Fixed & Variable Windows for SubarraysFixed and variable sliding windows for subarray/substring constraints in O(n).
- 4.Prefix Sums & Difference Arrays — Range Queries in O(1)Prefix sums and difference arrays for range sums and range updates.
Frontend
Browser and PWA internals: service workers, the rendering pipeline, and the multi-process model you can defend in interviews.
Browser Engines & PWAs
6 studies- 1.Browser Engines & PWAs — Service Workers, Rendering & Process ModelInterview map of browser architecture and the PWA platform: process model, rendering, Service Workers, and when a PWA beats a native shell.
- 2.Service Workers — Lifecycle, Install/Activate, Fetch & ClientsA Service Worker is an origin-scoped network proxy. Register, install, activate, skipWaiting, clients.claim, fetch, and scope explain real update bugs.
- 3.PWA Cache Strategies — Cache-First, Network-First, SWR & Workbox PatternsPick cache-first, network-first, or stale-while-revalidate per resource class. Workbox encodes the patterns. You still own versioning and update UX.
- 4.Web App Manifest, Installability, Push/Background Sync & Offline UXThe manifest names the installed app. Installability is HTTPS, a worker, icons, and engagement. App shell, push, and background sync still have platform gaps, especially on iOS.
Language Internals
JavaScript, Python, Rust, and Go runtimes, plus TypeScript and Python typing: ownership, event loops, the GIL, goroutines, and packaging.
Go Language Proficiency
11 studies- 1.Go — Beginner to Advanced (from TypeScript & Python)StudyBrief hub: Go from TypeScript and Python with Rosetta HTTP, errors, and fan-out.
- 2.Toolchain, Modules & Workspaces (Go 1.27)Toolchain, modules, workspaces, gofmt, vet, and staticcheck on Go 1.27.1.
- 3.Syntax, Types, Structs & MethodsSyntax, types, structs, methods, embedding, and zero values.
- 4.Interfaces, Generics & Type SetsImplicit interfaces, generics, type sets, and generic methods in Go 1.27.
JS/TS Language Proficiency
6 studies- 1.JavaScript & TypeScript — Intermediate to Advanced ProficiencyHub: intermediate–advanced JS/TS language proficiency — event loop, closures/this/prototypes, memory/GC, TS types, modules/emit — not React/Next.
- 2.Event Loop — Call Stack, Microtasks, Macrotasks & async/awaitCall stack, microtasks vs macrotasks, async/await scheduling, Node vs browser loop quirks.
- 3.Closures, Scope, this Binding & PrototypesLexical scope, closures, this binding rules, prototypes vs class sugar.
- 4.Memory, GC, WeakRef/WeakMap & Performance PitfallsGC reachability, WeakMap/WeakRef, retention leaks, language-level perf pitfalls.
Python Language Proficiency
6 studies- 1.Python — Intermediate to Advanced ProficiencyHub: intermediate–advanced Python proficiency — data model, asyncio vs threads vs processes, typing, packaging and profiling.
- 2.Data Model — dunder methods, protocols, descriptors & slotsDunder methods, protocols, descriptors, and slots in the CPython object model.
- 3.asyncio — Event Loop, Tasks, Cancellation & Concurrency PatternsThe asyncio event loop, Tasks, cancellation hygiene, and concurrency patterns on one thread.
- 4.GIL, Threading vs Multiprocessing vs asyncio — When Each WinsWhat the GIL does and does not protect, and when threads, processes, or asyncio win.
Rust Language Proficiency
11 studies- 1.Rust — Beginner to Advanced (from TypeScript & Python)Hub: beginner–advanced Rust for TS/Python engineers — ownership, traits, Result, async; Rosetta ownership/Result/concurrency.
- 2.Toolchain — rustup, Cargo, Editions & crates.iorustup, Cargo, editions, crates.io, toolchain files.
- 3.Ownership, Borrowing & LifetimesOwnership, moves, borrowing, lifetimes vs GC aliasing.
- 4.Types, Structs, Enums & Pattern MatchingStructs, enums, pattern matching, Option as enum.
Engineering practices
Timed interview debugging: find the signal, reproduce, ship the smallest fix, then design one component.
- 1.Interview Debugging & Implementation — Reproduce, Trace, Fix, ShipTimed interview loop: clarify, hypothesize, evidence, fix, verify, then a feature or one-component LLD. Logs versus a debugger versus reading code cold.
- 2.Finding the Signal — CloudWatch, Traces & Correlation IDsCloudWatch Logs Insights query craft, alarms as pointers, and X-Ray or OpenTelemetry correlation. The same loop on Cloud Logging and kubectl.
- 3.Reproduce & Hypothesize — Narrow, Bisect & Local ReproTighten the window and the tenant, change one variable, and keep a hypothesis log. Git bisect is named at a high level and taught in the Git series.
- 4.Fix, Verify & Roll Back — Smallest Change & ProofThe smallest correct diff, a characterization test, proof that the same query went quiet, and a rollback with a blast radius.
Operating systems
Virtual memory, paging, the page cache, and Linux I/O models: blocking calls, epoll, io_uring, and event-loop backpressure.
Linux I/O Models & Event Loops
6 studies- 1.Linux I/O Models - Blocking, Non-Blocking, epoll, io_uring & Event LoopsEvery server spends most of its life **waiting**: for a client to send bytes, for a disk, for a downstream service. An I/O model is simply the answer to "what does a thread do while it waits?". With **blocking I/O** the thread sleeps inside `read()` and you need one thread per in-flight connection. With **non-blocking I/O + readiness multiplexing** (`select`/`poll`/`epoll`/`kqueue`) one thread asks the kernel "which of my 50,000 sockets are ready?" and only touches those. With **completion-based I/O** (`io_uring` on Linux, IOCP on Windows) you hand the kernel the whole operation and buffer and collect results later. Runtimes such as Node/libuv, Nginx, Netty, Tokio and the Go netpoller are all built from these pieces; knowing which one sits under your framework explains its scaling limits, its failure modes ("someone blocked the event loop") and its tuning knobs.
- 2.File Descriptors & Non-Blocking I/O - Syscalls, EAGAIN, Partial Writes & FramingOn Unix, a socket, pipe, file, eventfd or timerfd is a **file descriptor (fd)**: a small integer indexing a per-process table that points at a kernel object. Every byte you move crosses the user/kernel boundary through a **syscall** (`read`, `write`, `recv`, `send`, `accept`). In **blocking** mode, `read()` on an empty socket puts your thread to sleep until data arrives. With `O_NONBLOCK` set (via `fcntl` or `SOCK_NONBLOCK`), the same call returns immediately with `-1` and `errno = EAGAIN` (also spelled `EWOULDBLOCK`). Non-blocking mode alone is useless (you'd spin); it's the building block that readiness APIs like epoll sit on. Two consequences every server must handle: **partial writes** (the kernel accepted only part of your buffer) and **arbitrary read boundaries** (TCP is a byte stream, so you need framing).
- 3.select vs poll vs epoll vs kqueue - Readiness, Level vs Edge Triggering & Thundering HerdsReadiness multiplexing lets one thread wait on many fds. **`select`** (1983, BSD) passes bitmaps of fds into the kernel on every call, is capped at `FD_SETSIZE` (1024 on glibc), and both kernel and app scan all of them. **`poll`** removes the cap with an array of `pollfd`, but still copies and scans the whole set every call: O(watched). **`epoll`** (Linux 2.6) keeps the interest set inside the kernel (`epoll_ctl` once per fd) and `epoll_wait` returns only the ready ones: O(ready). **`kqueue`** (FreeBSD/macOS) is the BSD equivalent and also handles timers, signals, process and file events through one API. On top of that you choose **level-triggered** (keep telling me while data remains, the default and the safest) or **edge-triggered** (`EPOLLET`, tell me once per change, and you must drain to `EAGAIN`). At multi-thread scale you also have to handle the **thundering herd** and accept distribution.
- 4.io_uring & Zero-Copy - Submission/Completion Rings, Batching, sendfile & splice**io_uring** (Linux 5.1, 2019, by Jens Axboe) is a completion-based I/O interface. The app and kernel share two ring buffers in memory: the **submission queue (SQ)** where you write operation descriptors (SQEs: read this fd into this buffer, accept, send, fsync, open...), and the **completion queue (CQ)** where the kernel posts results (CQEs with your `user_data` tag and a result code). One `io_uring_enter()` syscall can submit hundreds of operations; with **SQPOLL** a kernel thread polls the SQ and you may need no syscall at all. Unlike epoll it works for **regular files** as well as sockets, and supports registered files and fixed buffers to cut per-op overhead. The cost: a large, fast-moving kernel attack surface, so many platforms restrict it. Next to it sit the classic **zero-copy** tools: `sendfile`, `splice`, `MSG_ZEROCOPY`, and `mmap`, which reduce copies rather than syscalls.
- 1.Virtual Memory — Paging, Swapping, and Why Caches ExistEvery cache you will ever design is a smaller, faster copy of something bigger and slower, plus a rule for what to throw away. Virtual memory is the original version of that idea, built into the CPU and the kernel. A process sees a huge private address space; the kernel and MMU map small fixed-size pages of it onto physical RAM frames, keep a hot subset resident, and push the rest to disk or never materialize it at all. Once you understand page tables, the TLB, page faults, the page cache, eviction (CLOCK and working sets), huge pages, NUMA, and overcommit, three things you touch daily stop being magic: why Redis and a database buffer pool behave the way they do on a real box, why a container gets OOM-killed when the dashboard said "plenty of memory", and why vLLM chose to manage KV cache with block tables that look exactly like page tables. This hub gives the mental model and the map; five sibling lessons go deep.
- 2.Address Spaces — Paging, Segmentation, Paged Segmentation, and the TLBAn address space is a contract: the process sees contiguous private addresses, the hardware and kernel decide where the bytes actually live. Segmentation splits memory into variable-size logical regions with a base and a limit; paging splits it into fixed-size pages mapped through a page table; paged segmentation does both. Paging won because fixed-size units kill external fragmentation and make sharing, swapping, and lazy allocation trivial, but it costs page-table memory and a translation on every access. The TLB is the cache that makes that translation nearly free, and its limited reach is why random access over large heaps is slower than the big-O suggests. This lesson builds both translators in code and ends with the interview staples: multi-level tables, TLB shootdowns, ASIDs and PCIDs, and why these ideas reappear in database page ids and LLM block tables.
- 3.Faults, Swapping, and Thrashing — Working Set and the ClockA page fault is not an error; it is how the kernel implements laziness. Minor faults map a page that is already in RAM or hand out a zeroed frame in about a microsecond. Major faults read from disk or swap and cost tens of microseconds on NVMe or milliseconds on spinning disk, which is three to five orders of magnitude slower than a DRAM access. Old systems swapped whole processes; modern ones page on demand and reclaim individual pages with an approximation of LRU called CLOCK (Linux uses active and inactive lists, and newer kernels MGLRU). When the combined working sets of running processes exceed RAM, the system spends its time faulting instead of working: thrashing. This lesson builds FIFO, LRU, CLOCK, and Belady's OPT, shows the thrashing cliff, explains copy-on-write fork (Redis BGSAVE), and maps every idea onto cache design and KV-cache preemption.
- 4.The Page Cache — OS Cache vs Application Cache vs Buffer PoolLinux caches file contents in otherwise idle RAM: the page cache. Every buffered read and write, every mmap of a file, and every database that does not bypass it goes through this cache, keyed by file and offset and evicted by kernel policy. On top of it you usually stack more caches: a database buffer pool keyed by page id, an in-process LRU keyed by object, a shared Redis keyed by business identity, a CDN keyed by URL, and in LLM serving a KV cache keyed by token prefix. Each layer exists because it knows something the layer below cannot: object boundaries, transaction visibility, invalidation events, network locality, or the fact that a KV block can be recomputed. Stacking them carelessly double-caches the same bytes, wastes RAM, and lets one layer's eviction sabotage another. This lesson explains the read and write paths, mmap versus read versus direct I/O, why Postgres and InnoDB made opposite choices, why Redis is not the page cache, and how to pick the layer for each kind of data.
High-level design
HLD interview and machine-coding scale-up: capacity, APIs, data model, caching, sharding, and the path from a single-node solution to production.
SlotWise
3 studies- 1.SlotWise Reservation Service Incident - HLD Debug & Fix PathA reservation service starts double-booking slots under load. Reproduce, trace isolation and locking choices, fix with the right transaction boundaries, and ship a regression test.
- 2.SlotWise - SQLAlchemy Transaction Fix & TestsPartial unique index on active reservations, a transactional reserve that maps IntegrityError to 409, optional SELECT FOR UPDATE on the slot row, and a regression test that spawns concurrent reservers.
- 3.SlotWise - Load Repro & Isolation Decision ChartHow to reproduce the double-book under load, which isolation level still loses without a constraint, and a decision chart for constraint versus row lock versus serializable retries.
URL Shortener
3 studies- 1.URL Shortener - Spec, Flask API, SQLite & ConceptsMachine-coding URL shortener hub: Flask+SQLite spec, steps, HTTP contracts, concepts, create/redirect/delete/stats.
- 2.URL Shortener - Flask + SQLite Full Solution, Concurrency & TestsFull Flask+SQLite URL shortener with ownership delete, TTL, atomic hit UPDATE, concurrent hit tests, HTTP status contracts.
- 3.URL Shortener Scale-up HLD - ID Generation, Caching, Sharding & AnalyticsScale-up HLD: ID generation compared, Redis cache-aside, sharding, 301 vs 302, async analytics; cross-link existing pages.
Low-level design
OOP and SOLID class design for machine-coding LLD: interfaces, invariants, in-memory state, and concurrency you can implement and test.
ArchiveSweep
3 studies- 1.ArchiveSweep (File Deduplication) LLD - Spec, Pipeline & ConceptsMachine-coding LLD: recursively scan a directory tree, skip symlinks, collapse hard links to one inode, then find byte-identical duplicate groups with a staged pipeline (size -> sample prefix -> streaming SHA-256 -> byte verify).
- 2.ArchiveSweep - Streaming SHA Solution & TestsFull runnable stdlib implementation: archive_sweep(root, workers) with scan, sample, stream hash, byte verify, ThreadPoolExecutor, and unittest coverage for duplicates, near-misses, symlinks, hard links, and empty files.
- 3.ArchiveSweep - Concurrency, Symlinks & Hardlink Edge CasesWhich stages parallelize safely, how to aggregate errors without losing groups, symlink policy, hardlink reclaim math, and how to explain the design without overselling hash equality.
Comments and Replies
3 studies- 1.Comments and Replies LLD - Spec, Design Steps & ConceptsTwo-level in-memory comments/replies LLD: spec, structure choices, independent IDs, order, snapshots, concepts.
- 2.Comments and Replies - In-Memory Solution, Ordering & TestsFull CommentService Python+TS sketch with tests for order, validation, isolation, independent IDs, concurrency, snapshots.
- 3.Comments and Replies - Concurrent IDs, Snapshots & Post IsolationConcurrent ID allocation, defensive snapshots, post isolation, lock granularity decision chart, follow-ups.
Feature Flag Service
3 studies- 1.Feature Flag Service LLD - Targeting, Sticky Buckets & Kill SwitchesProvisional machine-coding LLD for an in-memory flag service: kill switch, disabled, segment gate, then a sticky SHA-256 percentage.
- 2.Feature Flag Service - Sticky Bucket + Kill Switch Solution & TestsRunnable FlagStore with upsert, sticky SHA-256 bucketing, segment gates, kill-switch precedence, and a concurrent upsert/evaluate test.
- 3.Feature Flag Service - Evaluation Order & ConcurrencyWhy the kill switch short-circuits before the percentage, why get copies the flag, and how this round connects to progressive delivery.
LLM Gateway Rate Limiter
3 studies- 1.Rate Limiter for an LLM Gateway - LLD Spec & ConceptsMachine-coding LLD for an in-process LLM gateway limiter: a per-tenant token bucket, an exact sliding-window log, and a composed decision with remaining and retry_after_ms.
- 2.Rate Limiter - Token Bucket + Sliding Window Solution & TestsRunnable TokenBucket, SlidingWindowLog, and GatewayLimiter, with a fake clock and an 80-thread burst flood.
- 3.Rate Limiter - Concurrency, Retry-After & CompositionWhy one RLock around the composed check still counts an RPM event when the burst denies, how retry_after_ms is computed, and how this round maps onto Redis Lua.
Reliability & Disaster Recovery
RTO and RPO, backups and PITR, multi-region failover, and cells with static stability you can defend in interviews.
Disaster Recovery & Multi-Region
6 studies- 1.Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-ActiveInterview hub: HA vs DR vs backup, RTO vs RPO, the four DR strategies (backup/restore, pilot light, warm standby, active-active) with cost tiers, RTO as a phase budget, hidden single-region dependencies; replication is not backup.
- 2.RTO, RPO & DR Strategies - Backup/Restore vs Pilot Light vs Warm Standby vs Active-ActiveBusiness impact analysis to tiers with RTO/RPO targets; each DR strategy in depth; expected-yearly-cost math that picks the strategy; measuring real RPO from p99 replication lag instead of averages; capacity and decision-time pitfalls.
- 3.Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore DrillsSnapshot vs incremental vs logical dump vs PITR vs object versioning; crash- vs application-consistent; runnable PITR replay stopping before a bad DELETE; 3-2-1-1-0 and immutable copies (Object Lock); GFS retention; GitLab 2017 and OVHcloud 2021; restore drills.
- 4.Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static StabilityHost/AZ/region/cell failure domains; availability math and why correlated regional dependencies cap it; static stability and data plane vs control plane; cell-based architecture and shuffle sharding (runnable); when multi-region is worth it.