Senior SWE maps
Interview study that behaves like a roadmap, not a blog.
Full lessons with gists, mobile-first diagrams, interview Q&A, voice readout, copy-ready LLM prompts, and in-page sandboxes. Skim a topic, run the snippet, then defend the trade-off.
Topics
Chunked paths you can actually finish on a phone.
System design
18 studiesCapacity, trade-offs, and request paths you can defend on a whiteboard.
rate-limiting · sharding · queues
APIs
24 studiesHTTP semantics, retries, idempotency, GraphQL tradeoffs, and contract design.
http · rest · graphql · idempotency
Caching
8 studiesLatency wins, stampede control, and invalidation that does not lie.
redis · ttl · cache-aside
Databases
42 studiesIndexes, isolation, storage engines, shard and partition keys, and zero-downtime migrations you can ship without a maintenance window.
indexes · transactions · storage · migrations · partitioning · shard-keys
Networking
30 studiesTCP, TLS, load balancers, and why the p99 lives in the handshake.
tcp · tls · lb
Concurrency
18 studiesLocks, isolation, and the bugs that only show up in production.
locks · async · races
Observability
17 studiesSLIs, traces, and the dashboard you would actually page on.
metrics · tracing · sli
AI / ML
59 studiesRetrieval, embeddings, vector indexes, evals, and serving patterns for senior interviews.
rag · embeddings · evals · serving
SQL
18 studiesQuery plans, joins, and the index the interviewer hopes you mention.
joins · plans · window-functions
Messaging
17 studiesKafka, event-driven architecture, outbox and CQRS, WebSockets, MQTT, and queues you can defend in interviews.
kafka · eda · outbox · queues
Data engineering
29 studiesPipelines, sketches, approximate aggregations, object storage, and stream processing with event time, windows, and exactly-once sinks.
sketches · pipelines · aggregations · object-storage · stream-processing
Distributed systems
41 studiesRaft consensus, replication, consistent hashing, saga-style distributed transactions, two-phase commit, and conflict-free replicated data types.
raft · consensus · replication · sagas · two-phase-commit · crdt
Security
31 studiesAuthN, AuthZ, OAuth/OIDC, OWASP web attacks (injection, XSS, CSRF, SSRF, CORS), secrets/KMS, tokens, and service identity you can defend in interviews.
oauth · oidc · jwt · mtls · authz · secrets · kms
DevOps
42 studiesKubernetes workloads, CI/CD pipelines, artifact digests, supply-chain controls, Git rebase, merge, and recovery, and Terraform state, modules, and safe change you can defend in interviews.
kubernetes · deployments · probes · rollouts · cicd · supply-chain · git · terraform · iac
Performance
6 studiesProfiling, flame graphs, latency budgets, and load tests you can defend in interviews.
profiling · flame-graphs · latency · p99
Testing
7 studiesPyramid, contracts, property-based tests, doubles, and flakes you can defend in interviews.
test-pyramid · contracts · property-based · flakes
DSA & Algorithms
37 studiesInterview pattern map — arrays to DP/backtracking — with Big-O, templates, and LeetCode drills.
dsa · algorithms · leetcode · patterns · big-o
Frontend
6 studiesBrowser and PWA internals: service workers, the rendering pipeline, and the multi-process model you can defend in interviews.
pwa · service-workers · browser-engines · rendering
Language Internals
34 studiesJavaScript, Python, Rust, and Go runtimes, plus TypeScript and Python typing: ownership, event loops, the GIL, goroutines, and packaging.
javascript · python · rust · go
Engineering practices
6 studiesTimed interview debugging: find the signal, reproduce, ship the smallest fix, then design one component.
debugging · cloudwatch · interview
Operating systems
12 studiesVirtual memory, paging, the page cache, and Linux I/O models: blocking calls, epoll, io_uring, and event-loop backpressure.
virtual-memory · paging · tlb
High-level design
6 studiesHLD interview and machine-coding scale-up: capacity, APIs, data model, caching, sharding, and the path from a single-node solution to production.
hld · url-shortener · system-design · machine-coding
Low-level design
12 studiesOOP and SOLID class design for machine-coding LLD: interfaces, invariants, in-memory state, and concurrency you can implement and test.
lld · ood · machine-coding · concurrency
Reliability & Disaster Recovery
6 studiesRTO and RPO, backups and PITR, multi-region failover, and cells with static stability you can defend in interviews.
disaster-recovery · rto · rpo · multi-region · failover
Studies
Diagrams, Q&A, and sandboxes — not gist stubs.
- Security
SOC 2 & ISO 27001 - Trust Services Criteria, Type I vs Type II, ISMS, Annex A, Statement of Applicability & Which to Choose
Cluster · Security & Data-Protection Compliance
SOC 2 (an AICPA attestation against the Trust Services Criteria, Type I vs Type II, observation windows, CUECs and bridge letters) vs ISO/IEC 27001:2022 (a certifiable ISMS with clauses 4 to 10, 93 Annex A controls, the Statement of Applicability and a surveillance cycle). Covers which to choose by market and how engineers produce evidence, with a runnable Type II sampling simulation and SoA builder.
Open study →- security
- compliance
- data-protection
- hipaa
- phi
- soc2
- iso27001
- pci-dss
- gdpr
- ccpa
- privacy-engineering
- audit-logs
- policy-as-code
- breach-notification
- interview
- Security
Security & Data-Protection Compliance - Why Frameworks Exist, the Shared Control Set & Choosing HIPAA, SOC 2, ISO 27001, PCI DSS, GDPR or CCPA
Cluster · Security & Data-Protection Compliance
Hub: why compliance frameworks exist (law vs contract vs market), the shared control set every framework asks for, and HIPAA vs SOC 2 vs ISO 27001 vs PCI DSS vs GDPR/CCPA (plus HITRUST, FedRAMP 20x, NIST CSF) compared by trigger, assessor, artifact and cadence. Includes a runnable control-mapping matrix and framework triage in Python and TypeScript, and an engineering-guidance, not legal-advice note.
Open study →- security
- compliance
- data-protection
- hipaa
- phi
- soc2
- iso27001
- pci-dss
- gdpr
- ccpa
- privacy-engineering
- audit-logs
- policy-as-code
- breach-notification
- interview
- Security
PCI DSS v4.0.1 - Cardholder Data, CDE Scoping, Segmentation, Tokenization, SAQ Types & Payment Page Scripts
Cluster · Security & Data-Protection Compliance
PCI DSS v4.0.1: cardholder data vs sensitive authentication data, CDE and connected-to scoping, segmentation and its testing, tokenization vs encryption, and how iframes, your own JS fields or a direct post lead to SAQ A, A-EP or D (including the January 2025 SAQ A change). Covers payment page script controls 6.4.3 and 11.6.1 and the future-dated requirements that are now mandatory, with a runnable token vault and scope calculator.
Open study →- security
- compliance
- data-protection
- hipaa
- phi
- soc2
- iso27001
- pci-dss
- gdpr
- ccpa
- privacy-engineering
- audit-logs
- policy-as-code
- breach-notification
- interview
- Security
HIPAA & PHI - Covered Entities, Business Associates, BAAs, Security Rule Safeguards, Breach Notification & De-identification
Cluster · Security & Data-Protection Compliance
HIPAA for builders: PHI vs health data that is not PHI, covered entities vs business associates, BAAs (including cloud and no-view providers), the Privacy Rule's minimum necessary standard, Security Rule safeguards and the status of the 2025 proposal, and breach notification (60 days, the 500 thresholds, the encryption safe harbor). Also covers Safe Harbor vs Expert Determination de-identification, tracking-technology guidance and enforcement cases, with a runnable de-identifier and PHI-safe logger.
Open study →- security
- compliance
- data-protection
- hipaa
- phi
- soc2
- iso27001
- pci-dss
- gdpr
- ccpa
- privacy-engineering
- audit-logs
- policy-as-code
- breach-notification
- interview
- Security
GDPR & CCPA Privacy Engineering - Lawful Bases, DSARs, Erasure in Practice, Data Transfers, DPIAs & Consent
Cluster · Security & Data-Protection Compliance
GDPR vs CCPA/CPRA in practice: controller vs processor, the six lawful bases, DSAR workflows and deadlines, erasure in backups, event logs, search indexes, warehouses and processors (with crypto-shredding), international transfers (EU-US DPF, SCCs, TIAs), DPIAs, consent and Global Privacy Control, and the 2026 CPPA regulations. Includes a runnable crypto-shredding demo and a DSAR orchestrator with legal holds.
Open study →- security
- compliance
- data-protection
- hipaa
- phi
- soc2
- iso27001
- pci-dss
- gdpr
- ccpa
- privacy-engineering
- audit-logs
- policy-as-code
- breach-notification
- interview
- Security
Engineering Compliance into Systems - Data Classification, Audit Logs, Retention, JIT Access, Policy-as-Code & Continuous Evidence
Cluster · Security & Data-Protection Compliance
Building compliance into the platform: data classification tags and lineage, key ownership, tamper-evident WORM audit logs, retention with legal hold, access reviews vs JIT access, policy-as-code gates, continuous evidence (how to evaluate Vanta, Drata and Secureframe), vendor risk, a breach notification runbook across GDPR 72h, HIPAA 60d, state laws, SEC 8-K and PCI, and compliance in CI/CD. Includes a runnable hash-chain audit log, retention engine and policy gate.
Open study →- security
- compliance
- data-protection
- hipaa
- phi
- soc2
- iso27001
- pci-dss
- gdpr
- ccpa
- privacy-engineering
- audit-logs
- policy-as-code
- breach-notification
- interview
- Data engineering
Workflow Orchestration - Scheduled DAGs vs Durable Execution (Airflow, Temporal & Friends)
Cluster · Workflow Orchestration
Concept hub: two families of orchestrators, scheduled batch DAGs (Airflow, Dagster, Prefect, Argo) vs durable execution (Temporal, Cadence, Step Functions, Durable Functions, Restate, Inngest); the shared kernel (durable state, scheduler, queue, workers, retries, heartbeats/leases, at-least-once units so idempotency matters); comparison tables; runnable task-level vs step-journal crash demo and lease/heartbeat kernel; decision chart.
Open study →- data-engineering
- workflow-orchestration
- airflow
- dags
- dagster
- prefect
- argo-workflows
- scheduling
- backfill
- idempotency
- interview
- Distributed systems
Vector Clocks vs Version Vectors - Detecting Concurrent Writes, Siblings & Dotted Version Vectors
Cluster · Time, Clocks & Ordering
Vector clock rules and four-way compare; runnable Dynamo-style sibling store with context and merge; vector clocks vs version vectors vs client-id vclocks vs dotted version vectors; runnable LWW vs per-server VV (lost write) vs DVV (siblings); size growth and pruning; LWW vs siblings vs CRDTs vs consensus; Dynamo, Riak 2.0, Cassandra.
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Time, Clocks & Ordering in Distributed Systems - Physical Clocks, Lamport, Vector Clocks, HLC & TrueTime
Cluster · Time, Clocks & Ordering
Interview hub: why no machine knows the real time; wall vs monotonic; the ladder from NTP wall clocks to Lamport, vector clocks, HLC and TrueTime with a decision flow; runnable LWW-on-skewed-clocks data loss vs Lamport vs vector clocks; comparison table incl. timestamp oracles; Cloudflare 2017 leap second, Spanner, CockroachDB, Dynamo, Snowflake/UUIDv7 anchors.
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Data engineering
Temporal Patterns in Production - Sagas, Child Workflows, Continue-As-New, Versioning & Worker Scaling
Cluster · Workflow Orchestration
Sagas with compensations (register compensation first, non-retryable errors, runnable trip saga), child workflows and parent close policy, continue-as-new and history limits (runnable), versioning with patching vs Worker Versioning (runnable patched-marker and unsafe-change demo), human-in-the-loop with signals/updates and timers, idempotent activities, worker scaling and sticky queues, persistence and visibility stores and history shards; shipping-a-change decision chart.
Open study →- data-engineering
- distributed-systems
- workflow-orchestration
- durable-execution
- temporal
- step-functions
- event-history
- replay
- sagas
- versioning
- interview
- SQL
Sorting, Aggregation, Spills & Parallel Query - External Merge, Top-N, HashAggregate vs GroupAggregate & work_mem Math
Cluster · SQL Query Execution & Optimizer
Sort strategies (quicksort, external merge, top-N heapsort, incremental sort) with a runnable external merge sort pass counter and top-N vs full sort comparisons; HashAggregate vs GroupAggregate; real PG 17 spills (external merge Disk, HashAggregate Batches/Disk Usage, hash_mem_multiplier) and a Gather/Partial Aggregate parallel plan; runnable work_mem worst-case sizing math; temp-file monitoring; MySQL/SQLite/DuckDB/warehouse contrasts.
Open study →- sql
- query-optimizer
- query-planner
- query-execution
- explain-analyze
- join-algorithms
- hash-join
- merge-join
- nested-loop
- cost-based-optimizer
- join-ordering
- cardinality-estimation
- extended-statistics
- work-mem
- parallel-query
- sargability
- postgresql
- interview
- SQL
Query Rewrites & SARGability - Unnesting, NOT IN vs NOT EXISTS, OR to UNION, Keyset Pagination & ORM N+1
Cluster · SQL Query Execution & Optimizer
Subquery unnesting to semi/anti joins vs the NOT IN NULL trap (real PG 17 plans and counts plus runnable sqlite3 repro); SARGability: functions and casts on columns vs half-open ranges, MySQL string-vs-number index rule; BitmapOr vs OR-to-UNION; OFFSET vs keyset pagination (150,020 vs 20 rows read in real PG); runnable ORM N+1 fingerprint detector; rewrite catalog table.
Open study →- sql
- query-optimizer
- query-planner
- query-execution
- explain-analyze
- join-algorithms
- hash-join
- merge-join
- nested-loop
- cost-based-optimizer
- join-ordering
- cardinality-estimation
- extended-statistics
- work-mem
- parallel-query
- sargability
- postgresql
- interview
- SQL
How a SQL Query Actually Executes - Parse, Rewrite, Plan, Execute, Volcano vs Vectorized & Reading Plan Trees
Cluster · SQL Query Execution & Optimizer
Interview hub: parse -> analyze -> rewrite -> plan -> execute; Volcano iterator vs vectorized vs compiled execution (runnable model counting next() calls, LIMIT pipelining); real PostgreSQL 17 run where one query shape gets index+nested-loop vs seq-scan+hash-join plans depending on the constant; reading plans top-down (control) vs bottom-up (data), inclusive vs exclusive time and q-error (runnable); PG vs MySQL vs SQLite vs DuckDB planner contrasts.
Open study →- sql
- query-optimizer
- query-planner
- query-execution
- explain-analyze
- join-algorithms
- hash-join
- merge-join
- nested-loop
- cost-based-optimizer
- join-ordering
- cardinality-estimation
- extended-statistics
- work-mem
- parallel-query
- sargability
- postgresql
- interview
- SQL
Join Algorithms - Nested Loop vs Index Nested Loop vs Hash Join vs Merge Join
Cluster · SQL Query Execution & Optimizer
Nested loop vs index nested loop vs hash join vs merge join: runnable work-count model, textbook I/O cost formulas and Postgres-style batch math (runnable), real PG 17 plans forcing each algorithm (incl. Memoize) and a hash join spilling to Batches: 8 under small work_mem; when each wins table, the planned-10-got-1M nested loop failure, MySQL (no merge join, hash join 8.0.18+), SQLite automatic indexes, DuckDB range joins.
Open study →- sql
- query-optimizer
- query-planner
- query-execution
- explain-analyze
- join-algorithms
- hash-join
- merge-join
- nested-loop
- cost-based-optimizer
- join-ordering
- cardinality-estimation
- extended-statistics
- work-mem
- parallel-query
- sargability
- postgresql
- interview
- AI / ML
Supervised Fine-Tuning (SFT) - Instruction Data, Chat Templates, Packing & Loss Masking
Cluster · LLM Post-Training
SFT data formats, chat templates and special tokens, assistant-only loss masking, packing with attention boundaries, truncation; data quality over volume; runnable masking + packing demos; template/masking bug failure paths.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
Reward Models & RLHF with PPO - Bradley-Terry, KL Penalty, Critics & Reward Hacking
Cluster · LLM Post-Training
Bradley-Terry reward models from pairwise preferences, PPO with KL penalty to the SFT reference, critic/GAE, four-models-in-memory cost, reward hacking and over-optimisation; runnable RM hacking demo + PPO step.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
Prefix & Prompt Caching Across Requests - Cache Keys, Routing, TTL & Hit-Rate Economics
Cluster · LLM Inference in Production
KV block hashing (vLLM APC) vs radix tree (SGLang); what goes into cache keys (adapter, tokenizer, images, salts); prefix-aware routing across replicas; provider prompt caching write premiums, read discounts, TTL; runnable fleet hit-rate sim + economics.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
Preference Optimization - DPO, IPO, ORPO, SimPO, KTO & GRPO Compared with PPO
Cluster · LLM Post-Training
DPO derivation from the RLHF objective; IPO, KTO, ORPO, SimPO variants (reference-free, unpaired data); GRPO group-relative advantages with verifiable rewards; all compared with PPO by signal, cost, failure mode; runnable losses + GRPO advantages.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- Distributed systems
Physical Clocks - NTP/PTP, Drift & Skew, Wall vs Monotonic Time & Leap Seconds
Cluster · Time, Clocks & Ordering
How clocks are kept in sync: oscillator drift (ppm), NTP four-timestamp offset/delay math and the delay/2 error bound (runnable), slew vs step, NTP vs PTP vs cloud time (ClockBound); wall vs monotonic APIs per language; runnable lease bug under an NTP step; leap seconds step vs smear; failure catalog (LWW, leases, TTL/JWT, negative durations).
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- AI / ML
Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & Memory
Cluster · LLM Inference in Production
Merged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- DevOps
Monorepo vs Polyrepo - Workspaces, Package Boundaries, Versioning & CI Fan-out
Cluster · Build Systems & Monorepos
Monorepo vs polyrepo trade-offs; workspaces as the base layer and task runners vs build systems on top; boundaries (tags, visibility, exports, CODEOWNERS) with a runnable boundary + CI fan-out + owners demo; fixed vs independent versioning with changesets (runnable); CI fan-out by scale from Turborepo-size repos to Google-scale.
Open study →- devops
- build-systems
- monorepo
- turborepo
- nx
- bazel
- buck2
- nix
- caching
- remote-cache
- remote-execution
- hermetic-builds
- reproducible-builds
- lockfiles
- dependency-resolution
- devcontainers
- interview
- AI / ML
LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging
Cluster · LLM Post-Training
LoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offs
Cluster · LLM Inference in Production
RTN and per-tensor/channel/group scales, outliers, WxAy notation; GPTQ vs AWQ (W4A16), SmoothQuant (W8A8), FP8 E4M3; calibration, accuracy checks, when it helps TTFT vs only TPOT; runnable quant basics, GPTQ/AWQ toy, format math.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview
- AI / ML
LLM Post-Training - From Base Model to Assistant: SFT, Preference Tuning, PEFT, Distillation & Evals
Cluster · LLM Post-Training
Hub: base model -> assistant pipeline (SFT, RLHF/PPO, DPO-family and GRPO, LoRA/QLoRA, distillation, eval gate); techniques compared by signal needed, cost and failure mode; runnable pipeline toy + technique picker; decision chart.
Open study →- ai-ml
- llm
- post-training
- sft
- instruction-tuning
- rlhf
- ppo
- reward-models
- dpo
- grpo
- orpo
- lora
- qlora
- peft
- distillation
- evals
- llm-as-judge
- interview
- AI / ML
LLM Latency, Cost & Capacity Planning - TTFT, TPOT/ITL, Goodput, $/1M Tokens & GPU Sizing
Cluster · LLM Inference in Production
TTFT vs TPOT vs ITL, goodput, percentiles from histograms; worked 70B FP8 fleet sizing (KV memory, prefill share, Little's law, headroom, $/1M); Erlang C TTFT hockey stick; autoscaling signals compared; runnable capacity plan + TTFT calculator.
Open study →- ai-ml
- llm
- inference
- serving
- quantization
- gptq
- awq
- fp8
- prefix-caching
- prompt-caching
- multi-lora
- s-lora
- capacity-planning
- ttft
- goodput
- gpu-sizing
- interview