Web Application Security - OWASP Top 10, Injection, XSS, CSRF & SSRF
Almost every web bug on the OWASP Top 10 is untrusted data parsed as code or as authority at a sink. Fix that sink with an API that keeps code and data apart, then add a second platform layer for the day the first control is skipped.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the root cause shared by SQL injection and XSS?
Answer
Untrusted data is concatenated into a string that an interpreter parses (SQL engine, HTML/JS parser), so data becomes code. The fix is an API that separates code from data: bound parameters for SQL, contextual auto-escaping and safe DOM APIs for HTML.
L2
Where should validation happen versus encoding?
Answer
Validate on input against what the field should be (type, length, allowlist). Encode on output for the specific sink. Do both; neither replaces the other.
L3
Is a WAF enough to say we're protected from injection?
Answer
No. It's a compensating control. Bypasses are routine, and it can't understand your query structure. Fix the code path and keep the WAF for defense in depth and emergency virtual patches.
L4
How would you review a new endpoint for security in 10 minutes?
Answer
Walk the five threat-model questions above, check authz on every object ID, check each sink uses the safe API, check which credentials ride along automatically, and confirm logging of security-relevant failures.
L5
Which OWASP category causes the most real breaches?
Answer
Broken access control (A01) is ranked first in 2021: missing per-object checks, IDOR, and privilege escalation. Injection is still common but frameworks have reduced it.
L6
What does "defense in depth" mean concretely for XSS?
Answer
Auto-escaping templates (primary), a strict nonce-based CSP (second), `HttpOnly` cookies so stolen script can't read the session, and Trusted Types to block dangerous DOM sinks.
L7
Why is storing pre-escaped HTML in the database a bad idea?
Answer
It ties data to one output context, double-escapes later, breaks other consumers, and still leaves JS/URL contexts unprotected.
Failure modes
WAF instead of a sink fix
Textbook payloads die. Encoded and second-order payloads do not.
Escape on input
Stored markup double-encodes later and still misses JavaScript and URL contexts.
Client-only validation
curl never runs the browser check.
Internal means trusted
SSRF and a stolen service credential make the private admin route public.
Misconceptions
Sanitize all input and the OWASP list is done.
Safety is a property of the sink. The same string is safe as a name and dangerous as SQL.
A WAF means we are protected from injection.
A WAF is a speed bump and a virtual patch. The query still has to be parameterized.
CSRF left the Top 10, so cookie sessions are fine.
SameSite and framework middleware made it rarer. Cookie-authenticated unsafe methods still need a proof of intent.
Interviewer traps
Reciting the ten category names with no control.
Give the one-line meaning and the primary control, then go deep on the sink this endpoint actually has.
Promising a single sanitizer function.
Name the interpreter. SQL gets parameters. HTML gets contextual encoding. URLs get an allowlist.
Design scenario
Same prompt for every reader.
Requirements
The note can contain an apostrophe. The callback must not reach cloud metadata. A missing check must not expose every tenant.
Traffic / scale
A few hundred sends per minute, with a public signup form in front.
Latency
The security checks run before the outbound fetch and add less than one extra DNS lookup.
Consistency
Authorization is decided on the server for that invoice id, not from a hidden form field.
Availability
A rejected callback is a 400 for that request. It does not take down the worker.
Failure assumptions
- The note is stored and later rendered in HTML.
- The callback URL redirects once.
Constraints
- No shelling out.
- Session cookie is the caller credential.
Prompt
A new POST /invoices/:id/send accepts a customer id, a note, and a callback URL. It writes SQL, renders an email, and fetches the callback.
API
Which fields are data, and which field is a URL the server will fetch?
Data
How is the note stored, and how is it encoded in the email body?
Architecture
Where do the SQL bind, the HTML encoder, the authz check, and the egress check sit?
Where the fix actually lives
Prefer
Safe API at the sink, plus a second layer
Bound parameters, contextual encoding, intent checks, and allowlists keep data as data. CSP, SameSite, and egress limits contain the miss.
- The interpreter never parses attacker bytes as syntax.
- A forgotten call site still hits the platform control.
- Stored data stays raw so each consumer encodes for itself.
Alternative
Strip bad characters, or trust the WAF
The same string is a name in one field and code in another sink. Blocklists and WAFs see yesterday's payload.
- O'Brien and encoded tags break the filter or the product.
- Escaping on the way in double-encodes on the way out.
- curl skips every check the browser enforced.
Five questions for a new endpoint
A 10-minute review. The rest of this cluster is the sink-level answer for injection, XSS, CSRF, SSRF, and CORS.
- 1
Who can call it?
Anonymous, a logged-in user, an admin, or another service. - 2
What untrusted data enters?
Body, query, headers, cookies, files, and URLs the server will fetch. - 3
Which sinks does it reach?
SQL, HTML, JavaScript, a shell, an outbound HTTP client, a redirect, or a log. - 4
What authority rides along?
Session cookies, mTLS identity, or the cloud role of the host. - 5
What is the blast radius?
One row, one tenant, or the cloud account. Zero layers on that path is a finding.
Overview
Almost every web vulnerability on the OWASP Top 10 comes from one mistake: untrusted data crosses a trust boundary and gets interpreted as something other than data. A quote turns into SQL syntax (injection). A <script> tag turns into code in someone else's browser (XSS). A cookie the browser attaches automatically turns into an action the user never meant to take (CSRF). A URL turns into a request from inside your network (SSRF). A permissive header turns into another site reading your users' data (CORS misconfiguration). The senior answer is never "sanitize all input". It is: know each sink, use the API that keeps code and data separate at that sink, and stack a second layer for when the first one is forgotten.
This cluster covers the five classes interviewers and security reviews ask about most, with the fix that actually works for each and what goes wrong if you pick a weaker one.
Broken access control is the Authorization lesson. Cryptographic failures and key handling are Secrets and KMS. Identification and session failures are OAuth 2.1 and OIDC. Known-CVE dependencies and unsigned updates are CI/CD pipelines. This page stays on the sink model.
The OWASP Top 10 (2021) in one table
| # | Category | One-line meaning | Primary control | Covered here |
|---|---|---|---|---|
| A01 | Broken Access Control | User reaches data or actions they shouldn't (IDOR, missing checks) | Server-side, per-object authorization | Authorization |
| A02 | Cryptographic Failures | Weak or missing encryption, bad key handling | TLS everywhere, KMS, modern algorithms | Secrets and KMS |
| A03 | Injection (includes XSS) | Data parsed as code by an interpreter | Parameterized APIs, contextual encoding | Injection, XSS |
| A04 | Insecure Design | Missing threat model, no abuse cases | Threat modeling, secure defaults | This hub |
| A05 | Security Misconfiguration | Default creds, verbose errors, permissive CORS/headers | Hardened baselines, config-as-code | CORS and headers |
| A06 | Vulnerable Components | Known-CVE dependencies | SCA scanning, lockfiles, patch cadence | CI/CD pipelines |
| A07 | Identification & Auth Failures | Weak sessions, credential stuffing | MFA, rate limits, secure cookies | OAuth 2.1 and OIDC |
| A08 | Software & Data Integrity Failures | Unsigned updates, unsafe deserialization | Signatures, SLSA, safe parsers | CI/CD pipelines |
| A09 | Logging & Monitoring Failures | Attacks go unnoticed | Security events, alerting, audit logs | Security events and audit logs |
| A10 | Server-Side Request Forgery | Server fetches attacker-chosen URLs | Egress allowlists, IMDSv2, IP pinning | SSRF |
CSRF dropped off the Top 10 list in 2017 because frameworks and SameSite cookies made it rarer. It still shows up in interviews and in any app with cookie sessions, so it has its own page: CSRF.
Flow
- 1
1. Untrusted input: form, header, URL, file, webhook
- next2. Trust boundary: your server or the victim's browser
- 2
2. Trust boundary: your server or the victim's browser
- next3. Sink: SQL, shell, HTML, JS, URL fetch, cookie-bearing request
- 3
3. Sink: SQL, shell, HTML, JS, URL fetch, cookie-bearing request
- next4. Interpreter parses data as code or as authority
- 4
4. Interpreter parses data as code or as authority
- next5. Fix at the sink: bind, encode, allowlist, or verify intent
- 5
5. Fix at the sink: bind, encode, allowlist, or verify intent
- next6. Second layer: CSP, SameSite, egress firewall, least privilege
- 6
6. Second layer: CSP, SameSite, egress firewall, least privilege
Lesson map
Web Application Security - OWASP Top 10, Injection, XSS, CSRF & SSRF
Almost every web bug on the OWASP Top 10 is untrusted data parsed as code or as authority at a sink. Fix that sink with an API that keeps code and data apart, then add a second platform layer for the day the first control is skipped.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Untrusted input: form, header, URL, file, webhook"] b["2. Trust boundary: your server or the victim's browser"] c["3. Sink: SQL, shell, HTML, JS, URL fetch, cookie-bearing request"] d["4. Interpreter parses data as code or as authority"] a -->|1. Untrusted input: form, header, URL, file, webhook to 2. Trust boundary: your server or the victim's browser| b b -->|2. Trust boundary: your server or the victim's browser to 3. Sink: SQL, shell, HTML, JS, URL fetch, cookie-bearing request| c c -->|3. Sink: SQL, shell, HTML, JS, URL fetch, cookie-bearing request to 4. Interpreter parses data as code or as authority| d
Same payload, five sinks
The point of this demo is that one string can be harmless or dangerous depending on where it lands, so the fix belongs at the sink, not at the front door.
# Same untrusted string, five different "sinks". Safety depends on the SINK,
# not on the input: this is why "sanitize all input" is the wrong mental model.
import html, json, sqlite3, shlex
from urllib.parse import quote
payload = "x'); DROP TABLE t;-- </script><img src=x onerror=alert(1)> $(id)"
# 1. SQL sink: a bound parameter is sent as DATA, never parsed as SQL.
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE t (name TEXT)")
db.execute("INSERT INTO t (name) VALUES (?)", (payload,))
print("sql ->", db.execute("SELECT count(*) FROM t").fetchone()[0], "row stored, table intact")
# 2. HTML body sink: encode <, >, &, quotes so the browser renders text.
print("html ->", html.escape(payload)[:70], "...")
# 3. JS string sink: JSON-encode AND break '</' so </script> can't close the tag.
print("js ->", json.dumps(payload).replace("</", "<\\/")[:70], "...")
# 4. URL query sink: percent-encode the value (never the whole URL).
print("url ->", quote(payload, safe="")[:70], "...")
# 5. Shell sink: best fix is "no shell"; if forced, quote one argument.
print("shell ->", shlex.quote(payload)[:70], "...")Output:
sql -> 1 row stored, table intact
html -> x'); DROP TABLE t;-- </script><img src=x onerror=alert(1 ...
js -> "x'); DROP TABLE t;-- <\/script><img src=x onerror=alert(1)> $(id)" ...
url -> x%27%29%3B%20DROP%20TABLE%20t%3B--%20%3C%2Fscript%3E%3Cimg%20src%3Dx%2 ...
shell -> 'x'"'"'); DROP TABLE t;-- </script><img src=x onerror=alert(1)> $(id)' ...Why "validate input" alone loses
| Approach | What it does | Why it is not enough by itself | Where it does belong |
|---|---|---|---|
Blocklist filtering (strip <script>, ', ;) | Removes "bad" characters | Endless bypasses: encodings, case, <svg onload>, \u003c; breaks legit names like O'Brien | Never as the main defense |
| Input validation (type, length, allowlist pattern) | Rejects malformed data early | The same valid string can still be dangerous in another sink | First layer, especially for IDs, enums, URLs |
| Output encoding / parameterization at the sink | Keeps data as data for that interpreter | Must be correct per context; easy to bypass with raw APIs (innerHTML, raw()) | Primary control |
| Platform second layer (CSP, SameSite, egress firewall, WAF) | Limits blast radius if the primary control is missed | Can be misconfigured; WAFs are bypassable | Defense in depth |
Defense in depth as a checklist
A useful review habit is to list each risk and the independent layers that stop it. A risk with zero layers is a finding, and a risk with one layer is a candidate for a second.
// Defense in depth: score a handler's controls per OWASP risk.
// One missing layer is tolerable; a risk with ZERO layers is a finding.
type Control = { name: string; covers: string[] };
const controls: Control[] = [
{ name: "parameterized queries (ORM)", covers: ["injection"] },
{ name: "template auto-escaping", covers: ["xss"] },
{ name: "CSP with nonces", covers: ["xss"] },
{ name: "SameSite=Lax session cookie", covers: ["csrf"] },
{ name: "object-level authz check", covers: ["broken-access-control"] },
// note: nothing covers "ssrf" in this service yet
];
const risks = ["broken-access-control", "injection", "xss", "csrf", "ssrf"];
for (const risk of risks) {
const layers = controls.filter((c) => c.covers.includes(risk)).map((c) => c.name);
const verdict = layers.length === 0 ? "FINDING (no layer)" : layers.length === 1 ? "single layer" : "defense in depth";
console.log(`${risk.padEnd(22)} ${String(layers.length)} -> ${verdict}`);
}Output:
broken-access-control 1 -> single layer
injection 1 -> single layer
xss 2 -> defense in depth
csrf 1 -> single layer
ssrf 0 -> FINDING (no layer)What happens if you choose differently
- Rely on a WAF instead of fixing code: WAFs catch the textbook payloads and miss encoded or business-logic variants. Attackers test against WAFs on purpose. Use one as a speed bump and for virtual patching while a fix ships.
- Escape on input and store escaped data: you double-encode on display, corrupt data used by non-HTML consumers (APIs, CSV, mobile), and still miss contexts like JS and URLs. Store raw, encode on output.
- Trust the frontend to validate: anything the browser enforces, an attacker with
curlskips. Client validation is UX; server validation is security. - Treat internal services as trusted: SSRF and lateral movement turn "internal only" endpoints into public ones. Authenticate service-to-service calls too.
Threat modeling in five minutes (STRIDE-lite)
For any new endpoint ask:
- Who can call it? Anonymous, logged-in user, admin, another service.
- What untrusted data enters? Body, query, headers, cookies, uploaded files, URLs to fetch.
- Which sinks does that data reach? DB query, template, shell, outbound HTTP, logs, redirects.
- What authority does the request carry automatically? Cookies, mTLS identity, cloud IAM role of the host.
- What is the blast radius if the check is missing? One user's data, every tenant, the cloud account.
Pros and cons of common framework defaults
| Default | Pro | Con / trap |
|---|---|---|
| ORM query builders | Parameterized by default | Raw SQL escape hatches (.raw(), text()), ORDER BY with user strings |
| React / Vue / Angular auto-escaping | Text interpolation is safe | dangerouslySetInnerHTML, v-html, bypassSecurityTrust*, href={userUrl} |
| Django / Rails CSRF middleware | Token checks on unsafe methods | Disabled for "API" routes that still use cookie sessions |
Browser SameSite=Lax default | Blocks most cross-site POST CSRF | Same-site subdomains are not cross-site; GET with side effects still exposed |
| Cloud IMDSv2 | Blocks naive SSRF to metadata | Only if v1 is disabled and hop limit is 1 |
Pick one route you shipped recently. Write the five threat-model answers in five lines. Mark any risk whose layer count is zero. That row is the finding.
Interview Q&A
What is the root cause shared by SQL injection and XSS?
Answer
Untrusted data is concatenated into a string that an interpreter parses (SQL engine, HTML/JS parser), so data becomes code. The fix is an API that separates code from data: bound parameters for SQL, contextual auto-escaping and safe DOM APIs for HTML.
Where should validation happen versus encoding?
Answer
Validate on input against what the field should be (type, length, allowlist). Encode on output for the specific sink. Do both; neither replaces the other.
Is a WAF enough to say we're protected from injection?
Answer
No. It's a compensating control. Bypasses are routine, and it can't understand your query structure. Fix the code path and keep the WAF for defense in depth and emergency virtual patches.
How would you review a new endpoint for security in 10 minutes?
Answer
Walk the five threat-model questions above, check authz on every object ID, check each sink uses the safe API, check which credentials ride along automatically, and confirm logging of security-relevant failures.
Which OWASP category causes the most real breaches?
Answer
Broken access control (A01) is ranked first in 2021: missing per-object checks, IDOR, and privilege escalation. Injection is still common but frameworks have reduced it.
What does "defense in depth" mean concretely for XSS?
Answer
Auto-escaping templates (primary), a strict nonce-based CSP (second), HttpOnly cookies so stolen script can't read the session, and Trusted Types to block dangerous DOM sinks.
Why is storing pre-escaped HTML in the database a bad idea?
Answer
It ties data to one output context, double-escapes later, breaks other consumers, and still leaves JS/URL contexts unprotected.