Engineering Compliance into Systems - Data Classification, Audit Logs, Retention, JIT Access, Policy-as-Code & Continuous Evidence
Building compliance into the platform: data classification tags and lineage, key ownership, tamper-evident WORM audit logs, retention with legal hold, access reviews vs JIT access, policy-as-code gates, continuous evidence (how to evaluate Vanta, Drata and Secureframe), vendor risk, a breach notification runbook across GDPR 72h, HIPAA 60d, state laws, SEC 8-K and PCI, and compliance in CI/CD. Includes a runnable hash-chain audit log, retention engine and policy gate.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why put classification in tags instead of a spreadsheet?
Answer
Tags live with the resource, so policy gates, retention jobs, erasure and scoping can read them; spreadsheets go stale within a quarter.
L2
What makes an audit log tamper-evident?
Answer
Write-once storage plus a hash chain with signed checkpoints shipped elsewhere, so edits, truncation and full rewrites are all detectable.
L3
Why separate audit events from debug logs?
Answer
Different schema, storage and IAM: the people being audited must not be able to edit the evidence.
L4
What always wins over a retention schedule?
Answer
A legal hold, which suspends deletion for matching data until it is released.
L5
How does JIT access help with access reviews?
Answer
With no standing privileges, reviews cover far less, and every elevation leaves a reason, approval and session log.
L6
Why is a policy gate result evidence?
Answer
The same check that blocks a bad change records that the control ran, with commit, pipeline and failures.
L7
What does a compliance automation platform not do?
Answer
It collects evidence and runs tests but does not design controls, and the SOC 2 opinion still comes from an independent CPA firm.
Failure modes
Audit logs in the app database
The people being audited can edit them, and an attacker with app credentials can erase their traces.
Retention that skips copies
Jobs delete from the primary store but leave replicas, backups and derived datasets behind.
Bypassable policy gates
Teams skip the gate with a label nobody reviews, so the control stops operating while still looking green.
Misconceptions
A compliance platform replaces control design.
It automates screenshots of whatever controls you have, including weak ones.
Standing access with quarterly reviews is enough.
Every review becomes a long exception list, and any compromised laptop reaches restricted data in between.
Storage TTLs handle retention.
They are a cheap backstop but cannot see legal holds or per-customer contracts; an application-level engine handles those.
Interviewer traps
Writing one breach deadline for every regime.
GDPR, HIPAA, state laws, SEC and contracts run different clocks; the runbook gathers facts fast and counsel decides what is notifiable.
Letting emergency changes bypass the pipeline.
Give them a documented fast path with after-the-fact review, not a bypass with no record.
Design scenario
Same prompt for every reader.
Requirements
Every production change and every restricted-data access is attributable, and retention runs automatically with legal holds.
Failure assumptions
- An engineer's laptop is compromised.
- A bucket is created without tags.
- A customer dispute triggers a legal hold.
Constraints
- Infrastructure as code in CI with required review.
- An auditor wants full-population evidence for a 12-month window.
Prompt
Design the delivery and data platform controls for a SaaS that holds PHI and EU personal data, so SOC 2, HIPAA and GDPR evidence collects itself.
API
How do JIT requests, approvals and session logs work for production and restricted data?
Data
How do tags, WORM audit storage and the retention engine fit together, and how are holds applied?
Architecture
Which pipeline gates (review, policy, signing, pipeline-only deploy) run, and how is their evidence exported?
Overview
Engineering guidance, not legal advice. Whether a law applies to your company, and what a contract obliges you to do, is a question for counsel and your auditors. Facts were checked against official sources on 2026-10-08.
Frameworks tell you what outcomes to prove; this page is about building systems where those outcomes are properties of the platform, so compliance is cheap and evidence collects itself. Eight building blocks do most of the work: data classification with tags and lineage, encryption with clear key ownership, tamper-evident audit logs, retention schedules with legal hold, access reviews and just-in-time access, policy-as-code gates, continuous evidence (tools such as Vanta, Drata and Secureframe compared), and vendor risk management. On top sits an incident response and breach notification runbook that can satisfy HIPAA, GDPR, state laws and SEC rules on different clocks, and compliance checks wired into CI/CD. Each section compares the approaches and what goes wrong with the alternative.
Not re-taught here (see Elsewhere in the library): KMS and envelope encryption, secrets leakage and audit trails, OPA/Rego vs Cedar, Terraform blast radius, supply-chain signing and SBOMs, SSM-based access, RBAC/ABAC design, backups and DR drills, observability pipelines.
1. Data classification, tagging and lineage
Every other control depends on knowing what data is where. A workable scheme has four or five levels and maps regulated types onto them:
| Class | Examples | Default controls |
|---|---|---|
| Public | Docs site, marketing | Integrity only |
| Internal | Metrics, non-sensitive config | Auth required |
| Confidential | Customer business data, source code | Encryption, access by role, retention |
| Restricted | PHI, PAN, government IDs, credentials, special-category personal data | Isolated stores, field-level encryption or tokens, purpose-based access, audit on every read, strict retention |
Classification is useless if it lives in a spreadsheet. Put it in the systems: column-level tags in the warehouse catalog, resource tags (data_class, residency, owner) on buckets and databases, schema annotations on API fields and event schemas. Then lineage (which jobs read which tagged columns and write where) lets you answer the questions every framework asks: where does PHI flow, which systems are in PCI scope, which stores must an erasure job touch.
Compared: manual inventories go stale within a quarter; automated discovery scanners find unknown copies but produce noise; schema-level tags enforced in CI are the most reliable for new data. Use tags as the source of truth and scanners as the safety net.
Just-in-time access, or standing access with periodic reviews?
Prefer
JIT elevation with expiry and session logs
No standing production or restricted-data access; engineers request time-boxed elevation with a reason.
- Access reviews shrink to the few standing grants that remain.
- Every elevation records reason, approval, expiry and the session.
- A stolen laptop does not carry standing access to restricted data.
Alternative
Standing access, reviewed every quarter
Engineers keep broad access, and managers certify it periodically.
- Reviews turn into long lists of exceptions.
- Access is open between reviews, for insiders and stolen credentials.
- Evidence shows who could access data, not who did.
A change from commit to production, condensed
Diagram 1 condensed.
- 1
Tag data classes
Tags tell every later control what data a resource holds. - 2
Review, policy gate and signed build
Peer review, policy-as-code and provenance block noncompliant or unreviewed changes. - 3
Deploy via the pipeline with JIT access only
No direct production writes and no standing privileges. - 4
Audit to WORM and run retention
Tamper-evident logs prove who did what; retention deletes on schedule unless a hold applies. - 5
Control test fails?
Continuous tests catch drift between audits; a failure is handled as an incident or finding, and evidence is collected automatically.
2. Encryption and key ownership
Encryption at rest and in transit is table stakes; the compliance questions are about keys: who can use them, who can disable them, where they live, and whether access is logged. Patterns, from least to most customer control: provider-managed keys, customer-managed KMS keys in your account (the default for regulated data), per-tenant keys (enables tenant-level crypto-shredding and offboarding), and customer-held keys (BYOK/HYOK) for buyers who require them. Separate the duty of key administration from data access, and log every decrypt of restricted data. The mechanics are covered in envelope-encryption-dek-kek-cmk and secrets-kms-envelope-encryption-rotation.
3. Immutable, tamper-evident audit logs
Audit logs answer "who did what to which record, when, from where". Every framework asks for them (HIPAA 164.312(b), SOC 2 CC7.2, ISO A.8.15, PCI Requirement 10, which keeps at least 12 months of history with 3 months immediately available). Design rules:
- Separate audit events from application debug logs: different schema, different storage, different IAM.
- Write-once storage (object lock in compliance mode, append-only ledger tables) in an account the application team cannot administer.
- Tamper evidence: hash-chain entries and ship signed checkpoints elsewhere, so edits, truncation and wholesale rewrites are all detectable.
- Alert on gaps: a log source that goes quiet is an incident signal.
"""Tamper-evident audit log: hash chain + signed checkpoints.
Concept first:
* Auditors (SOC 2 CC7.2, HIPAA 164.312(b), PCI 10.3.2/10.3.4, ISO A.8.15) want
audit logs protected from modification. "Only admins can edit" is weak: an
admin is exactly who an attacker becomes.
* A hash chain makes each entry commit to the previous one, so editing or
deleting any entry breaks every later hash. A periodic checkpoint (head hash)
signed and shipped to a separate account or WORM storage (e.g. S3 Object Lock
in compliance mode) means rewriting the whole chain is also detectable.
* Tamper-EVIDENT is not tamper-PROOF: you still need WORM storage, separate
credentials, and alerting on gaps.
"""
import hashlib, hmac, json
CHECKPOINT_KEY = b"held-by-a-different-account" # demo; real: KMS asymmetric signing key
def entry_hash(prev: str, body: dict) -> str:
return hashlib.sha256((prev + json.dumps(body, sort_keys=True)).encode()).hexdigest()
def append(log: list, body: dict) -> None:
prev = log[-1]["h"] if log else "0" * 64
log.append({"body": body, "prev": prev, "h": entry_hash(prev, body)})
def checkpoint(log: list) -> dict:
head = log[-1]["h"]
return {"n": len(log), "head": head, "sig": hmac.new(CHECKPOINT_KEY, head.encode(), "sha256").hexdigest()}
def verify(log: list, cp: dict) -> str:
prev = "0" * 64
for i, e in enumerate(log):
if e["prev"] != prev or entry_hash(prev, e["body"]) != e["h"]:
return f"BROKEN at entry {i}"
prev = e["h"]
if len(log) < cp["n"]:
return f"TRUNCATED: {cp['n'] - len(log)} entries missing since checkpoint"
if log[cp["n"] - 1]["h"] != cp["head"] or not hmac.compare_digest(
cp["sig"], hmac.new(CHECKPOINT_KEY, cp["head"].encode(), "sha256").hexdigest()):
return "REWRITTEN: chain does not match signed checkpoint"
return "OK"
log: list = []
for ev in [("u-77", "view", "pt_9f2c"), ("u-12", "export", "pt_9f2c"), ("admin-3", "role_grant", "u-12"),
("u-77", "view", "pt_1a7d")]:
append(log, {"actor": ev[0], "action": ev[1], "target": ev[2]})
cp = checkpoint(log)
print("clean log:", verify(log, cp))
tampered = json.loads(json.dumps(log)); tampered[1]["body"]["action"] = "view" # hide the export
print("edit one field:", verify(tampered, cp))
truncated = log[:3]
print("delete the tail:", verify(truncated, cp))
rewritten: list = [] # attacker rebuilds a consistent chain without the export
for e in log:
if e["body"]["action"] != "export":
append(rewritten, e["body"])
append(rewritten, {"actor": "u-77", "action": "view", "target": "pt_0000"}) # pad to same length
print("rebuild whole chain:", verify(rewritten, cp))Output:
clean log: OK
edit one field: BROKEN at entry 1
delete the tail: TRUNCATED: 1 entries missing since checkpoint
rebuild whole chain: REWRITTEN: chain does not match signed checkpoint4. Retention schedules and legal hold
Retention is two-sided: keep data long enough (legal floors such as PCI's 12-month log history or HIPAA's 6-year retention of required documentation) and no longer (GDPR storage limitation, PCI's minimal storage, smaller breach blast radius). A legal hold suspends deletion for matching data and always wins. Design the engine to run as a dry run first, act per data class, and write deletion evidence.
"""Retention schedule + legal hold engine (dry run first).
Concept first:
* A retention schedule says how long each DATA CLASS lives and why (law,
contract, business need). Keeping data "just in case" is itself a GDPR
storage-limitation problem and enlarges every breach.
* Some floors are external: e.g. PCI DSS 10.5.1 keeps audit log history for
at least 12 months (3 months immediately available); HIPAA 164.316(b)(2)
keeps required Security Rule documentation for 6 years. Medical-record
retention itself comes from state law, not HIPAA.
* A LEGAL HOLD (litigation, investigation) suspends deletion for matching
records and always wins over the schedule. Releasing the hold resumes it.
* Every deletion writes evidence: what class, how many, which rule, when.
Durations below are an example schedule, not advice for your company.
"""
from dataclasses import dataclass
from datetime import date, timedelta
SCHEDULE = { # class -> (keep_days, basis)
"app_debug_logs": (30, "business need; may contain identifiers"),
"security_audit_log": (400, "PCI DSS 10.5.1 floor 12 months + margin"),
"support_tickets": (365 * 2, "business need"),
"invoices": (365 * 7, "tax law (example jurisdiction)"),
"hipaa_policies": (365 * 6, "HIPAA 164.316(b)(2), from last effective date"),
"marketing_leads": (180, "consent-based; GDPR storage limitation"),
}
@dataclass
class Record:
id: str; cls: str; created: date; customer: str
@dataclass
class Hold:
name: str; customer: str | None; classes: set; active: bool = True
def covers(self, r: Record) -> bool:
return self.active and r.cls in self.classes and (self.customer in (None, r.customer))
def decide(r: Record, today: date, holds: list) -> tuple[str, str]:
keep, basis = SCHEDULE[r.cls]
for h in holds:
if h.covers(r):
return "KEEP", f"legal hold {h.name}"
if r.created + timedelta(days=keep) <= today:
return "DELETE", f"expired ({keep}d, {basis})"
return "KEEP", f"until {r.created + timedelta(days=keep)}"
today = date(2026, 10, 8)
records = [
Record("r1", "app_debug_logs", date(2026, 8, 1), "acme"),
Record("r2", "support_tickets", date(2024, 9, 1), "acme"),
Record("r3", "support_tickets", date(2024, 9, 1), "globex"),
Record("r4", "security_audit_log", date(2026, 2, 1), "acme"),
Record("r5", "marketing_leads", date(2026, 1, 15), "-"),
Record("r6", "invoices", date(2022, 3, 1), "globex"),
]
holds = [Hold("LH-2026-04 acme dispute", "acme", {"support_tickets", "invoices", "app_debug_logs"})]
def run(dry: bool) -> None:
print(f"--- {'DRY RUN' if dry else 'EXECUTE'} {today}")
evidence = []
for r in records:
action, why = decide(r, today, holds)
print(f" {r.id} {r.cls:19} {r.customer:7} {action:6} {why}")
if action == "DELETE" and not dry:
evidence.append({"record": r.id, "class": r.cls, "rule": why, "at": str(today)})
if not dry:
print(" evidence rows written:", len(evidence))
run(dry=True)
holds[0].active = False # counsel releases the hold
print("hold released ->")
run(dry=False)Output:
--- DRY RUN 2026-10-08
r1 app_debug_logs acme KEEP legal hold LH-2026-04 acme dispute
r2 support_tickets acme KEEP legal hold LH-2026-04 acme dispute
r3 support_tickets globex DELETE expired (730d, business need)
r4 security_audit_log acme KEEP until 2027-03-08
r5 marketing_leads - DELETE expired (180d, consent-based; GDPR storage limitation)
r6 invoices globex KEEP until 2029-02-27
hold released ->
--- EXECUTE 2026-10-08
r1 app_debug_logs acme DELETE expired (30d, business need; may contain identifiers)
r2 support_tickets acme DELETE expired (730d, business need)
r3 support_tickets globex DELETE expired (730d, business need)
r4 security_audit_log acme KEEP until 2027-03-08
r5 marketing_leads - DELETE expired (180d, consent-based; GDPR storage limitation)
r6 invoices globex KEEP until 2029-02-27
evidence rows written: 4Compared: TTLs at the storage layer (object lifecycle rules, table partitions dropped by date) are cheap and reliable but cannot see legal holds or per-customer contracts; an application-level engine sees holds and contracts but must reach every store. Most teams combine them: storage TTLs as the backstop, the engine for holds, exceptions and evidence (see object-storage-lifecycle-tiers-cdn-presigned and backups-restore-pitr-immutable-backups).
5. Access reviews and just-in-time access
Standing access is the biggest gap auditors find. Two complementary controls:
- Periodic access reviews (quarterly for sensitive systems; PCI 7.2.4 requires at least every six months): managers certify each person's access, with removals tracked to completion. Generate the review from the IdP and cloud IAM, not spreadsheets.
- Just-in-time (JIT) access: no standing production or restricted-data access; engineers request time-boxed elevation with a reason and ticket, approval for sensitive scopes, automatic expiry, and session logging (see
secure-access-ssm-session-manager).
JIT shrinks what access reviews must cover and gives a perfect audit trail. Break-glass accounts remain for emergencies, with alerts on every use.
6. Policy-as-code
Write compliance rules as code evaluated in CI and at admission time: buckets must be encrypted with customer-managed keys, restricted data needs tags, public access blocked, audit log retention at least 365 days, residency-tagged data stays in allowed regions. The same check that blocks a bad change is the evidence that the control operated.
// Compliance-as-code gate for CI: evaluate planned infrastructure against rules
// that are each mapped to framework controls, block the merge on violations,
// and emit an evidence record the auditor can sample later.
// Concept: the same policy that blocks a bad change IS the evidence that the
// control operated. Real stacks use OPA/Rego (conftest), Sentinel, Checkov or
// Cedar; the shape is identical: input document -> rules -> decision + evidence.
interface Resource { addr: string; type: string; attrs: Record<string, unknown> }
interface Rule { id: string; maps: string[]; applies: (r: Resource) => boolean; ok: (r: Resource) => boolean; msg: string }
const rules: Rule[] = [
{ id: "DATA-TAG", maps: ["ISO A.5.12", "SOC2 CC6.1", "GDPR Art.30"], msg: "storage must carry a data_class tag",
applies: (r) => ["bucket", "db"].includes(r.type), ok: (r) => typeof (r.attrs.tags as Record<string, string> | undefined)?.data_class === "string" },
{ id: "ENC-REST", maps: ["HIPAA 164.312(a)(2)(iv)", "PCI 3.5.1", "SOC2 CC6.1"], msg: "storage must be encrypted with a customer-managed KMS key",
applies: (r) => ["bucket", "db"].includes(r.type), ok: (r) => typeof r.attrs.kms_key === "string" },
{ id: "NO-PUBLIC", maps: ["SOC2 CC6.6", "ISO A.8.3"], msg: "public access must be blocked on buckets",
applies: (r) => r.type === "bucket", ok: (r) => r.attrs.public_access_block === true },
{ id: "LOG-RET", maps: ["PCI 10.5.1", "SOC2 CC7.2"], msg: "audit log retention must be >= 365 days",
applies: (r) => r.type === "log_group" && r.attrs.purpose === "audit", ok: (r) => Number(r.attrs.retention_days) >= 365 },
{ id: "RESIDENCY", maps: ["customer DPA", "GDPR Ch.V transfer policy"], msg: "data tagged residency=eu must stay in eu-* regions (contractual promise)",
applies: (r) => (r.attrs.tags as Record<string, string> | undefined)?.residency === "eu", ok: (r) => String(r.attrs.region).startsWith("eu-") },
];
const plan: Resource[] = [
{ addr: "aws_s3_bucket.patient_docs", type: "bucket", attrs: { region: "us-east-1", kms_key: "arn:kms:phi", public_access_block: true, tags: { data_class: "phi" } } },
{ addr: "aws_s3_bucket.exports", type: "bucket", attrs: { region: "us-east-1", public_access_block: false, tags: {} } },
{ addr: "aws_db.eu_customers", type: "db", attrs: { region: "us-west-2", kms_key: "arn:kms:eu", tags: { data_class: "personal", residency: "eu" } } },
{ addr: "aws_log_group.audit", type: "log_group", attrs: { purpose: "audit", retention_days: 90 } },
];
const results = plan.flatMap((r) => rules.filter((x) => x.applies(r)).map((x) => ({ rule: x.id, addr: r.addr, pass: x.ok(r), maps: x.maps, msg: x.msg })));
const fails = results.filter((x) => !x.pass);
for (const f of fails) console.log(`DENY ${f.rule.padEnd(10)} ${f.addr.padEnd(28)} ${f.msg} [${f.maps.join(", ")}]`);
console.log(`checks=${results.length} pass=${results.length - fails.length} fail=${fails.length} -> merge ${fails.length ? "BLOCKED" : "allowed"}`);
const evidence = { commit: "abc1234", pipeline: "pr-512", ts: "2026-10-08T15:00:00Z", checks: results.length, failures: fails.map((f) => `${f.rule}:${f.addr}`) };
console.log("evidence:", JSON.stringify(evidence));Output:
DENY DATA-TAG aws_s3_bucket.exports storage must carry a data_class tag [ISO A.5.12, SOC2 CC6.1, GDPR Art.30]
DENY ENC-REST aws_s3_bucket.exports storage must be encrypted with a customer-managed KMS key [HIPAA 164.312(a)(2)(iv), PCI 3.5.1, SOC2 CC6.1]
DENY NO-PUBLIC aws_s3_bucket.exports public access must be blocked on buckets [SOC2 CC6.6, ISO A.8.3]
DENY RESIDENCY aws_db.eu_customers data tagged residency=eu must stay in eu-* regions (contractual promise) [customer DPA, GDPR Ch.V transfer policy]
DENY LOG-RET aws_log_group.audit audit log retention must be >= 365 days [PCI 10.5.1, SOC2 CC7.2]
checks=10 pass=5 fail=5 -> merge BLOCKED
evidence: {"commit":"abc1234","pipeline":"pr-512","ts":"2026-10-08T15:00:00Z","checks":10,"failures":["DATA-TAG:aws_s3_bucket.exports","ENC-REST:aws_s3_bucket.exports","NO-PUBLIC:aws_s3_bucket.exports","RESIDENCY:aws_db.eu_customers","LOG-RET:aws_log_group.audit"]}ExpectedDENY DATA-TAG aws_s3_bucket.exports storage must carry a data_class tag [ISO A.5.12, SOC2 CC6.1, GDPR Art.30] DENY ENC-REST aws_s3_bucket.exports storage must be encrypted with a customer-managed KMS key [HIPAA 164.312(a)(2)(iv), PCI 3.5.1, SOC2 CC6.1] DENY NO-PUBLIC aws_s3_bucket.exports public access must be blocked on buckets [SOC2 CC6.6, ISO A.8.3] DENY RESIDENCY aws_db.eu_customers data tagged residency=eu must stay in eu-* regions (contractual promise) [customer DPA, GDPR Ch.V transfer policy] DENY LOG-RET aws_log_group.audit audit log retention must be >= 365 days [PCI 10.5.1, SOC2 CC7.2] checks=10 pass=5 fail=5 -> merge BLOCKED evidence: {"commit":"abc1234","pipeline":"pr-512","ts":"2026-10-08T15:00:00Z","checks":10,"failures":["DATA-TAG:aws_s3_bucket.exports","ENC-REST:aws_s3_bucket.exports","NO-PUBLIC:aws_s3_bucket.exports","RESIDENCY:aws_db.eu_customers","LOG-RET:aws_log_group.audit"]}
Press Run. Snippets must be self-contained — no network, files, or native modules.
Compared: OPA/Rego is general-purpose and the de facto standard for Kubernetes admission and conftest; Cedar is designed for application authorization with analyzable policies; Sentinel is tied to HashiCorp products; Checkov and similar scanners ship hundreds of prebuilt rules. Prebuilt rules give fast coverage; your own rules encode your data classes and contracts. See policy-engines-opa-rego-vs-cedar-vs-custom and terraform-blast-radius-policy-secrets-ci.
How compliance controls fit into delivery
Diagram 1: a change moving from commit to production with compliance built in, with the failure path.
Decisions
- 1
Step 1: tag data classes
- nextStep 2: code review approved
- 2
Step 2: code review approved
- nextStep 3: policy gate in CI
- 3
Step 3: policy gate in CI
- nextStep 4: signed build artifact
- 4
Step 4: signed build artifact
- nextStep 5: deploy via pipeline
- 5
Step 5: deploy via pipeline
- nextStep 6: JIT access only
- 6
Step 6: JIT access only
- nextStep 7: audit log to WORM
- 7
Step 7: audit log to WORM
- nextStep 8: retention job runs
- 8
Step 8: retention job runs
- nextStep 9: control test fails?
- ?
Step 9: control test fails?
- nextStep 10: evidence auto-collected
- nextFailure: drift or incident
- 10
Step 10: evidence auto-collected
- 11
Failure: drift or incident
- nextRun IR and notify on clock
- 12
Run IR and notify on clock
- nextStep 3: policy gate in CI
Lesson map
Engineering Compliance into Systems - Data Classification, Audit Logs, Retention, JIT Access, Policy-as-Code & Continuous Evidence
Building compliance into the platform: data classification tags and lineage, key ownership, tamper-evident WORM audit logs, retention with legal hold, access reviews vs JIT access, policy-as-code gates, continuous evidence (how to evaluate Vanta, Drata and Secureframe), vendor risk, a breach notification runbook across GDPR 72h, HIPAA 60d, state laws, SEC 8-K and PCI, and compliance in CI/CD. Includes a runnable hash-chain audit log, retention engine and policy gate.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: tag data classes"] s2["Step 2: code review approved"] s3["Step 3: policy gate in CI"] s4["Step 4: signed build artifact"] s5["Step 5: deploy via pipeline"] s6["Step 6: JIT access only"] s7["Step 7: audit log to WORM"] s8["Step 8: retention job runs"] s9["Step 9: control test fails?"] s10["Step 10: evidence auto-collected"] f1["Failure: drift or incident"] f2["Run IR and notify on clock"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s8 -->|continues| s9 s9 -->|continues| s10 s9 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s3
7. Continuous evidence: Vanta, Drata and Secureframe compared
Compliance automation platforms connect to your cloud accounts, IdP, HR system, code hosting, ticketing and endpoints; run automated tests against controls; map those controls to multiple frameworks; host policies and training; track vendors; and give your auditor read access to evidence. Vanta, Drata and Secureframe are the best-known examples and cover the same core: SOC 2, ISO 27001, HIPAA, PCI DSS, GDPR and more, automated tests, policy templates, auditor marketplaces and trust centers.
| Question to ask in an evaluation | Why it matters |
|---|---|
| Do they integrate with your cloud, IdP, HRIS, code host and MDM? | Untested integrations become manual screenshots |
| Can you write custom tests (via API or code) for your own controls? | Your policy gates and retention jobs should feed evidence directly |
| How do they handle multiple frameworks from one control set? | The whole point is control mapping |
| Can you choose any auditor, and what access does the auditor get? | Independence and price |
| What data do they collect, and is that itself in scope (PHI, PAN)? | The tool becomes a vendor in your own audit |
Differences between the vendors in integrations, custom-test depth, UI and pricing change frequently; trial each against your stack. What none of them change: the tool collects evidence, it does not design controls, and the SOC 2 opinion still comes from an independent CPA firm. The alternative, a GRC spreadsheet plus scripts, is cheaper and fully flexible but needs engineering time every audit cycle.
8. Vendor risk management
You remain accountable for data you give vendors. A lightweight program: an inventory of vendors and the data classes they receive; tiering (restricted data or production access means high tier); for high tier, review SOC 2 or ISO evidence yearly, including scope, exceptions and CUECs; contracts with the right instruments (BAA for PHI, DPA and SCCs for GDPR, PCI responsibility matrix and AOC for payment providers); a public subprocessor list where customers expect it; and offboarding that deletes data and revokes credentials. Engineering's part: egress allowlists and secrets inventories show which vendors actually receive data, which often differs from what procurement believes.
Incident response and the breach notification runbook
One incident can start several legal clocks. The runbook's job is to gather facts fast and route them to counsel, who decides what is a notifiable breach.
| Regime | Trigger | Clock | Who |
|---|---|---|---|
| GDPR Art. 33/34 | Personal data breach likely to result in risk | Regulator within 72 hours of awareness where feasible; individuals without undue delay if high risk | Controller; processors tell controllers without undue delay |
| HIPAA Breach Notification Rule | Breach of unsecured PHI | Individuals within 60 days of discovery; HHS and media per the 500 thresholds | Covered entity; BA tells covered entity within 60 days, usually sooner by contract |
| US state laws | Breach of defined personal information (all 50 states have laws) | Varies; several set fixed outer limits around 30 to 60 days, many require AG notice above thresholds | Data owner; service providers notify owners |
| SEC Form 8-K Item 1.05 | Material cybersecurity incident at a public company | 4 business days after determining materiality | Registrant |
| PCI DSS | Suspected compromise of account data | Immediately, per your acquirer's and card brands' rules | Merchant or service provider |
| Contracts | Defined in BAAs, DPAs, MSAs | Often 24 to 72 hours | Vendor to customer |
Engineering prerequisites: the data inventory (which subjects and data classes were in the affected system, and which states and countries they live in), audit logs that show what was actually accessed (the difference between "accessed" and "might have been accessed" decides notification scope), key status (encrypted data with uncompromised keys is often not reportable), and contact lists for regulators, acquirers and customers. Practice with tabletop exercises; NIST SP 800-61 Rev. 3 aligns incident response with CSF 2.0.
Compliance in CI/CD
The pipeline is where many controls actually operate, and where auditors sample: protected branches with required review (SOC 2 CC8.1, ISO A.8.32), policy-as-code checks on infrastructure, dependency and container scanning with remediation SLAs (PCI 6.3.3 for critical patches), signed artifacts and provenance (see supply-chain-signing-sboms-oidc-federation), secrets scanning, deploys only through the pipeline with no direct production write access, and automatic export of these results as evidence. Emergency changes get a documented fast path with after-the-fact review, not a bypass.
Decision chart: which control to build first?
Decisions
- 1
Q1
- nextClassify and tag data first
- nextQ2: can anyone read it without trace?
- 2
Classify and tag data first
- ?
Q2: can anyone read it without trace?
- nextAudit logs and JIT access
- nextQ3: kept longer than needed?
- 4
Audit logs and JIT access
- ?
Q3: kept longer than needed?
- nextRetention engine and TTLs
- nextQ4: checks manual each audit?
- 6
Retention engine and TTLs
- ?
Q4: checks manual each audit?
- nextPolicy-as-code and evidence automation
- nextQ5: breach clock rehearsed?
- 8
Policy-as-code and evidence automation
- ?
Q5: breach clock rehearsed?
- nextIR runbook and tabletop
- nextVendor reviews and continuous tests
- 10
IR runbook and tabletop
- 11
Vendor reviews and continuous tests
What happens if you choose otherwise
- Classify in a spreadsheet: new tables and buckets appear unclassified; erasure and scoping miss them.
- Keep audit logs in the app database: the people being audited can edit them, and an attacker with app credentials erases traces.
- Grant standing prod access to all engineers: every access review is a long list of exceptions, and any compromised laptop reaches restricted data.
- Buy a compliance platform before designing controls: you automate screenshots of weak controls.
Pitfalls
- Retention jobs that delete from the primary store but not from replicas, backups or derived datasets.
- Policy gates that teams bypass with a label nobody reviews.
- Incident runbooks with no owner for regulator communication.
- Evidence that depends on one person's laptop or memory.
How the code was checked
- The hash-chain audit log and the retention engine ran under Python 3.13 and the policy gate under
tsc --strictand Node 20. Their output blocks are the real captured output. - Resources, tenants and holds are example data; framework citations are abbreviated.
Interview Q&A
What is the first engineering control for any compliance program and why?
Answer
Data classification and inventory with tags in the systems themselves. Scoping, encryption, retention, erasure and breach assessment all depend on knowing which data is where.
How do you make audit logs trustworthy?
Answer
Separate them from app logs, write them to WORM storage in an account the app team cannot administer, hash-chain the entries and ship signed checkpoints elsewhere, and alert when a source goes quiet.
Why is a hash chain alone not enough?
Answer
An attacker who controls the log can recompute the whole chain. Signed checkpoints stored separately, plus WORM storage, make rewrites detectable and deletions impossible within retention.
How do retention and legal hold interact?
Answer
The schedule deletes data when its retention period expires; a legal hold suspends deletion for matching data until counsel releases it. Holds always win, and both actions should leave evidence.
Access reviews vs JIT access?
Answer
Reviews periodically certify standing access; JIT removes most standing access by granting time-boxed, approved, logged elevation. JIT reduces review scope and produces a better audit trail; you usually need both.
What is policy-as-code and why do auditors like it?
Answer
Compliance rules expressed as code and enforced in CI or admission control. The enforcement result itself is evidence that the control operated for every change, not just a sample.
What do compliance automation platforms do and not do?
Answer
They integrate with your stack, run automated control tests, map controls to frameworks, manage policies, vendors and evidence, and give auditors access. They do not design controls or replace the independent auditor.
An attacker exfiltrated a database with EU and California users, some PHI. Which clocks start?
Answer
GDPR 72 hours to the supervisory authority where feasible, HIPAA 60 days to individuals with HHS and media thresholds, California and other state notification laws, contractual notice to customers, and an SEC 8-K within four business days of a materiality determination if public. Counsel decides; engineering supplies scope fast.
Why does key ownership matter for compliance?
Answer
Breach safe harbors and scope reduction depend on who can decrypt. Customer-managed and per-tenant keys enable crypto-shredding, offboarding and proof of access, and separate key admins from data users.
What compliance controls live in CI/CD?
Answer
Required code review, policy-as-code on infrastructure, dependency and image scanning with SLAs, secrets scanning, signed artifacts with provenance, pipeline-only deploys and automatic evidence export.
Check yourself
Write three policy-as-code rules for your own infrastructure (tags present, customer-managed encryption, audit log retention of at least 365 days) and run them against your current Terraform plan. Count the failures and decide which ones block merges.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Policy Engines — OPA/Rego vs Cedar vs Custom, Blast Radius — Policy-as-Code, Secrets & CI Applies, Supply Chain Security — Signing, SBOMs & OIDC Federation, CI/CD Pipelines — Stages, Artifacts, Caching & Supply Chain, Secure Access Without SSH — AWS SSM Session Manager, Bastions & Alternatives, Secrets Threat Model — Leakage Paths, Side Channels & Audit Trails, Secrets & KMS — Envelope Encryption, Rotation & Blast Radius, Envelope Encryption — DEK, KEK & CMK Hierarchy, ABAC — Attributes, Policies, PDP & PEP, RBAC — Roles, Permissions & Role Explosion, Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore Drills, Lifecycle, Storage Classes, CDN & Presigned URLs, Region Failover in Practice - Runbooks, DNS/GLB Cutover, Failback & Game Days, Observability Triad — Metrics, Logs & Distributed Tracing.
Go Deeper
- NIST SP 800-53 Rev. 5: Security and Privacy Controls
- NIST SP 800-92: Guide to Computer Security Log Management
- NIST SP 800-88 Rev. 2: Guidelines for Media Sanitization
- NIST SP 800-61 Rev. 3: Incident Response Recommendations
- NIST Privacy Framework
- HHS: Submitting notice of a breach to the Secretary
- GDPR Art. 33: notification of a personal data breach
- California Attorney General: data breach reporting
- PCI SSC: Document library