Security
Part 5 of 6 · Secrets & KMSApp Secret Injection — Env, Sidecar/Agent, CSI Drivers & Runtime Fetch
How an app obtains a secret matters as much as where it is stored. This lesson compares environment variables, file mounts, Vault Agent and cloud sidecars, the Secrets Store CSI driver, and runtime SDK fetch — with blast radius and operational tradeoffs for Kubernetes and VMs.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Sidecar or agent versus an in-process SDK
Prefer
Pick from who must change on rotation
A sidecar keeps the app language-agnostic and renews in one place. An SDK keeps the pod smaller and makes retries explicit. Both fail if they log the value.
- Sidecar: centralized renew and templates. Extra container, noisy neighbor, another failure domain.
- SDK: fewer moving parts, explicit retries. Every language reimplements the cache.
- Either way, cache TTL is much shorter than the lease.
- The store still authorizes the workload identity, not the file mode alone.
Alternative
Bake the value into the image
Dockerfile ENV or COPY. Every registry pull and every old tag is a copy of the secret.
- Image scrapes keep the value forever.
- Rotation means a rebuild and a redeploy of every tag you forgot.
- Layer history survives the running pod.
- There is no production case for this path.
Overview
How an app obtains a secret matters as much as where it is stored. Environment variables, file mounts, Vault Agent or a cloud sidecar, the Secrets Store CSI driver, and a runtime SDK fetch have different blast radii. This lesson is that comparison for Kubernetes and for VMs.
The store choice is the previous store lesson. The overlap that makes a refreshed file or cache safe is rotation.
Injection options
| Method | Mechanism | Pros | Cons | Prefer |
|---|---|---|---|---|
| Env vars | kube env or process env | Simple | Leaks via proc, forks, dumps | Local only, or non-prod |
| File mount (tmpfs) | Secret volume mode 0600 | Not in argv | Still on a filesystem briefly | Twelve-factor file readers |
| Sidecar or agent | Renders files or wraps an API | Central renew, templates | Extra CPU and memory | Vault Agent or cloud agents |
| CSI driver | Mount from an external store | Can avoid a long-lived Secret copy | Driver and provider ops | Kubernetes plus Vault or cloud SM |
| Runtime fetch SDK | App calls the SM or Vault API | Short cache, identity-bound | App code and cold start | Services that already have IAM or workload identity |
| Bake into the image | Dockerfile ENV or COPY | None in production | Image scrapes forever | Never |
Request path
Sequence
- 1
Pod → CSI or Agent
Mount or fetch secret path
- 2
CSI or Agent → Vault or Cloud SM
Auth via ServiceAccount or IAM
- 3
Vault or Cloud SM → KMS
Decrypt underlying CMK if needed
- 4
KMS → Vault or Cloud SM
ok
- 5
Vault or Cloud SM → CSI or Agent
secret payload, short TTL
- 6
CSI or Agent → Pod
tmpfs file or localhost API
- 7
Pod
App reads. Cache TTL is less than the lease
Lesson map
App Secret Injection — Env, Sidecar/Agent, CSI Drivers & Runtime Fetch
How an app obtains a secret matters as much as where it is stored. This lesson compares environment variables, file mounts, Vault Agent and cloud sidecars, the Secrets Store CSI driver, and runtime SDK fetch — with blast radius and operational tradeoffs for Kubernetes and VMs.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB pod["Pod"] csi["CSI or Agent"] store["Vault or Cloud SM"] kms["KMS"] pod -->|Mount or fetch| csi csi -->|Auth via| store store -->|Decrypt| kms kms -->|ok| store store -->|secret payload,| csi csi -->|tmpfs file or| pod
Kubernetes, lightly
Prefer workload identity over static cloud keys in Secrets. How that identity is issued is the OAuth 2.1 & OIDC hub, not a second copy of the grant dance here.
CSI can sync as files without persisting a Secret object, depending on the provider config. If you must use native Secrets, enable etcd encryption, tight RBAC, and no standing get for developers in prod. Who is allowed that get is the Authorization hub.
Runtime fetch rules
- Authenticate with workload identity, not an embedded token.
- Cache in memory with a TTL much shorter than the lease, and jitter the refresh.
- Fail closed on renew errors after the grace period. Do not log secret values.
- Separate the bootstrap secret (how you reach the store) from application secrets.
The bootstrap problem is chicken-and-egg. Solve it with platform identity: a Kubernetes ServiceAccount or an instance role. A token checked into the repo is not a bootstrap design.
Runnable fetch client
The sandbox has no network. The dictionary stands in for an IAM-authenticated GET. The behavior under test is the cache: the second read does not fetch again until the TTL expires.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Prefer a file over the environment
Production reads a mode 0600 tmpfs file. This sandbox has no filesystem, so a map stands in for the file and a plain object stands in for the environment. The branch is the lesson: use the file when it exists, and warn when you fall back to the environment.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Sidecar versus SDK
Sidecar or agent. The app stays simple. Renew and templates live in one container. You pay an extra container to schedule, a noisy neighbor, and another failure domain.
SDK. Fewer moving parts in the pod, and retries are explicit in the service. Every language team reimplements caching, and it is easy to log the secret on an error path.
Same-pod sidecars share the network and filesystem trust domain. Harden with a minimal ServiceAccount, a read-only mount, and policy on the store. A sidecar can read what the mount allows. Store policy is the real boundary.
Rotate without a restart
An agent rewrites the file and the app re-reads on SIGHUP or on a poll. An SDK expires the cache and pulls the new version. Either path still needs the overlap window, because some processes will see the old file until the next poll.
Interview Q&A
Why are environment variables discouraged for secrets?
Answer
They are visible to process listings, inherited by children, and copied into diagnostics. Rotation usually means a restart. A tmpfs file or a runtime fetch narrows who sees the value and lets you refresh without rebuilding the image.
What does the Secrets Store CSI driver give you?
Answer
A mount of secrets from Vault or a cloud Secrets Manager into the pod filesystem, using the pod identity. That reduces static Secret sprawl. You still operate the driver and the provider, and you still set a sync period that fits the rotation overlap.
How does Vault Agent help?
Answer
It authenticates, renews tokens, and renders secret files for the app. The application can stay a file reader. The agent is the process that understands leases.
Runtime fetch versus a mount?
Answer
Fetch is fresher and puts more logic in the app. A mount is simpler for the app, and freshness depends on the agent or CSI sync period. Pick fetch when rotation speed matters more than a dumb file reader.
What is the bootstrap secret problem?
Answer
You need a credential to reach the store, and that credential cannot itself be a checked-in token. Platform identity — a Kubernetes ServiceAccount or an instance role — is the bootstrap. Application secrets come after that identity exists.
Can a sidecar steal secrets?
Answer
It shares the pod's network and filesystem trust domain, so a mount it can read is a mount it can exfiltrate. Limit the ServiceAccount, use a read-only mount, and enforce path policy on the store. The sidecar is not a security boundary against the rest of the pod.
How do you rotate without a restart?
Answer
The agent rewrites the file and the app re-reads on SIGHUP or a poll. Or the SDK cache expires and the next get pulls the new version. Processes that never re-read will keep the old value until they restart, which is why the overlap exists.
When do you still pick a sidecar over an SDK?
Answer
Many languages, a team that should not reimplement renew, and a file-shaped contract the app already understands. Pick the SDK when the pod budget is tight and one service owns its client. Do not pick either as a reason to log the payload.
Pitfalls
The database password rotates every day. The service must not restart to pick it up. Choose file-plus-agent or SDK cache, name the TTL, and say what the process does when the agent is crash-looping.