Kubernetes Storage - Volumes, PV, PVC, StorageClass, CSI, StatefulSets & Failure Modes
Storage: volume types, PV/PVC/StorageClass, static vs dynamic provisioning, Immediate vs WaitForFirstConsumer, provision -> bind -> attach -> mount via CSI, access modes (RWO/RWOP/RWX), reclaim policies, StatefulSets with volumeClaimTemplates, expansion and snapshots, zone pinning, Multi-Attach and node-loss failure timelines.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
PV vs PVC vs StorageClass?
Answer
A PVC is the namespaced request, a PV is the real storage, a StorageClass tells a provisioner how to create PVs.
L2
Static vs dynamic provisioning?
Answer
Static uses PVs an admin created; dynamic has the CSI driver create a PV per claim.
L3
Why WaitForFirstConsumer?
Answer
So a zonal volume is created where the scheduler already placed the pod.
L4
What does RWO actually mean?
Answer
Read-write by a single node; several pods on that node can share it. RWOP means a single pod.
L5
Retain vs Delete reclaim policy?
Answer
Delete removes the backing disk with the PVC; Retain keeps the disk for an admin to recover.
L6
What causes a Multi-Attach error?
Answer
An RWO volume is still attached to the old node, usually after that node died.
L7
How do you shorten recovery after a node loss?
Answer
Confirm the node is off, then apply the out-of-service taint so pods are force-deleted and the volume detaches immediately.
Failure modes
Volume node affinity conflict
Immediate binding created a zonal disk in a zone where the pod cannot run.
Multi-Attach after node loss
The RWO disk stays attached to the dead node, so the replacement waits about 12 minutes with defaults.
Data lost with a PVC delete
The StorageClass used the Delete reclaim policy for important data.
Misconceptions
RWO means one pod.
It means one node. Use RWOP for one pod.
A StatefulSet is a database operator.
It gives stable identity, storage and ordering. Replication, failover and backups are still your job.
A disk snapshot is a consistent backup.
It is crash-consistent. Quiesce or use the database's own tools for an application-consistent backup.
Interviewer traps
Adding the out-of-service taint without confirming the node is off.
If the node is still writing, two writers can corrupt the filesystem.
Using a RollingUpdate Deployment with an RWO PVC.
The new pod cannot attach while the old one holds the disk. Use Recreate or a StatefulSet.
Design scenario
Same prompt for every reader.
Requirements
One PVC per replica, no data loss on accidental deletes, and a documented node-loss runbook.
Failure assumptions
- A node powers off without draining.
- One zone runs out of capacity.
- Someone deletes a PVC by mistake.
Constraints
- RWO block storage only.
- Default kube-controller-manager timers.
Prompt
Run a three-replica Postgres on Kubernetes across three zones with zonal block storage, and keep recovery from a node loss under five minutes.
API
Which StorageClass settings and StatefulSet fields do you set?
Data
Which reclaim policy, snapshots and backups protect the data?
Architecture
How do replicas, zones, the attach/detach controller and the out-of-service taint interact in a node loss?
Overview
Kubernetes storage separates asking for storage from providing it, the same way pods ask for CPU without naming a machine. A pod references a PersistentVolumeClaim (PVC): "10 GiB, mounted read-write by one node, from the fast-ssd class". A PersistentVolume (PV) is the actual piece of storage, either created ahead of time by an admin (static) or created on demand by a CSI driver following a StorageClass (dynamic). A control loop binds one PVC to one PV. Then the attach/detach controller connects the disk to the node, and the kubelet mounts it into the pod.
The hard part is that storage, unlike CPU, has a location and an owner. A cloud block disk lives in one zone and can be attached to one node at a time. That single fact explains WaitForFirstConsumer, StatefulSets with stable identities, "volume node affinity conflict", "Multi-Attach error", and why a node failure can keep a database down for many minutes.
Volume types and the objects involved
| Thing | Lifetime | Who creates it | Example | Use for |
|---|---|---|---|---|
emptyDir | The pod | The pod spec | Scratch space, caches, sharing files between containers | Data you can lose on reschedule |
configMap, secret, projected, downwardAPI | The pod (content from API objects) | The pod spec | Config files, tokens, certificates | Read-only config injection |
hostPath | The node | The pod spec | Node agents reading /var/log | Node-level agents only; a security risk elsewhere |
| Generic ephemeral volume | The pod, but provisioned by a CSI driver | The pod spec creates a PVC owned by the pod | Large per-pod scratch on fast disks | Scratch bigger than node disk |
| PersistentVolumeClaim | Independent of pods | Users, or StatefulSet volumeClaimTemplates | data-postgres-0 | Anything that must survive pod restarts |
| PersistentVolume | Independent of pods and claims | Admin (static) or provisioner (dynamic) | A cloud disk, an NFS export | The real storage asset |
| StorageClass | Cluster-wide | Admin | gp3-encrypted, standard-rwo | Provisioner, parameters, reclaim policy, binding mode, expansion |
| VolumeSnapshot / VolumeSnapshotClass | Independent | Users / admin | Point-in-time copy of a PVC | Backups, cloning environments |
Immediate binding or WaitForFirstConsumer?
Prefer
WaitForFirstConsumer
Schedule the pod first, then create the volume in that zone.
- The scheduler picked node-b1 in zone-b.
- The disk was then created in zone-b.
- The pod's affinity and tolerations are respected.
Alternative
Immediate binding
Provision the moment the PVC is created, before any pod is scheduled.
- The disk was created in zone-a.
- The pod was unschedulable: volume node affinity conflict.
- Fixing it means deleting and re-creating the claim.
From a claim to a mounted volume
Diagram 1 condensed: provision, bind, attach, mount, and the node-loss exit.
- 1
Claim storage
The pod references a PVC with size, access mode and class. - 2
Provision and bind
The CSI driver creates a PV and the binder links it to the claim. - 3
Attach to the node
The attach/detach controller and external-attacher call ControllerPublishVolume. - 4
Stage and mount
The kubelet calls NodeStageVolume and NodePublishVolume. - 5
Node loss
The RWO disk stays attached to the dead node until a forced detach or the out-of-service taint.
The lifecycle: provision, bind, attach, mount
Decisions
- 1
Step 1: user creates a PVC (size, accessModes, storageClassName)
- nextStep 2: StorageClass volumeBindingMode?
- ?
Step 2: StorageClass volumeBindingMode?
- nextStep 3a: external-provisioner calls CSI CreateVolume right away, in a zone chosen without the pod
- nextStep 3b: PVC waits; when a pod uses it, the scheduler picks a node and annotates the selected node
- 3
Step 3a: external-provisioner calls CSI CreateVolume right away, in a zone chosen without the pod
- nextStep 5: PV created and bound one-to-one to the PVC (claimRef)
- nextFailure path: volume node affinity conflict, pod Pending
- 4
Step 3b: PVC waits; when a pod uses it, the scheduler picks a node and annotates the selected node
- nextStep 4: provisioner calls CreateVolume with that node's topology
- 5
Step 4: provisioner calls CreateVolume with that node's topology
- nextStep 5: PV created and bound one-to-one to the PVC (claimRef)
- 6
Step 5: PV created and bound one-to-one to the PVC (claimRef)
- nextStep 6: attach/detach controller creates a VolumeAttachment, external-attacher calls ControllerPublishVolume
- 7
Step 6: attach/detach controller creates a VolumeAttachment, external-attacher calls ControllerPublishVolume
- nextStep 7: kubelet calls the CSI node plugin: NodeStageVolume (format, mount to staging), NodePublishVolume (bind mount into pod)
- nextFailure path: Multi-Attach error until detach or force detach
- 8
Step 7: kubelet calls the CSI node plugin: NodeStageVolume (format, mount to staging), NodePublishVolume (bind mount into pod)
- nextStep 8: containers start with the volume mounted
- 9
Step 8: containers start with the volume mounted
- 10
Failure path: volume node affinity conflict, pod Pending
- 11
Failure path: Multi-Attach error until detach or force detach
Lesson map
Kubernetes Storage - Volumes, PV, PVC, StorageClass, CSI, StatefulSets & Failure Modes
Storage: volume types, PV/PVC/StorageClass, static vs dynamic provisioning, Immediate vs WaitForFirstConsumer, provision -> bind -> attach -> mount via CSI, access modes (RWO/RWOP/RWX), reclaim policies, StatefulSets with volumeClaimTemplates, expansion and snapshots, zone pinning, Multi-Attach and node-loss failure timelines.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB p["Step 1: user creates a PVC (size, accessModes, storageClassName)"] w["Step 2: StorageClass volumeBindingMode?"] pr["Step 3a: external-provisioner calls CSI CreateVolume right away, in a zone chosen without the pod"] sch["Step 3b: PVC waits when a pod uses it, the scheduler picks a node and annotates the selected node"] pr2["Step 4: provisioner calls CreateVolume with that node's topology"] b["Step 5: PV created and bound one-to-one to the PVC (claimRef)"] at["Step 6: attach/detach controller creates a VolumeAttachment, external-attacher calls ControllerPublishVolume"] mo["Step 7: kubelet calls the CSI node plugin: NodeStageVolume (format, mount to staging), NodePublishVolume (bind mount into pod)"] run["Step 8: containers start with the volume mounted"] f1["Failure path: volume node affinity conflict, pod Pending"] f2["Failure path: Multi-Attach error until detach or force detach"] p -->|continues| w w -->|continues| pr w -->|continues| sch sch -->|continues| pr2 pr -->|continues| b pr2 -->|continues| b b -->|continues| at at -->|continues| mo mo -->|continues| run pr -->|continues| f1 at -->|continues| f2
Static vs dynamic provisioning
| Static | Dynamic | |
|---|---|---|
| Who makes the disk | An admin, ahead of time, creates PVs | The CSI driver on demand, per PVC |
| Matching | Binder finds a PV whose size, class, access modes (and selector) satisfy the claim; you may get a bigger one | The new PV is always bound to the PVC that triggered it |
| Good for | Pre-existing data, NFS exports, importing a disk into the cluster | Almost everything else |
| Pitfall | A 100 GiB claim never binds if only 50 GiB PVs exist; claims with storageClassName: "" only match class-less PVs | A missing or wrong default StorageClass leaves PVCs Pending |
Binding modes and why WaitForFirstConsumer exists
With Immediate (the default when unset), binding and provisioning happen as soon as the PVC is created, without knowledge of the pod's scheduling requirements. For zonal storage that can create a disk in a zone where the pod can never run. WaitForFirstConsumer delays binding and provisioning until a pod using the claim is scheduled, so the volume follows the pod's resource requests, node selectors, affinity rules and tolerations.
"""Simulation: volume binding modes and StatefulSet volume identity.
Part 1 models the documented difference between volumeBindingMode Immediate
and WaitForFirstConsumer for zonal block storage.
Part 2 models StatefulSet volumeClaimTemplates: claim names are
<template>-<statefulset>-<ordinal>, and PVCs are retained on scale-down by default.
"""
NODES = {"node-a1": "zone-a", "node-b1": "zone-b"} # node -> zone
NODE_HAS_ROOM = {"node-a1": False, "node-b1": True} # zone-a node is full
def provision(zone, claim):
print(f" provisioner: created disk for {claim} in {zone} (PV nodeAffinity zone={zone})")
return {"zone": zone}
def schedule(pod, pv=None):
feasible = [n for n, z in NODES.items() if NODE_HAS_ROOM[n] and (pv is None or pv["zone"] == z)]
if not feasible:
why = "volume node affinity conflict" if pv else "no room"
print(f" scheduler: {pod} unschedulable: {why}")
return None
print(f" scheduler: {pod} -> {feasible[0]} ({NODES[feasible[0]]})")
return feasible[0]
print("Immediate: the PVC is provisioned the moment it is created")
pv = provision("zone-a", "data-db") # provisioner picks a zone with no knowledge of the pod
schedule("db-pod", pv)
print("\nWaitForFirstConsumer: the scheduler picks a node first, then the volume follows")
node = schedule("db-pod") # VolumeBinding plugin defers the decision
pv = provision(NODES[node], "data-db")
# ---------------- Part 2: StatefulSet identity --------------------------------
print("\nStatefulSet web (replicas 3 -> 1 -> 3), podManagementPolicy OrderedReady")
pvcs = {}
def scale(replicas, current):
if replicas > current:
for i in range(current, replicas): # create in order 0..N-1
claim = f"data-web-{i}"
status = "re-attached existing" if claim in pvcs else "created new"
pvcs.setdefault(claim, f"pv-{i}")
print(f" web-{i} Running+Ready, {status} PVC {claim} -> {pvcs[claim]}")
else:
for i in reversed(range(replicas, current)): # terminate in reverse order
print(f" web-{i} terminated; PVC {f'data-web-{i}'} kept (whenScaled: Retain)")
return replicas
n = scale(3, 0)
n = scale(1, n)
n = scale(3, n)
print(" PVCs left in namespace:", sorted(pvcs))Output:
Immediate: the PVC is provisioned the moment it is created
provisioner: created disk for data-db in zone-a (PV nodeAffinity zone=zone-a)
scheduler: db-pod unschedulable: volume node affinity conflict
WaitForFirstConsumer: the scheduler picks a node first, then the volume follows
scheduler: db-pod -> node-b1 (zone-b)
provisioner: created disk for data-db in zone-b (PV nodeAffinity zone=zone-b)
StatefulSet web (replicas 3 -> 1 -> 3), podManagementPolicy OrderedReady
web-0 Running+Ready, created new PVC data-web-0 -> pv-0
web-1 Running+Ready, created new PVC data-web-1 -> pv-1
web-2 Running+Ready, created new PVC data-web-2 -> pv-2
web-2 terminated; PVC data-web-2 kept (whenScaled: Retain)
web-1 terminated; PVC data-web-1 kept (whenScaled: Retain)
web-1 Running+Ready, re-attached existing PVC data-web-1 -> pv-1
web-2 Running+Ready, re-attached existing PVC data-web-2 -> pv-2
PVCs left in namespace: ['data-web-0', 'data-web-1', 'data-web-2']Access modes and reclaim policies
| Access mode | Meaning (per the Kubernetes docs) | Typical backends |
|---|---|---|
| ReadWriteOnce (RWO) | Read-write by a single node; several pods on that node can still share it | Cloud block disks (EBS, Persistent Disk, Azure Disk) |
| ReadOnlyMany (ROX) | Read-only by many nodes | NFS, some CSI drivers, snapshots cloned read-only |
| ReadWriteMany (RWX) | Read-write by many nodes | NFS, CephFS, cloud file services (EFS, Filestore, Azure Files) |
| ReadWriteOncePod (RWOP) | Read-write by a single pod in the whole cluster; CSI only; stable since v1.29 | CSI drivers that support it |
Access modes are used for matching and attach rules; they do not make a mounted filesystem read-only by themselves. A volume is mounted with one access mode at a time.
| Reclaim policy | After the PVC is deleted | Use when |
|---|---|---|
| Delete (default for dynamic PVs, inherited from the StorageClass) | PV and the backing disk are deleted | Disposable or easily rebuilt data |
| Retain | PV becomes Released; the disk and data remain until an admin cleans up or re-binds | Databases and anything you must not lose to a mistaken delete |
| Recycle | Deprecated basic scrub; only nfs and hostPath in 1.37 | Do not use; prefer dynamic provisioning |
Protection layers you get for free: the kubernetes.io/pvc-protection and pv-protection finalizers keep an in-use PVC or a bound PV from disappearing under a running pod, and the external-provisioner.volume.kubernetes.io/finalizer makes sure a Delete PV is removed only after its backing volume is deleted.
CSI: how a storage driver plugs in
A CSI driver runs as two parts. The controller plugin (a Deployment or StatefulSet) talks to the storage backend's API; the node plugin (a DaemonSet) does the work on each node. Kubernetes-maintained sidecars translate API objects into CSI gRPC calls.
| CSI call | Part | Triggered by | Sidecar or caller |
|---|---|---|---|
CreateVolume / DeleteVolume | Controller | PVC created / PV deleted | external-provisioner |
ControllerPublishVolume / ControllerUnpublishVolume (attach and detach) | Controller | VolumeAttachment objects from the attach/detach controller | external-attacher |
ControllerExpandVolume | Controller | PVC size increased | external-resizer |
CreateSnapshot | Controller | VolumeSnapshot | external-snapshotter (plus the snapshot controller) |
NodeStageVolume / NodePublishVolume | Node | Pod scheduled to the node | kubelet |
NodeExpandVolume | Node | Filesystem resize after the disk grew | kubelet |
| Registration | Node | Driver start | node-driver-registrar |
The CSI spec requires ControllerUnpublishVolume (detach) to happen only after every NodeUnpublishVolume and NodeUnstageVolume for that volume has succeeded. That rule protects data, and it is exactly what makes a dead node a problem: nobody can confirm the unmount.
StatefulSets: stable identity for stateful pods
A StatefulSet gives each replica a stable name (web-0, web-1), a stable network identity through a headless Service (web-0.web.ns.svc), and stable storage through volumeClaimTemplates, which create one PVC per ordinal named <template>-<statefulset>-<ordinal> (data-web-0). If web-1 is rescheduled, it comes back as web-1 and re-attaches data-web-1.
| Setting | Default | Alternatives | Effect |
|---|---|---|---|
podManagementPolicy | OrderedReady: create 0..N-1 one at a time, each Running and Ready first; delete in reverse | Parallel | Parallel is faster but breaks "replica 0 is up before replica 1" assumptions |
updateStrategy | RollingUpdate in reverse ordinal order, partition for canaries | OnDelete | maxUnavailable (beta since v1.35) allows updating several pods at once |
persistentVolumeClaimRetentionPolicy | whenDeleted: Retain, whenScaled: Retain (stable since v1.32) | Delete for either | Retain protects data; Delete avoids paying for orphaned disks |
ordinals.start | 0 | Any number | Useful when splitting or migrating a StatefulSet across clusters |
A StatefulSet is not a database operator. It gives you identity and ordered rollout; replication, failover, backups and fencing are still your job or an operator's. See Failover & Split Brain for why "at most one writer" matters.
Expansion and snapshots
- Expansion (stable since v1.24) requires
allowVolumeExpansion: trueon the StorageClass. Edit the PVC's requested size; the existing volume grows (no new PV), and the filesystem is resized online if supported (XFS, ext3, ext4). You cannot shrink below current capacity. If an expansion is too large for the backend, you can retry with a smaller value that is still above current capacity. - Snapshots (stable since v1.20, CSI only) are
VolumeSnapshotobjects. Restore by creating a new PVC withdataSourcepointing at the snapshot. A crash-consistent disk snapshot is not an application-consistent backup; quiesce or use the database's own tools. Backups That Actually Restore covers the difference.
Failure modes and what the clock looks like
| Symptom | Cause | Fix or prevention |
|---|---|---|
PVC Pending, no events | No default StorageClass, or storageClassName typo | Set a default class; check events on the PVC |
| Pod Pending: volume node affinity conflict | Zonal PV in zone a, pod can only run in zone b (often Immediate binding or a full zone) | WaitForFirstConsumer; capacity in every zone; topology spread aligned with storage |
Multi-Attach error for volume | RWO disk still attached to the old node | Wait for detach, use the out-of-service taint after confirming shutdown, or use RWX storage if truly needed |
Pod stuck ContainerCreating with mount timeouts | Node plugin down, stale device, or filesystem check | Check the CSI node DaemonSet and kubelet logs |
| StatefulSet replica never comes back after node loss | Old pod object stuck Terminating, so the controller will not create a duplicate identity | Non-graceful node shutdown procedure |
| Deleted a PVC and lost data | Reclaim policy Delete | Retain for important classes; snapshots |
// Simulation: a node holding a ReadWriteOnce volume powers off.
// Uses documented defaults where they exist; ATTACH_MOUNT_S is an example value.
// node-monitor-grace-period (kube-controller-manager) 50s
// default-unreachable-toleration-seconds (apiserver) 300s
// forced detach when pod deletion has not succeeded 6 min, if node unhealthy
const NODE_GRACE_S = 50;
const UNREACHABLE_TOLERATION_S = 300;
const FORCE_DETACH_S = 360;
const ATTACH_MOUNT_S = 30; // example: cloud attach + filesystem mount
type Step = [number, string];
const fmt = (s: number) => `${String(Math.floor(s / 60)).padStart(2, "0")}:${String(s % 60).padStart(2, "0")}`;
const show = (title: string, steps: Step[]) => {
console.log(title);
for (const [t, msg] of steps) console.log(` t=${fmt(t)} ${msg}`);
};
function defaultPath(): Step[] {
const notReady = NODE_GRACE_S;
const evict = notReady + UNREACHABLE_TOLERATION_S;
const detach = evict + FORCE_DETACH_S;
return [
[0, "node-a powers off; kubelet stops renewing its Lease"],
[notReady, "node controller marks node-a NotReady/Unknown, adds unreachable:NoExecute taint"],
[evict, "pod tolerationSeconds expires -> pod deleted, stuck Terminating (kubelet cannot confirm)"],
[evict + 1, "Deployment creates replacement on node-b -> Multi-Attach error: volume still attached to node-a"],
[detach, "6 min without successful deletion and node unhealthy -> attach/detach controller force-detaches"],
[detach + ATTACH_MOUNT_S, "volume attached to node-b, mounted, container Running"],
];
}
function outOfServicePath(operatorAt: number): Step[] {
return [
[0, "node-a powers off"],
[NODE_GRACE_S, "node NotReady; on-call confirms the VM is really off (not rebooting)"],
[operatorAt, "kubectl taint node node-a node.kubernetes.io/out-of-service=nodeshutdown:NoExecute"],
[operatorAt + 1, "pods on node-a force-deleted, volume detach starts immediately"],
[operatorAt + 1 + ATTACH_MOUNT_S, "replacement attached and Running on node-b"],
];
}
const a = defaultPath();
const b = outOfServicePath(180);
show("Default timers (Deployment, RWO volume):", a);
show("\nNon-graceful shutdown with the out-of-service taint:", b);
console.log(`\ndowntime: default ~${fmt(a[a.length - 1][0])}, with out-of-service taint ~${fmt(b[b.length - 1][0])}`);
console.log("StatefulSet note: the controller will not create web-1 again while the old web-1 object exists,");
console.log("so without the taint (or a manual force delete) the replacement never starts at all.");Output:
Default timers (Deployment, RWO volume):
t=00:00 node-a powers off; kubelet stops renewing its Lease
t=00:50 node controller marks node-a NotReady/Unknown, adds unreachable:NoExecute taint
t=05:50 pod tolerationSeconds expires -> pod deleted, stuck Terminating (kubelet cannot confirm)
t=05:51 Deployment creates replacement on node-b -> Multi-Attach error: volume still attached to node-a
t=11:50 6 min without successful deletion and node unhealthy -> attach/detach controller force-detaches
t=12:20 volume attached to node-b, mounted, container Running
Non-graceful shutdown with the out-of-service taint:
t=00:00 node-a powers off
t=00:50 node NotReady; on-call confirms the VM is really off (not rebooting)
t=03:00 kubectl taint node node-a node.kubernetes.io/out-of-service=nodeshutdown:NoExecute
t=03:01 pods on node-a force-deleted, volume detach starts immediately
t=03:31 replacement attached and Running on node-b
downtime: default ~12:20, with out-of-service taint ~03:31
StatefulSet note: the controller will not create web-1 again while the old web-1 object exists,
so without the taint (or a manual force delete) the replacement never starts at all.ExpectedDefault timers (Deployment, RWO volume): t=00:00 node-a powers off; kubelet stops renewing its Lease t=00:50 node controller marks node-a NotReady/Unknown, adds unreachable:NoExecute taint t=05:50 pod tolerationSeconds expires -> pod deleted, stuck Terminating (kubelet cannot confirm) t=05:51 Deployment creates replacement on node-b -> Multi-Attach error: volume still attached to node-a t=11:50 6 min without successful deletion and node unhealthy -> attach/detach controller force-detaches t=12:20 volume attached to node-b, mounted, container Running Non-graceful shutdown with the out-of-service taint: t=00:00 node-a powers off t=00:50 node NotReady; on-call confirms the VM is really off (not rebooting) t=03:00 kubectl taint node node-a node.kubernetes.io/out-of-service=nodeshutdown:NoExecute t=03:01 pods on node-a force-deleted, volume detach starts immediately t=03:31 replacement attached and Running on node-b downtime: default ~12:20, with out-of-service taint ~03:31 StatefulSet note: the controller will not create web-1 again while the old web-1 object exists, so without the taint (or a manual force delete) the replacement never starts at all.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The 6-minute forced detach is documented behavior, and it can be disabled with the disable-force-detach-on-timeout setting in kube-controller-manager. After that, an unhealthy node's volumes are only released through the non-graceful shutdown procedure. Only add the node.kubernetes.io/out-of-service taint after verifying the node is really off. If a "dead" node is actually still writing, two writers on one disk can corrupt the filesystem.
Decision chart: choosing storage for a workload.
Decisions
- 1
A
- nextemptyDir or a generic ephemeral volume
- nextStep 2: do several pods on different nodes write the same files?
- 2
emptyDir or a generic ephemeral volume
- ?
Step 2: do several pods on different nodes write the same files?
- nextRWX file storage (NFS, CephFS, cloud file service), or redesign around object storage
- nextStep 3: does each replica need its own disk and a stable name?
- 4
RWX file storage (NFS, CephFS, cloud file service), or redesign around object storage
- ?
Step 3: does each replica need its own disk and a stable name?
- nextStatefulSet with volumeClaimTemplates, RWO or RWOP, WaitForFirstConsumer
- nextDeployment with replicas 1, strategy Recreate, RWO PVC
- 6
StatefulSet with volumeClaimTemplates, RWO or RWOP, WaitForFirstConsumer
- nextStep 4: is it a database you must fail over?
- 7
Deployment with replicas 1, strategy Recreate, RWO PVC
- ?
Step 4: is it a database you must fail over?
- nextUse an operator or a managed database; StatefulSet alone is not HA
- nextRetain reclaim policy plus scheduled snapshots
- 9
Use an operator or a managed database; StatefulSet alone is not HA
- 10
Retain reclaim policy plus scheduled snapshots
What happens if you choose an alternative
| Choice | Instead of | Consequence |
|---|---|---|
| Immediate binding with zonal disks | WaitForFirstConsumer | Disks created in zones the pod cannot use; Pending pods |
| Deployment with RollingUpdate on an RWO PVC | Recreate or a StatefulSet | New pod cannot attach while the old one holds the disk; rollout stalls |
| RWX network file storage for a database | RWO block storage | Higher latency and locking semantics many databases do not support |
| Running the database in-cluster | Managed database service | You own backups, failover, upgrades and storage failure modes |
| Delete reclaim policy everywhere | Retain for critical classes | One mistaken kubectl delete pvc destroys data |
| hostPath for app data | PVCs | Data tied to one node, no scheduling awareness, security exposure |
Pitfalls
- Editing a PV's capacity by hand can stop automatic resizing: the controller concludes nothing needs to change.
- Two default StorageClasses: Kubernetes uses the newest default; surprising after a cluster add-on installs its own.
- Spreading StatefulSet pods across zones with zonal disks works only if each zone keeps capacity; a full zone strands its replica.
- Assuming RWO means one pod. Two pods on the same node can both mount an RWO volume; use RWOP when you need one pod.
- Forgetting about PVCs after scale-down: retained PVCs keep costing money until you delete them.
Interview Q&A
Explain PV, PVC and StorageClass.
Answer
A PVC is a namespaced request for storage (size, access mode, class). A PV is a cluster-scoped piece of real storage. A StorageClass tells a provisioner how to create PVs dynamically: driver, parameters, reclaim policy, binding mode and whether expansion is allowed. A controller binds each PVC to exactly one PV.
Why use WaitForFirstConsumer?
Answer
With zonal or node-local storage, provisioning before scheduling can put the volume where the pod cannot run. Waiting lets the scheduler choose a node first, considering all the pod's constraints, and then the volume is created in that topology.
What causes a Multi-Attach error?
Answer
An RWO volume is still attached to another node, usually because the old pod's node died or is slow to unmount, so the detach has not happened. Kubernetes waits for safe detach; with defaults it force-detaches after six minutes of failed deletion on an unhealthy node, or immediately once you apply the out-of-service taint to a node you have confirmed is shut down.
How do StatefulSets differ from Deployments for storage?
Answer
StatefulSets give each replica a stable ordinal name, DNS identity and its own PVC from volumeClaimTemplates that follows the replica across reschedules, with ordered creation and updates. Deployments' pods are interchangeable and share whatever PVC the template references.
Walk through the CSI calls when a pod with a new PVC starts.
Answer
CreateVolume (provisioner), ControllerPublishVolume to attach to the chosen node (attacher, via VolumeAttachment), then on the node NodeStageVolume to format and mount to a staging path and NodePublishVolume to bind-mount into the pod.
How do you grow a PVC?
Answer
Set allowVolumeExpansion on the StorageClass, then raise the PVC's requested size; the existing volume grows and the filesystem is resized online if supported. You cannot shrink it.
Why does a StatefulSet replica never come back after its node dies?
Answer
The old pod object is stuck Terminating because the kubelet cannot confirm it stopped, and the controller will not create a second pod with the same identity. Apply the out-of-service taint after confirming the node is off, or force delete the pod.
What does persistentVolumeClaimRetentionPolicy control?
Answer
Whether StatefulSet PVCs are kept or deleted when the StatefulSet is deleted or scaled down. Both default to Retain.
Check yourself
In a test cluster, create a PVC with an Immediate zonal StorageClass and a pod pinned to another zone. Read the scheduler event, then repeat with WaitForFirstConsumer.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Backups That Actually Restore - Snapshots, PITR, Immutable Copies & Restore Drills, Disaster Recovery & Multi-Region - RTO/RPO, Backups, Pilot Light to Active-Active, Multi-AZ vs Multi-Region - Blast Radius, Cell Architecture & Static Stability, Failover & Split Brain — Detection, Promotion, Fencing & Lost Writes, App Secret Injection — Env, Sidecar/Agent, CSI Drivers & Runtime Fetch, Secret Stores Compared — Vault, Cloud Secrets Manager & Kubernetes Secrets, Zero-Downtime Database Migrations — Expand/Contract, Dual Write & Online DDL.