Data engineering
Part 1 of 6 · Object StorageObject Storage — Buckets, Keys, Consistency & Scale
Object storage is the default durable store for lakes, backups, media, and model artifacts: a flat bucket plus key over HTTP, not a POSIX mount. This hub maps block versus file versus object, then consistency, multipart, lifecycle, and the IAM threat model. Metadata partition keys and change streams stay on the sharding and CDC hubs.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Where the bytes should live
Prefer
Object store for blobs, lakes, and artifacts
A bucket and a key, whole-object replace, and lifecycle tiers. Scale is aggregate throughput, not a mounted disk.
- HTTP PUT, GET, DELETE, and LIST. No POSIX rename.
- Multipart when the object is large or the network is flaky.
- Private bucket, short-lived presigns, CDN with origin access control.
Alternative
Block volume or a file share for everything
Right for a database or a multi-writer POSIX app. Wrong as the landing zone for 40 GB checkpoints and parquet parts.
- Block gives random IOPS and in-place pages. You provision the volume.
- A file share gives paths and locks. It does not give cheap cold tiers.
- Mounting an object bucket as a filesystem hides the API and the failure modes.
Overview
Object storage (S3, Cloud Storage, Azure Blob) is the default durable store for data lakes, backups, media, and model artifacts. The namespace is flat: a bucket plus a key. The API is HTTP. You do not mount it and seek like a disk.
Interviewers probe four things. Can you choose block versus file versus object from the workload? Do you understand keys, versioning, ETags, and listing? Can you design multipart, lifecycle, and CDN or presigned flows? Do you treat IAM, encryption, and public access as a threat model?
The prompt to answer out loud: if a client uploads a 40 GB checkpoint and the network flaps, how do you finish safely and prove integrity? Multipart, per-part retries, a checksum, then a metadata pointer. That path is the rest of this cluster.
This page does not re-teach database page layout. A write-ahead log on a block volume is a different durability story from an object that appears only after a successful PUT or a completed multipart upload. It also does not re-teach cache invalidation. Edge cache in front of a bucket is the lifecycle lesson.
Two neighbors stay one sentence each. A metadata database that points at object keys is a shard-key problem: Database Sharding & Partitioning. Landing change events as lake objects is a capture problem: Change Data Capture. Do not fold those lessons in here.
What this cluster covers
- Data Model — Buckets, Objects, Keys, Versioning & Metadata — flat keys, prefixes, ETags, delete markers.
- Consistency — Read-after-write, Listing & Conditional Writes — modern strong reads, and the layers that are still stale.
- Multipart Uploads, Parallelism & Throughput — parts, complete, abort, checksums.
- Lifecycle, Storage Classes, CDN & Presigned URLs — tiers, origin access control, short TTL.
- Security — IAM, Encryption, Public Buckets & Threat Model — roles, block public access, key loss.
Flow
- 1
1. Choose block file or object
- next2. Design bucket and key
- 2
2. Design bucket and key
- next3. PUT or multipart plus checksum
- 3
3. PUT or multipart plus checksum
- next4. Commit a metadata pointer
- 4
4. Commit a metadata pointer
- next5. Lifecycle tier and private access
- 5
5. Lifecycle tier and private access
- next6. IAM encryption and audit
- 6
6. IAM encryption and audit
Lesson map
Object Storage — Buckets, Keys, Consistency & Scale
Object storage is the default durable store for lakes, backups, media, and model artifacts: a flat bucket plus key over HTTP, not a POSIX mount. This hub maps block versus file versus object, then consistency, multipart, lifecycle, and the IAM threat model. Metadata partition keys and change streams stay on the sharding and CDC hubs.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Choose block file or object"] b["2. Design bucket and key"] c["3. PUT or multipart plus checksum"] d["4. Commit a metadata pointer"] a -->|1. Choose block file or object| b b -->|2. Design bucket and key| c c -->|3. PUT or multipart plus checksum| d
Block versus file versus object
| Dimension | Block storage | File storage | Object storage |
|---|---|---|---|
| Address | LBA or volume | Path in a tree | Bucket plus key |
| API | SCSI or NVMe | POSIX open, read, write | HTTP PUT, GET, DELETE, LIST |
| Mutability | In-place blocks | In-place bytes | Whole-object replace, plus versions |
| Scale unit | Volume size | Share or tree | Per object and prefix parallelism |
| Consistency | Strong on the volume | Share and lease dependent | Strong read-after-write on modern S3; see the consistency lesson |
| Best for | Databases, VMs, seek-heavy logs | Shared home dirs, lift-and-shift | Lakes, backups, media, artifacts |
| Cost shape | Provisioned IOPS and GB | Capacity plus ops | Storage plus requests plus egress |
Rule: databases and their logs need block (or a managed database). Multi-writer POSIX apps need file. Bulk durable blobs and lake tables need object.
Same model, three clouds
| Concept | Amazon S3 | Cloud Storage | Azure Blob |
|---|---|---|---|
| Container | Bucket | Bucket | Container in a storage account |
| Object id | Key | Object name | Blob name |
| Region | Region, optional replication | Location or dual-region | Region or geo-redundant options |
| IAM | IAM, bucket policy; avoid ACLs | IAM, uniform bucket access | RBAC plus SAS |
| Large upload | Multipart upload | Resumable or compose | Block blobs, stage then commit |
| CDN | CloudFront | Cloud CDN | Azure CDN or Front Door |
Say the shared model first: a container, a name, and HTTP. Then name the auth surface (roles versus SAS) and the upload protocol. Do not memorize every SKU.
Decisions
- 1
Need durable storage
- nextRandom IOPS or POSIX
- ?
Random IOPS or POSIX
- DB or VMBlock volume
- paths or blobsPOSIX multi-writer
- 3
Block volume
- ?
POSIX multi-writer
- yesFile share
- noObject store
- 5
File share
- 6
Object store
- nextObject size
- ?
Object size
- under 100MBSingle PUT plus checksum
- large or flakyMultipart or resumable
- 8
Single PUT plus checksum
- 9
Multipart or resumable
Decisions
- 1
Object is the store
- nextWho calls it
- ?
Who calls it
- private appIAM role private bucket
- browser or partnerPresign with short TTL
- 3
IAM role private bucket
- 4
Presign with short TTL
- nextHot public reads
- ?
Hot public reads
- yesCDN plus OAC
- noLifecycle and versioning
- 6
CDN plus OAC
- nextLifecycle and versioning
- 7
Lifecycle and versioning
Happy path and the two failures
A producer uploads parts, completes with checksums, and only then writes a pointer in the metadata database. Two failures matter more than the happy path.
- Complete succeeded, database commit failed. The object is durable and invisible to readers who trust the pointer. Reconcile the orphan or retry the pointer. Do not upload a second copy under a new key unless the id is new.
- CDN still serves a deleted object. Origin consistency does not purge the edge. Use a version or content hash in the URL, or a short TTL. The lifecycle lesson owns the details.
Decisions
- 1
1. Upload parts
- next2. Complete with checksums
- 2
2. Complete with checksums
- next3. ETag and version id
- 3
3. ETag and version id
- next4. Pointer commit
- ?
4. Pointer commit
- yes5. CDN fetches on miss
- noReconcile the orphan object
- 5
5. CDN fetches on miss
- next6. Edge still cached
- 6
Reconcile the orphan object
- ?
6. Edge still cached
- staleVersion the URL or purge
- freshServe current bytes
- 8
Version the URL or purge
- 9
Serve current bytes
Default order
Pick the store from the access pattern. Then make the object visible only after the bytes and the pointer agree.
- 1
Name the workload
Random IOPS, POSIX multi-writer, or large immutable blobs. That choice is block, file, or object. - 2
Design the key
Tenant, dataset, and day as a prefix. A leading slash is a character, not a folder. - 3
Upload with a checksum
Single PUT under about 100 MB. Multipart when the body is large or the network flaps. - 4
Commit the pointer
The metadata row is the catalog. The object store is the bytes. A split is an orphan, not a success. - 5
Close the threat
Private bucket, short presign, lifecycle abort for incomplete uploads, encryption you can still decrypt.
Key layout you can run
Prefixes are a convention for listing and lifecycle filters. They are not inodes. A single prefix such as all/ under extreme PUT rates can still hotspot even when the service auto-partitions. Spread by tenant and day when you measure throttling. The metadata database that stores these keys is a separate design: shard that table on the sharding hub, not by inventing a second object store.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The Python key is a teaching layout, not a metastore. The TypeScript check is the policy you apply before you call a cloud SDK. Neither snippet talks to a bucket.
A lake you can defend
Fifty terabytes a day, about ten thousand concurrent uploads, global reads for hot assets, ninety days hot, then infrequent, then a deep archive for compliance.
- No public bucket. CDN uses origin access control for hot assets.
- Multipart plus checksums. A lifecycle rule aborts incomplete uploads.
- Prefix by tenant, then date, so list and lifecycle filters stay bounded.
- The catalog of pointers lives in a database. Partition that catalog with a real shard key. The object key is not a substitute for one.
- Change events that become parquet parts are produced by CDC, then written here. Capture semantics stay on the CDC hub.
Interview Q&A
Why not put the OLTP database on object storage?
Answer
OLTP wants low-latency random writes, transactions, and usually block semantics. Object storage wants large objects and high aggregate throughput. Page-level mutation is the wrong API.
Are folders real?
Answer
No. Keys are flat. A slash is a character people use as a separator. Listing by prefix is a query, not a directory inode. The data-model lesson is the full version.
Give a one-sentence landing zone.
Answer
Land raw objects under tenant and date prefixes with multipart and checksums, register pointers in a metastore, lifecycle colder tiers, serve hot reads through a private origin and a CDN, and never open a public ACL.
When do you pick block, file, or object?
Answer
Block for database and VM IOPS. File for POSIX multi-writer paths. Object for durable blobs, media, backups, and lake files. If you cannot name the access pattern, you are not ready to name the product.
What is the same across S3, Cloud Storage, and Azure Blob?
Answer
A container, an object name, and an HTTP API. Auth and the large-upload protocol differ: IAM versus SAS, multipart versus resumable versus staged blocks. Interview the model, not the price sheet.
What broke in the old S3 consistency story?
Answer
Before December 2020, a new PUT was read-after-write consistent and overwrites plus listings were the eventual part people memorized. Modern S3 is strongly consistent for those API operations in-region. CDNs, other regions, and caches are not. Depth is the consistency lesson.
When is a single PUT enough?
Answer
Small objects, on the order of 100 MB or less, when a full retry is cheap. Larger bodies and flaky networks want multipart so you retry a part, not the whole file.
What is the public-bucket failure?
Answer
An ACL or policy that allows the world to GetObject or ListBucket. Encryption does not close that hole when the service decrypts for anyone allowed to read. Block public access, then use roles and short presigns.
How is a prefix like a shard key?
Answer
Both are partitioning choices: they decide which requests pile up together and which filters stay cheap. The prefix shapes list and lifecycle on the object store. The shard key shapes the metadata database that points at those keys. Design them separately.
Where does CDC fit?
Answer
CDC decides how a database change becomes an event. This cluster decides how that event, or the file it produces, is stored. Do not recap connectors here.
Pitfalls
On a whiteboard, three columns: block, file, object. For a 40 GB model checkpoint, a Postgres volume, and a shared build directory, put each workload in one column and say the API you would call. Then add the failure: network flap on the checkpoint, and name multipart plus checksum without opening the other two lessons.