Data engineering
Part 2 of 6 · Object StorageData Model — Buckets, Objects, Keys, Versioning & Metadata
The object data model is a bucket, bytes plus system metadata, a UTF-8 key, optional user metadata, and optional versions. Interviews fail when prefixes are treated as directories, ETags and version ids are ignored, or one hot prefix throttles PUTs and listings.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Rename is a copy
Prefer
Copy to the new key, then delete the old
The store has no inode rename. Readers that need a stable name should follow a pointer you update after the copy succeeds.
- The new key is a new object until you delete the source.
- Versioning keeps the old generation if you overwrite in place instead.
- A metadata row can flip from the old key to the new key in one commit.
Alternative
Treat the prefix as a directory and rename it
Consoles draw folders. The API does not move a subtree in one call. A loop of copy and delete is the operation, with a partial-failure window.
- A crash mid-loop leaves both keys or neither, depending on where it died.
- Listings during the loop show a mix of old and new prefixes.
- Permissions and lifecycle rules follow the key string, not an inode.
Overview
The model looks small and then surprises people. A bucket (or container) is the namespace and the policy boundary: region, ownership, public-access block, default encryption. An object is bytes plus system metadata, replaced as a whole unless you are using a special append or compose API. A key is the name, unique inside the bucket, up to 1024 bytes of UTF-8 on S3. Optional user metadata and tags hang off the object. Optional versioning keeps prior generations under that same key.
Interviews fail in three places. Candidates call prefixes directories. They ignore ETag and version id when talking about overwrite and delete. They design a keyspace so hot that listings and PUTs throttle.
This is not a POSIX lesson. There is no directory inode, no hard link, and no cheap rename.
Entities
| Entity | What it is | Interview note |
|---|---|---|
| Bucket or container | Namespace and policy boundary | Region, ownership, block public access |
| Key | Unique name in the bucket | Flat; slash is convention only |
| Object | Bytes plus metadata | Whole-object replace, unless an append API |
| Version | Immutable generation of a key | Version id; delete markers |
| ETag | Integrity or change token | Often MD5 for a single PUT; not for multipart |
| Prefix | Leading substring of keys | List filter; the folder illusion |
| Tag | Separate key and value | Lifecycle and IAM conditions |
Flow
- 1
Bucket policy and region
- nextlake/acme/dt/a.parquet
- nextmodels/acme/v3/weights
- 2
lake/acme/dt/a.parquet
- nextbytes ETag version id
- 3
models/acme/v3/weights
- nextbytes checksum storage class
- 4
bytes ETag version id
- 5
bytes checksum storage class
Lesson map
Data Model — Buckets, Objects, Keys, Versioning & Metadata
The object data model is a bucket, bytes plus system metadata, a UTF-8 key, optional user metadata, and optional versions. Interviews fail when prefixes are treated as directories, ETags and version ids are ignored, or one hot prefix throttles PUTs and listings.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB b["Bucket policy and region"] k1["lake/acme/dt/a.parquet"] k2["models/acme/v3/weights"] m1["bytes ETag version id"] b -->|Bucket policy and region| k1 b -->|Bucket policy and region| k2 k1 -->|lake/acme/dt/a.p| m1
Prefix design
A prefix is the leading part of a key you pass to LIST. lake/acme/events/dt=2026-10-01/ returns the objects whose names start that way. The console draws a folder. The service answered a string filter.
| Pattern | Example | Strength | Cost |
|---|---|---|---|
| Date first | dt=2026-10-01/tenant/... | Easy time-based lifecycle | Per-tenant scans fan out |
| Tenant first | tenant/acme/dt=... | Isolation and per-tenant IAM | A hot tenant concentrates load |
| Hash shard | ab/cd/ plus a uuid | Spreads extreme PUT rate | Harder for humans to browse |
| Hive style | table/dt=.../part-*.parquet | Friendly to query engines | Needs compaction discipline |
Modern S3 partitions prefixes for you, so a random hash is not a default requirement the way 2010s rate-limit lore claimed. It is still a knob when you measure 503 Slow Down on one prefix, and it is still how you bound a list or a lifecycle rule. Pick the prefix for operations you run, then spread only if the metrics say so.
That choice rhymes with a database partition key and is not the same design. The object prefix shapes list, lifecycle, and request rate on the blob store. The key that shards the catalog of pointers lives on Database Sharding & Partitioning. Do not invent a storage-engine lesson to explain a string prefix.
Versioning and delete markers
Turn versioning on when an overwrite or a delete must be recoverable. Every PUT creates a new immutable version id. A DELETE without a version id does not erase bytes. It inserts a delete marker that becomes the current version. A GET with no version id then returns 404. A GET with the previous version id still returns those bytes.
Flow
- 1
1. PUT report.pdf creates v1
- next2. PUT same key creates v2
- 2
2. PUT same key creates v2
- next3. DELETE inserts a marker
- 3
3. DELETE inserts a marker
- next4. GET without version id is 404
- 4
4. GET without version id is 404
- next5. GET version id v2 returns bytes
- 5
5. GET version id v2 returns bytes
Rules that belong in the answer:
- Versioning protects accidents and increases storage until you expire noncurrent versions.
- A delete marker is not a byte object. It hides the latest version for the default GET.
- Listing current objects and listing all versions are different calls.
- Object Lock (WORM) and MFA Delete show up in compliance interviews. Name them here. The security lesson is where the threat model sits.
Unversioned DELETE removes the object. There is no previous generation to GET.
Metadata layers
- System metadata. Size, content type, last-modified, storage class, encryption state. The service maintains these.
- User metadata. Custom headers in the
x-amz-meta-style. Size limits apply. You set them on PUT. They come back on HEAD and GET. - Object tags. A separate small set of key-value pairs. Lifecycle rules and IAM conditions often match tags, not user metadata.
- External metastore. Iceberg, Hive, or a database of pointers. That system is the source of truth for tables. Objects remain files.
User metadata is not a database. It is not indexed for rich queries. If you need to find objects by an attribute, put that attribute in a real index or in the key prefix you list.
What to store where
Bytes in the object. Searchable facts in a metastore. Policy facts in tags.
- 1
Bytes and content type
The object body and system metadata. Cache-Control belongs here when a CDN will read it. - 2
Small user metadata
A pipeline version or a checksum the writer wants echoed back. Not a document you will query. - 3
Tags for rules
Lifecycle and IAM match tags. Keep the set small. - 4
Metastore for tables
The catalog says which keys make a dataset. The bucket does not.
Keys you can validate
A leading slash becomes part of the name and surprises every tool that strips it. An empty key, a key that is only a trailing slash, and a doubled slash are how consoles and humans disagree. S3 allows a large character set; the allowlist below is a teaching guard for lake keys, not the full S3 grammar.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
private, max-age=0 is the right default for data that a CDN should not pin. Immutable static assets, in the lifecycle lesson, take a long cache and a content hash in the URL instead.
Interview Q&A
What is the difference between a versioned and an unversioned delete?
Answer
Unversioned DELETE removes the object. Versioned DELETE without a version id adds a delete marker. Prior versions remain until you delete that version id or a lifecycle rule expires noncurrent versions. A default GET then 404s.
Why can an ETag differ from the MD5 of the file?
Answer
A single-part PUT often uses the MD5 as the ETag. A multipart ETag is commonly a hash of the part hashes plus a part count, not the MD5 of the concatenated bytes. Some encryption modes also stop the ETag from being an MD5. When integrity matters, send an explicit checksum such as CRC32C or SHA-256.
How do you rename an object?
Answer
Copy to the new key, then delete the old key, or overwrite under versioning and expire the old version. There is no cheap inode rename. If readers use a stable name, update a pointer after the copy succeeds.
Is a prefix a folder?
Answer
No. It is a leading substring used as a list filter and as a lifecycle or IAM scope. Two keys that share a prefix do not share an inode.
What is a bucket for, beyond holding keys?
Answer
Region, block public access, default encryption, bucket policy, and ownership. It is the blast-radius boundary. A key does not cross buckets without a copy.
Where do Iceberg or Hive fit?
Answer
They are the table catalog. They record which objects belong to a snapshot. The object store still stores bytes. Query planning does not scan user metadata headers.
User metadata versus tags?
Answer
User metadata rides on the object headers and comes back on GET. Tags are a separate pair list that lifecycle and IAM can match. Neither is a secondary index.
What does a delete marker cost?
Answer
It hides the current version and it is its own version. The hidden bytes remain billable until you expire noncurrent versions or delete them by id. Versioning without an expiry rule is a silent storage leak.
Why can one tenant prefix be a hotspot?
Answer
All PUTs and LISTs for that tenant share a prefix. Automatic partitioning helps, and a single noisy tenant can still concentrate requests. Hashing the prefix trades browse-ability for spread. Measure before you hash everything.
What is the 1024-byte rule?
Answer
On S3 the key is at most 1024 bytes of UTF-8, not 1024 characters. A long Unicode name can blow the limit earlier than a quick character count suggests.
Pitfalls
Write one key, report.pdf. PUT twice, then DELETE. Mark which GET returns 404 and which GET returns v2. Then say what still costs money, and which lifecycle action removes it. That action is named properly on the lifecycle page.