Data engineering
Part 4 of 6 · Object StorageMultipart Uploads, Parallelism & Throughput
Large objects should not ride one fragile HTTP PUT. Multipart splits bytes into parts, uploads them in parallel, and publishes one object at complete. Interviews expect the state machine, part size, checksums, and abort hygiene so incomplete uploads do not bill forever.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
One PUT versus parts
Prefer
Multipart for large or flaky transfers
Retry the failed part. Parallelize across the NIC. Complete once, with checksums, and the key appears.
- A crash keeps the upload id and the finished part ETags.
- The last part may be smaller than 5 MiB. The others may not.
- You must abort or the parts keep billing.
Alternative
Single PUT for a small object
One request, and the ETag is often the MD5. A retry sends the entire body again. That is fine under about 100 MB and painful at 40 GB.
- No upload id to resume.
- A timeout wastes the bytes already on the wire.
- Simpler clients should stay here until size or failure rate forces parts.
Overview
A 40 GB checkpoint on one HTTP PUT dies with the connection and starts over. Multipart upload on S3, resumable or compose on Cloud Storage, and staged blocks on Azure Blob split the body, send parts in parallel, then commit one object.
The interview is the state machine, the knobs (part size and concurrency), the checksum, and the hygiene: incomplete uploads are stored and billed until you abort them.
State machine
- CreateMultipartUpload with the key and metadata. The response is an upload id. Nothing is visible at the key yet.
- UploadPart for each numbered part. Parts can run in parallel. Each response carries a part ETag. Persist those.
- CompleteMultipartUpload with the part numbers and ETags, sorted by part number. The object becomes visible. The final ETag is not a whole-file MD5.
- On failure, AbortMultipartUpload if you still have the upload id. Also set a bucket lifecycle rule that aborts incomplete uploads after a few days, because clients crash and forget.
Decisions
- 1
1. CreateMultipartUpload
- next2. Receive an upload id
- 2
2. Receive an upload id
- next3. Upload parts in parallel
- 3
3. Upload parts in parallel
- next4. Every part stored
- ?
4. Every part stored
- yes5. Complete with part ETags
- no7. Abort or lifecycle abort
- 5
5. Complete with part ETags
- next6. Object is visible
- 6
6. Object is visible
- 7
7. Abort or lifecycle abort
Lesson map
Multipart Uploads, Parallelism & Throughput
Large objects should not ride one fragile HTTP PUT. Multipart splits bytes into parts, uploads them in parallel, and publishes one object at complete. Interviews expect the state machine, part size, checksums, and abort hygiene so incomplete uploads do not bill forever.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["1."] u["2. Receive an upload id"] p["3. Upload parts in parallel"] d["4. Every part stored"] c -->|1. to 2. Receive an upload id| u u -->|2. Receive an upload id| p p -->|3. Upload parts in parallel| d
S3 limits to remember in one breath: part numbers run from 1 to 10,000; each part except the last is at least 5 MiB; the object can be up to 5 TiB. A too-small non-tail part fails. List the parts you want, in ascending part number, at complete. Parts you never complete stay billed until you abort or a lifecycle rule does. A successful complete assembles the parts you listed and ends that upload.
Upload strategies
| Approach | When | Strength | Cost |
|---|---|---|---|
| Single PUT | Under about 100 MB | Simple; ETag often equals body MD5 | Retry resends the body |
| Multipart, parallel | Large files or a WAN | Resume parts; throughput | ETag is not the MD5; more states |
| Resumable session | Flaky clients on Cloud Storage | The session remembers progress | You manage the session |
| Stage blocks, then commit | Azure block blobs | Commit chooses the block list | Different names, same idea |
| Compose or copy-concat | Merge objects server-side | Avoids pulling bytes back | Extra operations and quotas |
Say the idea, then the product word. Do not pretend CreateMultipartUpload is the Azure API.
Knobs that move throughput
- Part size. Larger parts mean fewer requests and less per-request overhead. Smaller parts mean more parallelism and cheaper retries. Many tools pick 8 to 64 MiB. Stay at or above 5 MiB except for the tail.
- Concurrency. Match the NIC and memory. Too many in-flight parts throttle or exhaust the client. A worker pool with a fixed width beats unbounded tasks.
- Checksums. Prefer CRC32C or SHA-256 per part and for the object when the API allows it. Do not treat the final ETag as a full-file MD5. It is often the MD5 of the concatenated part MD5s, with a hyphen and the part count, like an ETag ending in
-8. - Path to the endpoint. Same region, or a transfer-acceleration endpoint when the client is far away. Egress and RTT dominate before your part-size tweak matters.
- Prefix spread. Many writers on one prefix can throttle. That is prefix design on the data model, not a reason to shrink parts.
Decisions
- 1
1. Split a 40GB file
- next2. Worker pool of size W
- 2
2. Worker pool of size W
- next3. UploadPart with retries
- 3
3. UploadPart with retries
- next4. All parts accepted
- ?
4. All parts accepted
- yes5. Complete and verify checksum
- no6. Retry only failed parts
- 5
5. Complete and verify checksum
- 6
6. Retry only failed parts
- next5. Complete and verify checksum
After the crash
The upload id is the resume token. Losing it means abort if you can, or wait for the lifecycle rule.
- 1
Persist upload id and part ETags
A local file or a metadata row. Memory-only state dies with the process. - 2
Retry missing parts
Completed parts stay. Do not resend the whole 40 GB. - 3
Complete in part-number order
The complete call lists number and ETag. The object appears only after this succeeds. - 4
Abort the unknown ones
If the upload id is lost or the file changed, abort. A lifecycle rule is the backstop.
Incomplete uploads are a cost leak
Parts you never complete are still stored. Lakes discover this as a bill, not as an error.
- Set a lifecycle rule to abort incomplete multipart uploads after a small number of days. The lifecycle lesson shows where that rule sits next to tiering.
- Abort explicitly when the client still has the upload id.
- Alarm on incomplete multipart count and bytes. A quiet metric is the leak.
A complete that succeeds and a metadata pointer that fails is an orphan object, not an incomplete upload. That object is visible. Reconcile the pointer. Do not confuse it with parts that never completed.
Plan the parts
The planner below refuses a non-tail part under 5 MiB when the object is larger than one part. The last slice may be short. The sum of sizes equals the file.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
S3 wants the complete list in ascending part number. Clients that finish part 3 before part 1 still send 1, 2, 3. The TypeScript sort is that requirement with the network removed.
Interview Q&A
When is multipart mandatory?
Answer
Not by a law at small sizes. It becomes the practical requirement for multi-gigabyte objects, for flaky networks, and for SDK thresholds that switch automatically. Below that, a single PUT is easier to reason about.
Why does the final ETag look like a hash with a hyphen and a count?
Answer
Multipart ETags often hash the part hashes and append the part count. They will not match a local MD5 of the full file. Compare a checksum header you sent, not that ETag, when you need content integrity.
How do you resume after a crash?
Answer
Persist the upload id and the completed part ETags. Upload only the missing parts. Then complete. If the upload id or the bytes are unknown, abort and start over. A lifecycle abort covers the case where nobody still holds the id.
When does the object become visible?
Answer
At successful complete, not at create, and not as parts land. GET and LIST before complete do not return the final object. That is the consistency rule applied to this state machine.
What is the 5 MiB rule?
Answer
Every part except the last must be at least 5 MiB on S3. The last part can be smaller. Planning a 1 MiB part size for a 100 MiB file fails. The tail of a file that is not divisible by the part size is the legal short part.
What happens to parts you uploaded but never completed?
Answer
They are not the object, and they are billed until AbortMultipartUpload or a lifecycle abort. A successful complete assembles the parts you listed and ends that upload. Losing the upload id is why the lifecycle rule exists.
How do you pick concurrency?
Answer
Start from bandwidth and memory. Each in-flight part holds a buffer. Past the NIC or the request-rate limit, more workers add throttling, not throughput. Bound the pool.
GCS resumable and Azure block blobs in one line?
Answer
Same shape, different words. A session or a staged block stands in for the part. A compose or a commit block list stands in for complete. Do not mix the verbs in one API call.
Where do checksums sit relative to ETag?
Answer
Send CRC32C or SHA-256 if the API offers it, per part and for the full object. Use the ETag as the part token you must echo at complete. Do not use the multipart ETag as the file's MD5.
What metric tells you the leak?
Answer
Incomplete multipart upload count and bytes, by bucket. A lifecycle abort rule that never fires, or a client that never calls abort, shows up there before anyone files a ticket.
Pitfalls
Write create, eight parts, a timeout on part 5, a retry of part 5 only, then complete. Mark the moment the key becomes GET-able. Then cross out the upload id and say which lifecycle action deletes the parts.