Language Internals
Part 4 of 6 · Python Language ProficiencyGIL, Threading vs Multiprocessing vs asyncio — When Each Wins
What the GIL does and does not protect, and when threads, processes, or asyncio win.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does the GIL serialize?
Answer
Python bytecode in one process. It does not serialize your own locks, and it does not cover code that released it.
L2
Why can threads still speed up downloads?
Answer
Blocking I/O releases the GIL. Other threads run Python while one thread waits in the kernel.
L3
Why can CPU-bound threads be slower than one thread?
Answer
They take turns on the same interpreter and pay for switching. Wall time goes up without extra bytecode throughput.
L4
Why can NumPy use more than one core from threads?
Answer
The heavy work runs in C or Fortran that releases the GIL for the duration of the compute.
L5
When is a process pool slower than staying in-process?
Answer
When the task is short or the arguments are huge. Startup and pickle dominate the arithmetic.
L6
What extra rule does a process entry point need?
Answer
Guard it with if __name__ == '__main__' so spawn on Windows and macOS does not re-import and re-launch forever.
L7
Does a free-threaded build change the interview answer?
Answer
Mention PEP 703 as an experimental direction. Start from the classic GIL unless the role runs a nogil build.
Failure modes
GIL treated as an application lock
Two threads update a dict and lose updates. The interpreter lock does not make compound operations atomic.
Process pool for tiny tasks
Pickle and startup cost more than the work. Batch or stay in process.
fork plus threads
Forking a multithreaded process is unsafe. Prefer the spawn start method in complex apps.
CPU loop on the asyncio thread
The loop cannot poll sockets while pure Python holds the GIL.
Misconceptions
The GIL makes Python single-threaded.
Many threads exist. Only one runs bytecode at a time. Waiting threads are still a useful pool.
Multiprocessing is always the CPU answer.
Measure. Short tasks and fat objects lose to pickle.
asyncio is a third way to get CPU parallelism.
It is one thread. Pair it with a process pool when the transform is pure Python.
Interviewer traps
Explaining the GIL as a distributed lock or a database lock.
One process, one interpreter mutex. Then name threads, processes, or the loop.
Claiming threads cannot help any CPU work.
Native code that releases the GIL is the exception. Pure Python is not.
Re-teaching the JavaScript event loop as if it had a GIL.
JS has one call stack per realm and no GIL. Say that in a sentence and return to CPython.
Design scenario
Same prompt for every reader.
Requirements
Sockets stay on one loop. The transform runs off that thread in another interpreter. Shared dicts on the API side stay locked or immutable.
Traffic / scale
Hundreds of concurrent requests, each with a transform of tens of milliseconds or more.
Latency
The loop stays responsive while a transform runs.
Consistency
A request reads a stable snapshot of config, not a dict another thread is editing.
Availability
A worker crash fails that job. It does not take the parent interpreter with it.
Failure assumptions
- The transform is scheduled with asyncio.to_thread and is pure Python.
- Workers are forked after threads already exist.
- Arguments are multi-megabyte objects pickled per call.
Constraints
- Choose the tool from the table. Cancellation details stay on the asyncio page.
- Name pickle cost and the spawn guard.
Prompt
An API awaits many HTTP calls and then runs a pure-Python feature transform that pegs one core.
Overview
The GIL is a mutex around CPython's interpreter state. One thread runs bytecode. That shapes the whole menu: asyncio for many waits, threads for sync I/O and for C that releases the lock, processes for CPU-bound Python.
Free-threaded builds (PEP 703) are a real experiment. The interview answer still starts from the classic GIL unless the job says otherwise.
JavaScript does not have this lock. The event-loop page and the JS/TS hub describe one call stack and two queues. This page is the CPython choice.
Who holds the lock
Flow
- 1
Thread holds GIL
- nextRuns Python bytecode
- 2
Runs Python bytecode
- blocking I/OReleases GIL
- 3
Releases GIL
- nextAnother thread acquires
- 4
Another thread acquires
- nextThread holds GIL
Lesson map
GIL, Threading vs Multiprocessing vs asyncio — When Each Wins
What the GIL does and does not protect, and when threads, processes, or asyncio win.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB hold["Thread holds GIL"] byte["Runs Python bytecode"] rel["Releases GIL"] next["Another thread acquires"] hold -->|Thread holds GIL to Runs Python bytecode| byte byte -->|blocking I/O| rel rel -->|Releases GIL to Another thread acquires| next next -->|Another thread acquires| hold
- It is a mutex around interpreter internals.
- It is not your application lock. Compound updates of a list or dict still race.
- Blocking I/O typically releases it, so other Python threads run.
- Pure-Python CPU holds it. Extra threads mostly add switching.
Decision table
| Workload | Prefer | Why |
|---|---|---|
| Many idle sockets or HTTP fan-out | asyncio | Cheap waits, explicit awaits |
| Sync SDK with no async port | threading | I/O releases the GIL |
| Pure-Python CPU (encode, parse loops) | process pool | Separate interpreters |
| NumPy or similar native compute | threading often | C releases the GIL |
| API plus a CPU transform | asyncio plus a process pool | I/O on the loop, CPU elsewhere |
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def cpu_bound(n):
total = 0
for i in range(n):
total += i * i
return total
def io_bound_sim(url):
return "fetched:" + url
def main():
with ProcessPoolExecutor() as cpu, ThreadPoolExecutor(max_workers=32) as io:
fut_cpu = cpu.submit(cpu_bound, 200_000)
fut_io = [io.submit(io_bound_sim, f"https://ex/{i}") for i in range(20)]
return fut_cpu.result(), [f.result() for f in fut_io]Spawn on Windows and macOS re-imports the module. Guard the entry point with if __name__ == "__main__".
Press Run. Snippets must be self-contained — no network, files, or native modules.
Sharing state
| Mechanism | Risk |
|---|---|
| Threads plus a list or dict | Races. Use queue.Queue, a lock, or hand off immutable data |
Process Queue | Pickle cost. Large objects dominate |
Manager proxies | Convenient and slow. Not a hot path |
| asyncio plus threads | Only loop-safe calls from other threads, via call_soon_threadsafe |
Side by side
| asyncio | threading | multiprocessing | |
|---|---|---|---|
| Parallel pure-Python CPU | No | No, the GIL | Yes |
| Concurrent I/O | Excellent | Good | Usually overkill |
| Memory | Shared | Shared | Isolated, copy or pickle |
| Cancellation | Strong | Weak | Join or kill |
| Debugging | Await stacks | Thread dumps | Several processes |
Interview Q&A
Why can NumPy use multiple cores with threads?
Answer
The arithmetic runs in native code that releases the GIL. Python threads can sit in that native section on different cores.
Is a process pool always faster for CPU?
Answer
No. Pickle and process startup can dwarf a short function. Batch inputs and measure.
Can the GIL deadlock by itself?
Answer
Lock-order deadlocks are still yours. The GIL does not order application locks. A common case is joining a thread that waits for a lock you still hold.
What is wrong with to_thread for a pure-Python CPU loop?
Answer
The worker thread still needs the GIL. The loop thread cannot run Python either while that worker holds it. Use a process for that work.
Why does fork get dangerous after threads exist?
Answer
The child copies one thread and the locks other threads held. Prefer spawn in programs that already use threads.
Where do application locks still matter?
Answer
Any shared mutable object: a dict of inflight requests, a counter, a cache. The GIL can switch between bytecodes of a compound update.
What would you say about nogil builds?
Answer
PEP 703 makes the lock optional in experimental builds. Classic CPython is still the default assumption. Ask which runtime the service ships.
How does this relate to JavaScript workers?
Answer
JS worker threads are separate realms, closer to processes than to CPython threads. One sentence, then come back to who holds the GIL.
Pitfalls
- A process pool created in a context that cannot spawn.
forkafter threads.- Expecting asyncio to speed a CPU loop.
- Shipping megabyte objects through a process queue one call at a time.
- Treating the GIL as a substitute for
threading.Lock.
A ticket says "parallelize this Python loop with threads." Ask whether the loop is pure Python or native. Write the tool you would switch to if the profile shows the time inside a Python for body.