I haven’t written anything in a while. With all the changes AI has brought, I wasn’t sure if it was worth it. But recently, I learned something interesting and wanted to share it.
This blog isn’t meant for experts. I wrote it because I needed to explain these concepts to my wife, so I thought it would be helpful to write them down here too.
So my blog is divided like this:
- What is PEP?
- What is PEP 703?
- What is GIL?
- Why GIL was implemented?
- Notes Async execution models
- Notes on Concurrency, Parallelism, and Event Loop
- What problems does PEP 703 solve?
- Why Now?
- How to test, try and use PEP 703?
What is PEP?
A Python Enhancement Proposal (PEP) is a formal design document that informs the Python community, proposes major new features, or outlines changes to Pythonâs processes or environment. Maintained in a versioned repository, PEPs serve as the primary mechanism for collecting community feedback, documenting technical specifications, and recording the design decisions that govern how the language evolves.
What is PEP 703?
PEP 703 â Making the Global Interpreter Lock Optional in CPython Thatâs it in one line, Now Iâll try to explain this in the way I explained it to my wife: what the GIL is, why we implemented it, the bottleneck it introduced, how we tried to solve it pre-PEP 703 and why we suddenly implemented it
What is GIL?
The Python Global Interpreter Lock, or GIL, in simple terms, is a mutex (or a lock) that allows only one thread to control the Python interpreter.
This means that only one thread can be in a state of execution at any point in time. The impact of the GIL isnât visible to developers who execute single-threaded programs, but it can be a performance bottleneck in CPU-bound and multi-threaded code.
Since the GIL allows only one thread to execute at a time even in a multi-threaded architecture with more than one CPU core, the GIL has gained a reputation as an âinfamousâ feature of Python. Python uses reference counting for memory management. This means objects created in Python have a reference count variable that tracks how many references point to them. When this count reaches zero, the memory occupied by the object is released.
The problem was that this reference count variable needed protection from race conditions where two threads increase or decrease its value simultaneously. If this happens, it can cause either leaked memory that is never released or, even worse, incorrectly release the memory while a reference to that object still exists. This can cause crashes or other âweirdâ bugs in your Python programs.
This reference count variable can be kept safe by adding locks to all data structures that are shared across threads so that they are not modified inconsistently.
But adding a lock to each object or groups of objects means multiple locks will exist, which can cause another problemâDeadlocks (deadlocks can only happen if there is more than one lock). Another side effect would be decreased performance caused by the repeated acquisition and release of locks.
The GIL is a single lock on the interpreter itself, which adds a rule that execution of any Python bytecode requires acquiring the interpreter lock. This prevents deadlocks (as there is only one lock) and doesnât introduce much performance overhead. But it effectively makes any CPU-bound Python program single-threaded.
So the benefits of the GIL are
- Faster single-threaded performance
- Single-threaded programs run faster because the interpreter does not waste time constantly acquiring and releasing individual locks on every single data structure or variable
- Easy C library integration
- It makes it simple to integrate non-thread-safe C and C++ libraries (like NumPy) into Python without risking memory corruption.
- Simpler implementation
- Having one global lock is much easier for the language creators to build and maintain than managing complex, fine-grained locking or lock-free memory systems
- No deadlocks from internal locks
- Because there is only one master lock protecting Pythonâs internal reference counting, it prevents deadlocks that usually happen when juggling multiple locks
Now the problems the GIL introduce
- No True Multi-Core Parallelism
- The biggest disadvantage of the GIL is that multi-threaded Python code cannot run in parallel across multiple CPU cores. Even if you run a Python script on a machine with a 16-core processor, a multi-threaded program will effectively be choked down to utilising just a single CPU core
- Major Bottleneck for CPU-Bound Tasks
- For applications requiring heavy mathematical calculations, image processing, or data analysis, the GIL acts as a massive bottleneck. Instead of speeding up execution, adding more threads to a CPU-bound application often degrades performance because of the overhead generated by threads constantly fighting to acquire and release the lock
- Code Complexity via Workarounds
- Because threading is limited, developers are forced to use alternative concurrency models to bypass the GIL. This usually means using the multiprocessing module instead of threading. However, multiprocessing comes with its own set of downsides
- High Memory Overhead: Each process spawns its own isolated instance of the Python interpreter, multiplying memory consumption.
- Inter-Process Communication (IPC): Sharing data between processes is significantly more complex and slower than sharing memory between threads.
Here is a detailed blog explaining more about the GIL. Do read it out.
How we used to overcome the limitation of GIL?
Bypass or work around Pythonâs Global Interpreter Lock (GIL) by matching your concurrency strategy, such as multiprocessing, asyncio, or threading, to whether your task is CPU-bound or I/O-bound
- Multiprocessing
- Creates separate OS processes, each with its own independent Python interpreter and memory space
- Completely bypasses the GIL because each process has its own lock.
- Best for: CPU-bound tasks (e.g., heavy math, data processing, machine learning) needing true parallel execution across multi-core CPUs.
- Trade-off: High memory overhead and costly Inter-Process Communication (IPC).
- Threading
- Uses standard OS-level threads that share the same memory space.
- Does not bypass the GIL for pure Python bytecode execution. However, the GIL is automatically released when threads wait for blocking I/O operations (like network or disk reads).
- Best for I/O-bound tasks where waiting on external systems allows other threads to run smoothly.
- Trade-off: Useless for parallelising pure Python CPU-bound workloads.
- Asyncio
- Uses a single-threaded event loop to cooperatively switch between tasks using async/await syntax
- Avoids the GIL entirely by staying within a single thread and yielding control during pauses.
- Best for High-concurrency I/O-bound tasks (e.g., handling thousands of simultaneous network requests, chat servers, or API calls)
- Trade-off: Does not speed up CPU-heavy operations.
Note on Concurrency, Parallelism and Event Loop
- Concurrency vs. Parallelism
- While often used interchangeably, concurrency is about structure, whereas parallelism is about execution.
- Concurrency is dealing with multiple things at once. It means managing multiple tasks by switching back and forth between them (interleaving). A single-core processor can achieve concurrency by pausing one task to work on another, giving the illusion of simultaneous progress.
- Parallelism is doing multiple things at once. It requires physical hardware with multiple CPU cores. Tasks literally execute at the same millisecond, independent of one another.
- The Event Loop
- The Event Loop is an architectural pattern used to handle concurrent, non-blocking I/O operations using a single thread.
- It acts as a continuous loop that monitors an execution queue. When an asynchronous task (like a database query or network request) is triggered, the event loop hands the task off to the operating system kernel or a background thread pool and immediately moves to the next task.
- Once the OS finishes the I/O operation, it pushes a callback event back into the event loopâs queue to be processed when the thread is free.
- It eliminates the high memory and CPU overhead of creating thousands of OS threads, allowing a single thread to manage thousands of simultaneous open connections efficiently.
The Problem PEP 703 solves
Even though we have advanced abstractions like multiprocessing, concurrency, and event loops, PEP 703 solves the core structural problems of overhead, memory inefficiency, and data sharing that these workarounds introduced. Before PEP 703, every parallel execution strategy in Python required choosing between two major compromises: high operating system overhead or severe performance bottlenecks.
- Eliminates High Memory Overhead (The Multiprocessing Problem)
- To achieve true multi-core parallelism, developers were forced to use multiprocessing or Gunicorn workers. This meant spinning up entirely separate OS processes, each loading its own copy of the Python interpreter, standard libraries, and application state into memory.
- PEP 703 Solution: It creates a âfree-threadedâ build of Python. Multiple CPU cores can now execute raw Python bytecode simultaneously within a single process. This dramatically drops memory footprints since threads inherently share the same memory space.
- Removes Serialisation Bottlenecks (The Inter-Process Communication Problem)
- Because processes do not share memory, passing data between them (like sending a large NumPy array or a complex machine learning model from a master process to a worker) requires serialisation (pickling). For massive data pipelines, the time spent converting data into bytes and back frequently wiped out the performance gains of parallel processing.
- PEP 703 Solution: Threads can directly read and write to the same objects in memory instantly, without any pickling or serialisation overhead.
- Erases the âGIL-Bouncingâ Delay in Native Extensions (The AI/ML Problem)
- Heavy computation libraries (like NumPy, PyTorch, or C extensions) can release the GIL when executing low-level C++ or CUDA code. However, the moment that background thread completes and needs to pass the result back into a Python object or callback, it has to wait to re-acquire the GIL. In high-throughput AI/ML pipelines, threads frequently bottlenecked simply waiting in line to talk back to Python.
- PEP 703 Solution: Native C extensions and Python threads can transition data back and forth concurrently without ever stalling for a global lock.
- Unifies Async and Parallelism (The Event Loop Problem)
- An event loop (asyncio) is highly efficient for waiting on I/O (network requests), but it runs entirely on a single thread. If an async application suddenly needs to perform a heavy CPU computation (like hashing a password or parsing a massive JSON payload), it blocks the entire event loop, freezing all other concurrent network requests. Resolving this required complex hybrid architectures (e.g., executing the async loop, but offloading math tasks to a ProcessPoolExecutor).
- PEP 703 Solution: You can run multi-threaded processing pools inside your async applications safely without worrying about blocking the main loop or incurring heavy process-spawning penalties.
Why PEP 703 is inevitable now?
PEP 703 became inevitable because single-core processor speeds have stopped growing (the end of Mooreâs Law), and Python has become the universal orchestration layer for massive multi-core AI, machine learning, and data workloads where old workarounds like multiprocessing cause unbearable memory bloat and serialisation delays.
- The Death of Single-Core CPU Scaling
- The AI, ML, and Data Science Boom
- The Collapse of âMultiprocessingâ as a Sustainable Patch
How PEP 703 Replaced the GIL Without Breaking Everything?
Simply deleting the GIL would cause data corruption due to thread race conditions. To make Python thread-safe without the global lock, PEP 703 introduced three primary engineering innovations directly to Pythonâs internals.
Biased Reference Counting (BRC): Instead of using expensive, slow atomic locks every time a variable is accessed, Python checks which thread owns the object. If itâs the creating thread, it uses lightning-fast normal increments. It only switches to atomic locking if another thread accesses it.
Immortal Objects: Global constants that are accessed millions of times (like None, True, False, or static integers) are marked as âimmortal.â Their reference counters never change, preventing massive thread contention over shared constants.
Deferred Reference Counting: For objects that are highly active across multiple threads (like functions and modules), reference count updates are entirely deferred to a separate Garbage Collector cycle to keep thread execution clean and fast.
How to test: try using PEP 703
Testing and using the free-threaded build (PEP 703) is straightforward. The feature was introduced experimentally in Python 3.13 and graduated to officially supported status in Python 3.14, though it remains an optional, opt-in build. Free-threaded installations are designated by a t suffix in their version tags (e.g., 3.13t or 3.14t)
Step 1 Install
# Install Python 3.14 free-threaded variant
uv python install 3.14t
# Or Python 3.13 free-threaded variant
uv python install 3.13t
Step 2: Verify Your Setup
To confirm that you are running the true GIL-free environment, invoke your executable with the -VV flag or run a quick runtime check:
# Check using uv
uv run --python 3.14t python -VV
# Expected output text includes: "free-threading build"
# Or check programmatically via Python script
uv run --python 3.14t python -c "import sys; print(sys._is_gil_enabled())"
# Expected output: False
Step 3: Run code and test your script.
To see PEP 703 in action, write a basic script that spins up pure Python threading tasks on a CPU-heavy problem (e.g., calculations or processing a massive array).
import threading
import time
import sys
def cpu_heavy_task():
# Simple CPU-burning math loop
count = 0
for i in range(10_000_000):
count += i
print(f"GIL enabled: {sys._is_gil_enabled()}")
start_time = time.time()
# Spin up 4 threads to utilise multiple CPU cores
threads = []
for _ in range(4):
t = threading.Thread(target=cpu_heavy_task)
threads.append(t)
t.start()
for t in threads:
t.join()
print(f"Execution Time: {time.time() - start_time:.4f} seconds")
If you execute this code using standard Python 3.14, the 4 threads will stall sequentially against the GIL, yielding an identical or worse execution time than a single-threaded approach.
When run under python3.14t, all 4 threads will light up 4 CPU cores simultaneously, executing significantly faster.
Notes
- Third-Party C-Extensions Might Turn the GIL Back On: If you import an older external C library (like an un-updated data science package) that does not explicitly declare it supports free-threading, Python will automatically re-enable the GIL silently for safety. Keep an eye out for runtime environment warnings.
- Force-Disabling the GIL: If a package throws a warning or attempts to turn the GIL back on, you can explicitly force it off at your own risk using an environment variable or flag:
python3.14t -Xgil=0 my_script.pyorexport PYTHON_GIL=0 - Ecosystem Readiness: You can track which major open-source tools have native, safe multi-threaded support by checking the community-maintained Python Free-Threading Guide Tracker
Further Reading