The GIL Has Been Enabled to Load Module: Free-Threaded Python's Silent Fallback
Python · Intermediate · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: You installed free-threaded Python 3.14t, but one import quietly turns the GIL back on and your threads stop scaling. Here's how to find and fix the culprit.
What Changed
Python 3.14 (October 2025) is the first release where the free-threaded build is officially supported rather than experimental — PEP 779 moved it to phase two of the PEP 703 plan. You install it as a separate interpreter, python3.14t, and it ships with the GIL compiled as optional.
The part nobody warns you about: the GIL can come back at runtime. The first time you import a C extension that hasn't declared free-threading safety, CPython prints
<frozen importlib._bootstrap>:488: RuntimeWarning: The global interpreter lock (GIL) has been enabled to load module '_brotli', which has not declared that it can run safely without the GIL. To override this behavior and keep the GIL disabled (at your own risk), run with PYTHON_GIL=0 or -Xgil=0.
…and then re-enables the GIL for the entire process. Your import succeeds, your tests pass, your threads run — and you get exactly zero of the parallelism you upgraded for. If your logging config filters warnings, you never even see the line.
The Old Way vs The New Way
Before — you assumed "running python3.14t" meant "no GIL," and measured only wall-clock:
# bench.py
import time
from concurrent.futures import ThreadPoolExecutor
import heavy_lib # innocuous-looking C extension
def burn(n):
x = 0
for _ in range(n):
x += 1
return x
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as ex:
list(ex.map(burn, [10_000_000] * 4))
print(f"{time.perf_counter() - start:.2f}s")
# python3.14t bench.py -> 4.1s. Same as 3.14. Now what?
After — you make the interpreter state part of the assertion, and treat the import warning as a build failure:
# conftest.py / app startup
import sys, warnings
warnings.filterwarnings(
"error",
message=r"The global interpreter lock \(GIL\) has been enabled",
category=RuntimeWarning,
)
import heavy_lib # now raises, with a traceback naming the offender
assert not sys._is_gil_enabled(), "GIL re-enabled — free-threading benefit lost"
Run it as python3.14t -W error::RuntimeWarning -m pytest in CI and the silent fallback becomes a loud, attributable failure.
Why It Was Added
The fallback exists because removing the GIL doesn't magically make twenty years of C extensions thread-safe. Extension code has always been allowed to mutate module-level statics, use non-atomic refcount tricks, and cache borrowed pointers, precisely because the GIL serialized it. Running that code with the GIL off doesn't produce a clean error; it produces heap corruption, wrong results, and segfaults minutes later in unrelated code.
So CPython requires extensions to opt in via the Py_mod_gil slot. Anything that doesn't opt in is treated as "un-audited," and the interpreter chooses correctness over speed: it turns the GIL back on and tells you. The class of bug this eliminates is the worst kind — nondeterministic memory corruption in a dependency you didn't write, surfacing in production under load.
The cost of that safety is the confusing symptom in this article: a performance regression instead of a crash.
How It Works Underneath
Extension modules using multi-phase init declare slots. A free-threading-ready one includes:
static PyModuleDef_Slot slots[] = {
{Py_mod_exec, mymod_exec},
{Py_mod_gil, Py_MOD_GIL_NOT_USED}, // the declaration CPython looks for
{0, NULL}
};
During import, CPython inspects the loaded module for that slot. Missing (or Py_MOD_GIL_USED) means the import machinery calls into the eval loop to enable the GIL — which requires stopping every running thread, attaching the GIL, and resuming. From that point the process is effectively a regular CPython, with free-threading's extra overhead (biased reference counting, per-object locks, no specializing adaptive interpreter in the same form) and none of the payoff.
flowchart TD
A["import foo (python3.14t)"] --> B["Load extension .so"]
B --> C{"PYTHON_GIL / -Xgil set?"}
C -->|"=0 (override)"| G["GIL stays disabled<br/>no warning, no safety net"]
C -->|"unset"| D{"Module declares<br/>Py_mod_gil = Py_MOD_GIL_NOT_USED?"}
D -->|yes| E["GIL stays disabled<br/>sys._is_gil_enabled() -> False"]
D -->|no| F["Emit RuntimeWarning<br/>stop the world, attach GIL"]
F --> H["Whole process now GIL-bound<br/>threads serialize again"]
E --> I["Threads run on multiple cores"]
G --> I
Two practical consequences fall straight out of the diagram. First, this is process-global: one bad transitive dependency ruins it for every thread. Second, PYTHON_GIL=0 doesn't fix anything — it just deletes the branch that was protecting you, so the un-audited C code now runs concurrently for real.
Build-system hooks that emit the declaration for you:
- pybind11 (2.13+):
m.def_submodule(...)aside, passpy::mod_gil_not_used()to the module definition — and note it's compiled conditionally, so the free-threaded headers must actually be in use. - Cython (3.1+):
# cython: freethreading_compatible = True. - Rust/PyO3: expose the module with the free-threaded feature enabled; check the crate version publishes
cp314twheels.
Should You Adopt It Yet
Honest verdict: yes for CPU-bound workloads with a small, well-known dependency set; not yet for large ML/data stacks.
What you're signing up for:
- A second wheel tag (
cp314t) to source for every binary dependency. Coverage is genuinely good now for numpy, cffi, cryptography and friends, but the long tail — codecs, database drivers, vendored SDKs — still ships GIL-using builds. Real-world reports of this exact warning name things liketriton._C.libtriton,onnxruntime,_brotli. - Single-threaded overhead. Free-threaded builds are measurably slower per-thread than the default build; if your workload is I/O-bound or already scales with processes, you may lose.
- Your own Python code becomes genuinely concurrent. Code that was accidentally safe because bytecode ran under the GIL —
dictmutation from multiple threads, check-then-act on shared counters, lazily-initialized singletons — now needs real locks.
Wait if you depend on a heavy native stack you don't control, or if your team has never written locking-correct Python. Go if you have a pinned dependency set you can verify, and a CPU-bound hot path that currently pays multiprocessing pickling costs.
Migration Notes
Install both interpreters and run side by side. With uv:
uv python install 3.14 3.14t, thenuv venv --python 3.14t. Keep the standard build as the fallback deployment target until the parallel one beats it on your real benchmark.Audit before you commit. One command tells you if you're wasting your time:
python3.14t -W error::RuntimeWarning -c "import your_app" 2>&1 | tail -20
Each failure is one dependency to upgrade, replace, or rebuild.
Grep your own C/Cython/Rust sources for
PyModuleDef_SlotandPyModule_Create. Single-phase init (PyModule_Create) has no place to put the slot — those modules must move to multi-phase init before they can declare compatibility. Also grep forPy_BEGIN_ALLOW_THREADS,PyGILState_Ensure, and module-levelstaticmutable state; that's where the real audit work is.Add the guard permanently, not just during migration. A dependency bump six months from now can reintroduce a non-declaring wheel and quietly halve your throughput:
import sys
if sysconfig.get_config_var("Py_GIL_DISABLED") and sys._is_gil_enabled():
raise RuntimeError("GIL re-enabled at import time")
Don't reach for
PYTHON_GIL=0to make the warning go away. The only legitimate use is a deliberate experiment on code you have personally audited. In CI or production it converts a visible performance problem into an invisible memory-safety one.Re-test your Python-level concurrency. Run your test suite under
python3.14twith high thread counts and, if you can, ThreadSanitizer-built extensions. Races that the GIL hid for years surface here first as flaky tests, not crashes.
Key takeaway: On free-threaded Python, assert `sys._is_gil_enabled() is False` in CI and turn that import RuntimeWarning into an error — otherwise a single un-audited C extension silently restores the GIL and your parallel speedup disappears.
Real-world challenge
A data team moved an image-processing worker to python3.14t expecting ~4x on 4 cores. Locally it scales fine; in the Docker image built from the same requirements.txt it doesn't. Logs show only `<frozen importlib._bootstrap>:488: RuntimeWarning: The global interpreter lock (GIL) has been enabled to load module '_some_codec'`, and warnings are filtered to 'ignore' in production config anyway. Diagnose it.
1. Confirm the runtime state, not the interpreter name.
python3.14t -c "import app.worker, sys; print(sys._is_gil_enabled())"
# True -> something re-enabled it during import
2. Find the exact import that flipped it. Warnings-as-errors gives you a traceback pointing at the import site:
python3.14t -W error::RuntimeWarning -c "import app.worker"
3. Explain the local/Docker split. Locally pip found a cp314t wheel or your cache had one; in the slim image pip fell back to a wheel or sdist built without the Py_mod_gil declaration. Check with pip debug --verbose / the wheel filename actually installed — *-cp314t-*.whl vs *-cp314-*.
4. Fix, in order of preference: upgrade the dependency to a version publishing free-threaded wheels; replace it with a pure-Python or already-audited alternative; or, if you own the extension, add the slot:
static PyModuleDef_Slot slots[] = {
{Py_mod_gil, Py_MOD_GIL_NOT_USED},
{0, NULL}
};
5. Prevent regression. Add a startup guard that fails fast instead of silently degrading:
import sys
if hasattr(sys, "_is_gil_enabled") and sys._is_gil_enabled():
raise RuntimeError("GIL re-enabled at import time; free-threading benefit lost")
Do not paper over it with PYTHON_GIL=0 — that keeps the GIL off while running C code nobody audited for concurrent access.