The GIL Has Been Enabled to Load Module: Free-Threaded Python's Silent Fallback

Python · Intermediate · 7 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: You installed free-threaded Python 3.14t, but one import quietly turns the GIL back on and your threads stop scaling. Here's how to find and fix the culprit.

What Changed

Python 3.14 (October 2025) is the first release where the free-threaded build is officially supported rather than experimental — PEP 779 moved it to phase two of the PEP 703 plan. You install it as a separate interpreter, python3.14t, and it ships with the GIL compiled as optional.

The part nobody warns you about: the GIL can come back at runtime. The first time you import a C extension that hasn't declared free-threading safety, CPython prints

<frozen importlib._bootstrap>:488: RuntimeWarning: The global interpreter lock (GIL) has been enabled to load module '_brotli', which has not declared that it can run safely without the GIL. To override this behavior and keep the GIL disabled (at your own risk), run with PYTHON_GIL=0 or -Xgil=0.

…and then re-enables the GIL for the entire process. Your import succeeds, your tests pass, your threads run — and you get exactly zero of the parallelism you upgraded for. If your logging config filters warnings, you never even see the line.

The Old Way vs The New Way

Before — you assumed "running python3.14t" meant "no GIL," and measured only wall-clock:

# bench.py
import time
from concurrent.futures import ThreadPoolExecutor
import heavy_lib  # innocuous-looking C extension

def burn(n):
    x = 0
    for _ in range(n):
        x += 1
    return x

start = time.perf_counter()
with ThreadPoolExecutor(max_workers=4) as ex:
    list(ex.map(burn, [10_000_000] * 4))
print(f"{time.perf_counter() - start:.2f}s")
# python3.14t bench.py -> 4.1s. Same as 3.14. Now what?

After — you make the interpreter state part of the assertion, and treat the import warning as a build failure:

# conftest.py / app startup
import sys, warnings

warnings.filterwarnings(
    "error",
    message=r"The global interpreter lock \(GIL\) has been enabled",
    category=RuntimeWarning,
)

import heavy_lib  # now raises, with a traceback naming the offender

assert not sys._is_gil_enabled(), "GIL re-enabled — free-threading benefit lost"

Run it as python3.14t -W error::RuntimeWarning -m pytest in CI and the silent fallback becomes a loud, attributable failure.

Why It Was Added

The fallback exists because removing the GIL doesn't magically make twenty years of C extensions thread-safe. Extension code has always been allowed to mutate module-level statics, use non-atomic refcount tricks, and cache borrowed pointers, precisely because the GIL serialized it. Running that code with the GIL off doesn't produce a clean error; it produces heap corruption, wrong results, and segfaults minutes later in unrelated code.

So CPython requires extensions to opt in via the Py_mod_gil slot. Anything that doesn't opt in is treated as "un-audited," and the interpreter chooses correctness over speed: it turns the GIL back on and tells you. The class of bug this eliminates is the worst kind — nondeterministic memory corruption in a dependency you didn't write, surfacing in production under load.

The cost of that safety is the confusing symptom in this article: a performance regression instead of a crash.

How It Works Underneath

Extension modules using multi-phase init declare slots. A free-threading-ready one includes:

static PyModuleDef_Slot slots[] = {
    {Py_mod_exec, mymod_exec},
    {Py_mod_gil, Py_MOD_GIL_NOT_USED},   // the declaration CPython looks for
    {0, NULL}
};

During import, CPython inspects the loaded module for that slot. Missing (or Py_MOD_GIL_USED) means the import machinery calls into the eval loop to enable the GIL — which requires stopping every running thread, attaching the GIL, and resuming. From that point the process is effectively a regular CPython, with free-threading's extra overhead (biased reference counting, per-object locks, no specializing adaptive interpreter in the same form) and none of the payoff.

flowchart TD
    A["import foo (python3.14t)"] --> B["Load extension .so"]
    B --> C{"PYTHON_GIL / -Xgil set?"}
    C -->|"=0 (override)"| G["GIL stays disabled<br/>no warning, no safety net"]
    C -->|"unset"| D{"Module declares<br/>Py_mod_gil = Py_MOD_GIL_NOT_USED?"}
    D -->|yes| E["GIL stays disabled<br/>sys._is_gil_enabled() -> False"]
    D -->|no| F["Emit RuntimeWarning<br/>stop the world, attach GIL"]
    F --> H["Whole process now GIL-bound<br/>threads serialize again"]
    E --> I["Threads run on multiple cores"]
    G --> I

Two practical consequences fall straight out of the diagram. First, this is process-global: one bad transitive dependency ruins it for every thread. Second, PYTHON_GIL=0 doesn't fix anything — it just deletes the branch that was protecting you, so the un-audited C code now runs concurrently for real.

Build-system hooks that emit the declaration for you:

Should You Adopt It Yet

Honest verdict: yes for CPU-bound workloads with a small, well-known dependency set; not yet for large ML/data stacks.

What you're signing up for:

Wait if you depend on a heavy native stack you don't control, or if your team has never written locking-correct Python. Go if you have a pinned dependency set you can verify, and a CPU-bound hot path that currently pays multiprocessing pickling costs.

Migration Notes

  1. Install both interpreters and run side by side. With uv: uv python install 3.14 3.14t, then uv venv --python 3.14t. Keep the standard build as the fallback deployment target until the parallel one beats it on your real benchmark.

  2. Audit before you commit. One command tells you if you're wasting your time:

python3.14t -W error::RuntimeWarning -c "import your_app" 2>&1 | tail -20

Each failure is one dependency to upgrade, replace, or rebuild.

  1. Grep your own C/Cython/Rust sources for PyModuleDef_Slot and PyModule_Create. Single-phase init (PyModule_Create) has no place to put the slot — those modules must move to multi-phase init before they can declare compatibility. Also grep for Py_BEGIN_ALLOW_THREADS, PyGILState_Ensure, and module-level static mutable state; that's where the real audit work is.

  2. Add the guard permanently, not just during migration. A dependency bump six months from now can reintroduce a non-declaring wheel and quietly halve your throughput:

import sys
if sysconfig.get_config_var("Py_GIL_DISABLED") and sys._is_gil_enabled():
    raise RuntimeError("GIL re-enabled at import time")
  1. Don't reach for PYTHON_GIL=0 to make the warning go away. The only legitimate use is a deliberate experiment on code you have personally audited. In CI or production it converts a visible performance problem into an invisible memory-safety one.

  2. Re-test your Python-level concurrency. Run your test suite under python3.14t with high thread counts and, if you can, ThreadSanitizer-built extensions. Races that the GIL hid for years surface here first as flaky tests, not crashes.

Key takeaway: On free-threaded Python, assert `sys._is_gil_enabled() is False` in CI and turn that import RuntimeWarning into an error — otherwise a single un-audited C extension silently restores the GIL and your parallel speedup disappears.

Real-world challenge

A data team moved an image-processing worker to python3.14t expecting ~4x on 4 cores. Locally it scales fine; in the Docker image built from the same requirements.txt it doesn't. Logs show only `<frozen importlib._bootstrap>:488: RuntimeWarning: The global interpreter lock (GIL) has been enabled to load module '_some_codec'`, and warnings are filtered to 'ignore' in production config anyway. Diagnose it.

1. Confirm the runtime state, not the interpreter name.

python3.14t -c "import app.worker, sys; print(sys._is_gil_enabled())"
# True  -> something re-enabled it during import

2. Find the exact import that flipped it. Warnings-as-errors gives you a traceback pointing at the import site:

python3.14t -W error::RuntimeWarning -c "import app.worker"

3. Explain the local/Docker split. Locally pip found a cp314t wheel or your cache had one; in the slim image pip fell back to a wheel or sdist built without the Py_mod_gil declaration. Check with pip debug --verbose / the wheel filename actually installed — *-cp314t-*.whl vs *-cp314-*.

4. Fix, in order of preference: upgrade the dependency to a version publishing free-threaded wheels; replace it with a pure-Python or already-audited alternative; or, if you own the extension, add the slot:

static PyModuleDef_Slot slots[] = {
    {Py_mod_gil, Py_MOD_GIL_NOT_USED},
    {0, NULL}
};

5. Prevent regression. Add a startup guard that fails fast instead of silently degrading:

import sys
if hasattr(sys, "_is_gil_enabled") and sys._is_gil_enabled():
    raise RuntimeError("GIL re-enabled at import time; free-threading benefit lost")

Do not paper over it with PYTHON_GIL=0 — that keeps the GIL off while running C code nobody audited for concurrent access.