Repository navigation
Python 3.13 breaks circular imports through single-phase extension modules using non-isolated subinterpreters #138045
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Aug 22, 2025 - addedextension-modulesC modules in the Modules dirC modules in the Modules dir3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixes3.15bugs and security fixesbugs and security fixes
on Aug 22, 2025 gh-117953 says:
For the main interpreter and non-isolated subinterpreters, nothing different would happen from now; there would be no switching.
However, what was actually implemented in gh-118157 contains this comment:
/* We *could* leave in place a legacy interpreter here * (one that shares obmalloc/GIL with main interp), * but there isn't a big advantage, we anticipate * such interpreters will be increasingly uncommon, * and the code is a bit simpler if we always switch * to the main interpreter. */I think this discrepancy is the crux of the issue.
Thanks for doing some digging. The repro (with
_interpreters.create) creates an isolated subinterpreter, not a legacy one, and we definitely don't support loading single-phase extensions in those. I'll see what we can do to fix the crash (there should be an exception instead), but first, do things work as intended if you switch that to a shared-GIL subinterpreter?- Ah, I wasn't aware that the module used isolated mode. As far as I understand it, the crash happens either way, it's just that in the isolated case it's supposed to cleanly bail later (but it doesn't get that far here). The whole reentrancy thing happens before the interpreter even has a chance to figure out what kind of module this is. I don't have a working environment with Python 3.12 or older right now, but my understanding is all working deployments of the Ceph manager with the diskprediction plugin must be using that (and it works). As far as I know ceph-mgr doesn't set any isolation flags (I grepped the source and didn't find any) so it must be using non isolated mode, which is why it works on older versions. The reentrancy behavior definitely repros the same with the full ceph-mgr workload, the repro was just an attempt to minimize it. - marcanReacted by Peter Bierma
Here's a test that uses the non-isolated subinterpreters:
#include <Python.h> #include <stdio.h> int main() { Py_Initialize(); PyThreadState *tstate = Py_NewInterpreter(); PyObject *numpy = PyImport_ImportModule("numpy"); if (PyErr_Occurred()) { fprintf(stderr, "Error occured\n"); PyErr_Print(); return 1; } return 0; }
On Python 3.13.7 and numpy 2.2.6, it actually manages to bail out instead of crashing (it's still reentering the extension init, numpy just catches it before doing anything that crashes and explodes):
Error occured RuntimeError: CPU dispatcher tracer already initlizedOn py 3.13.5 and numpy 1.26.4 (F41), it segfaults while printing the error (this is what happens in ceph-mgr):
Error occured [1] 32792 segmentation fault (core dumped) ./pytestOn an older machine (Gentoo) with Python 3.11 and numpy 1.26.4, it works with a warning:
# gcc -o pytest -lpython3.11 -I/usr/include/python3.11 pytest.c && ./pytest sys:1: UserWarning: NumPy was imported from a Python sub-interpreter but NumPy does not properly support sub-interpreters. This will likely work for most users but might cause hard to track down issues or subtle bugs. A common user of the rare sub-interpreter feature is wsgi which also allows single-interpreter mode. Improvements in the case of bugs are welcome, but is not on the NumPy roadmap, and full support may require significant effort to achieve.This warning is issued by checking the thread state, so it does not indicate the extension or module was initialized twice (that's a different warning). So I do expect it to work if numpy is only ever used from a single subinterpreter.
- added a commit that references this issue
on Sep 11, 2025 I've tried the following patch which only switches to the main interpreter when the current thread state has its own GIL:
diff --git a/Python/import.c b/Python/import.c index 64048a4ef91..510f01eb03e 100644 --- a/Python/import.c +++ b/Python/import.c @@ -1975,21 +1975,22 @@ import_run_extension(PyThreadState *tstate, PyModInitFunction p0, * and then continue loading like normal. */ bool switched = false; - /* We *could* leave in place a legacy interpreter here - * (one that shares obmalloc/GIL with main interp), - * but there isn't a big advantage, we anticipate - * such interpreters will be increasingly uncommon, - * and the code is a bit simpler if we always switch - * to the main interpreter. */ - PyThreadState *main_tstate = switch_to_main_interpreter(tstate); - if (main_tstate == NULL) { - return NULL; - } - else if (main_tstate != tstate) { - switched = true; - /* In the switched case, we could play it safe - * by getting the main interpreter's import lock here. - * It's unlikely to matter though. */ + PyThreadState *main_tstate = tstate; + /* For non-isolated, legacy interpreters they share the GIL and allocator + * with the main thread anyway, so there's no need to switch to the main + * interpreter. If we would still switch in that case in would result in + * calling the init function of single-phase modules that have cyclic + * references twice, see issue #138045. */ + if (tstate->interp->ceval.own_gil) { + main_tstate = switch_to_main_interpreter(tstate); + if (main_tstate == NULL) { + return NULL; + } else if (main_tstate != tstate) { + switched = true; + /* In the switched case, we could play it safe + * by getting the main interpreter's import lock here. + * It's unlikely to matter though. */ + } } struct _Py_ext_module_loader_result res; @@ -2173,9 +2174,12 @@ clear_singlephase_extension(PyInterpreterState *interp, /* We must use the main interpreter to clean up the cache. * See the note in import_run_extension(). */ PyThreadState *tstate = PyThreadState_GET(); - PyThreadState *main_tstate = switch_to_main_interpreter(tstate); - if (main_tstate == NULL) { - return -1; + PyThreadState *main_tstate = tstate; + if (tstate->interp->ceval.own_gil) { + main_tstate = switch_to_main_interpreter(tstate); + if (main_tstate == NULL) { + return -1; + } } /* Clear the cached module def. */
It fixes your test @marcan , but breaks
test.test_import.SinglephaseInitTests.test_basic_multiple_interpreters_deleted_no_reset. I don't really understand what is being tested there exactly :/Also in a more complex scenario (that worked with Python <= 3.12) importing numpy still crashes with
_PyThreadState_Attach: non-NULL old thread state.Any ideas?
@ericsnowcurrently is my check for the shared GIL correct?
@jan-hasse At a glance, that fix makes sense to me. The
test_basic_multiple_interpreters_deleted_no_resetseems to be a known issue at the moment: #131229I've fixed the test failure: There was a check for the main interpreter in
update_global_state_for_extensionwhich now had to be removed.FWIW, numpy 2.4.0 converted its extensions to multi-phase and this no longer repros on that version.
pip install --user numpy==2.2.6is enough to bring back the repro though.It seems that #144602 fixed this issue.
Using a local debug build of 3.15 on NumPy 2.2.6, I see the following output instead of a crash:
Exception while importing from subinterpreter: Traceback (most recent call last): File "/home/peter/develop/cpython/.venv/lib/python3.15/site-packages/numpy/__init__.py", line 114, in <module> from numpy.__config__ import show_config File "/home/peter/develop/cpython/.venv/lib/python3.15/site-packages/numpy/__config__.py", line 4, in <module> from numpy._core._multiarray_umath import ( ...<3 lines>... ) File "/home/peter/develop/cpython/.venv/lib/python3.15/site-packages/numpy/_core/__init__.py", line 23, in <module> from . import multiarray File "/home/peter/develop/cpython/.venv/lib/python3.15/site-packages/numpy/_core/multiarray.py", line 10, in <module> from . import overrides File "/home/peter/develop/cpython/.venv/lib/python3.15/site-packages/numpy/_core/overrides.py", line 7, in <module> from numpy._core._multiarray_umath import ( add_docstring, _get_implementing_args, _ArrayFunctionDispatcher) RuntimeError: CPU dispatcher tracer already initlizedI think I understand what was going on:
Python reentrantly attempts to initialize the (single-phase!) extension module in the main interpreter again
KaboomThe "Kaboom" is doing a lot of heavy lifting here, so I missed it. NumPy noticed that it was being initialized again, which caused an exception. Prior to my fix in #144602, this would crash the interpreter, because the exception from the subinterpreter would be used in the main interpreter, which isn't allowed.
@jan-hasse's patch avoided the problem by preventing switching when the interpreter shared the GIL, so the exception was owned by the correct interpreter. But, this problem exists for own-GIL subinterpreters as well, so the patch from that PR still fails for
_interpreters.run_string(_interpreters.create(), "import numpy").Anyways, I think we can close this now!
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Bug report
Bug description:
I believe this regression was caused by the changes made to fix gh-117953.
Repro:
Tested on Fedora 41 with Python 3.13.5 and numpy 1.26.4, and Fedora 42 with 3.13.7 and 2.2.6, respectively.
What happens is that:
numpymultiarrayextensionnumpy, now in the main interpretermultiarrayextension again, this time already in the main interpreter, so no switching happensThe easiest way to see the carnage in action is to run it under gdb and break on
PyInit__multiarray_umath. You'll see it gets called twice, the second time reentrantly (with a 276-frame backtrace going through numpy initialization twice). If you don't set any breakpoints, you'll get crashes trying to process exceptions and other craziness when the original reentrant execution backtrace is long gone.numpy does not support subinterpreters, but I would expect it to work when it is used in a single subinterpreter without isolation.
This is the root cause for:
CPython versions tested on:
3.13
Operating systems tested on:
Linux
Linked PRs