Repository navigation
Python 3.13.0a6 freethreading on s390x: test.test_io.CBufferedReaderTest.test_constructor crash with Floating point exception #117755
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Apr 11, 2024 - changed the title
[-]3.13.0a6: `test.test_io.CBufferedReaderTest.test_constructor` ends with `Fatal Python error: Floating point exception`[/-][+]3.13.0a6: `test.test_io.CBufferedReaderTest.test_constructor` ends with `Fatal Python error: Floating point exception` on s390x[/+]on Apr 11, 2024 - changed the title
[-]3.13.0a6: `test.test_io.CBufferedReaderTest.test_constructor` ends with `Fatal Python error: Floating point exception` on s390x[/-][+]3.13.0a6: `test.test_io.CBufferedReaderTest.test_constructor` ends with `Fatal Python error: Floating point exception` on s390x with freethreading[/+]on Apr 11, 2024 - added a commit that references this issue
on Apr 12, 2024 The root issue is a division by zero in mimalloc when requested memory is huge: 0x7fffffffffffffff bytes.
It can be reproduced on s390x without test_io:
$ ./python Python 3.13.0a6+ (heads/main-dirty:396b831, Apr 12 2024, 04:17:51) [GCC 8.5.0 20210514 (Red Hat 8.5.0-20)] on linux >>> import sys >>> size = 0x7fffffffffffffff >>> len = (size - sys.getsizeof(b'')) >>> b'x'*len Floating point exception (core dumped) $ uname -r 4.18.0-513.18.1.el8_9.s390xReacted by Miro HrončokI can reproduce the issue with a Python built with:
./configure --disable-gil CFLAGS="-O0 -g -ggdb" time make -j6I disable all compiler optimizations (
-O0) to ease debugging in gdb.gdb logs when the bug occurs:
$ gdb -args ./python bug.py (gdb) run Program received signal SIGFPE, Arithmetic exception. 0x000000000112b508 in mi_page_init (tld=<optimized out>, block_size=<optimized out>, page=0x40000000168, heap=0x1595ce8 <_PyRuntime+295592>) at Objects/mimalloc/page.c:696 696 page->reserved = (uint16_t)(page_size / block_size); (gdb) p page_size $1 = 0 (gdb) p block_size $2 = 0 (gdb) up #1 0x0000000001172d9e in mi_page_fresh_alloc (heap=0x16efaa8 <_PyRuntime+295592>, pq=0x16f0590 <_PyRuntime+298384>, block_size=9223372036854775808, page_alignment=0) at Objects/mimalloc/page.c:295 295 mi_page_init(heap, page, full_block_size, heap->tld); (gdb) p full_block_size $3 = 0 (gdb) p block_size $4 = 9223372036854775808 (gdb) p /x block_size $5 = 0x8000000000000000 (gdb) p pq == 0 $6 = 0 (gdb) p mi_page_queue_is_huge(pq) $7 = true (gdb) p mi_page_block_size(page) $8 = 0 (gdb) p page->xblock_size $9 = 0(gdb) b PyObject_Malloc Breakpoint 1 at 0x11806c6: file Objects/obmalloc.c, line 1288. (gdb) condition 1 size >= 0x7fffffffffffffffPyObject_Malloc():
- mi_find_page()
- mi_large_huge_page_alloc()
- mi_page_fresh_alloc()
- _mi_segment_page_alloc()
- mi_segment_huge_page_alloc(): mi_segment_os_alloc() and _mi_segment_page_start()
mi_segment_huge_page_alloc() is called with size=0x8000000000000000:
- mi_segment_alloc(size) returns 0x40000000000
uint8_t* start = _mi_segment_page_start(segment, page, &psize);withslice->slice_count = 0setspsizeto 0.
The problem is that
slice->slice_countis 0: integer overflow.mi_page_t.slice_counttype isuint32_t, whereas on s390x, we try to set it to 0x8000_00000000 (140737488355328) which doesn't fit:static mi_page_t* mi_segment_span_allocate(mi_segment_t* segment, size_t slice_index, size_t slice_count, mi_segments_tld_t* tld) { ... slice->slice_count = (uint32_t)slice_count; ... return page; }
gdb:
Breakpoint 5, mi_segment_span_allocate (segment=0x40000000000, slice_index=1, slice_count=140737488355328, tld=0x16f1eb0 <_PyRuntime+304816>) at Objects/mimalloc/segment.c:711- added a commit that references this issue
on Apr 12, 2024 On x86-64, we don't reach this bug, since mmap() fails before.
strace on the big mmap() call:
- x86-64, fail with ENOMEM:
mmap(NULL, 9223372036863164416, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_NORESERVE, -1, 0) = -1 ENOMEM - s390x, success:
mmap(NULL, 9223372036858970112, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_NORESERVE, -1, 0) = 0x3ffdcf80000.
On x86-64, gdb traces around mmap():
- _mi_arena_alloc_aligned(size=0x8000000000410000) returns NULL.
- _mi_prim_alloc()
- unix_mmap() with size=0x8000000000800000, try_alignment=0x2000000, protect_flags=PROT_WRITE | PROT_READ, allow_large=true => it fails with errno=ENOMEM and return NULL.
- x86-64, fail with ENOMEM:
- changed the title
[-]3.13.0a6: `test.test_io.CBufferedReaderTest.test_constructor` ends with `Fatal Python error: Floating point exception` on s390x with freethreading[/-][+]Python 3.13.0a6 freethreading: `test.test_io.CBufferedReaderTest.test_constructor` crash with `Floating point exception` on s390x[/+]on Apr 12, 2024 - changed the title
[-]Python 3.13.0a6 freethreading: `test.test_io.CBufferedReaderTest.test_constructor` crash with `Floating point exception` on s390x[/-][+]Python 3.13.0a6 freethreading on s390x: `test.test_io.CBufferedReaderTest.test_constructor` crash with `Floating point exception`[/+]on Apr 12, 2024 One thing I'm curious about: even with overcommit, I think that
mmap()still needs to create the page mappings. With 4KB pages, that's trillions of pages. The page table itself is too bit to store in memory.Why does it succeed on s390x? Is mmap able to use much bigger pages on IBM Z? Something else?
Why does it succeed on s390x? Is mmap able to use much bigger pages on IBM Z? Something else?
Sorry, I have no idea 🤷🏻♂️
- added 5 commits that reference this issue
on Apr 15, 2024 - added 3 commits that reference this issue
on Apr 17, 2024
Bug report
Bug description:
Since #114331 was solved, we once again attempted to build Python with freethreading on s390x Fedora Linux.
test.test_io.CBufferedReaderTest.test_constructorfails. The traceback:Is this freethreading related? I don't know. Hoping to raise visibility and pointers as where the issue comes from. cc @vstinner
CPython versions tested on:
3.13
Operating systems tested on:
Linux
Linked PRs