test: harden test_vmm_allocate_zero_size against parallel-run flakes - #3010
Merged
Merged
Conversation
test_vmm_allocate_zero_size asserts that grown.close() waits on the recorded deallocation stream (done.is_done after a 200ms nanosleep). It ran as PARALLEL under the free-threaded (py3.14t) job and failed on aarch64 / A100 / CUDA 13.0.2 with assert done.is_done == False, while its sibling test_vmm_close_synchronizes_recorded_streams (which is marked thread_unsafe and wraps the close in assert_no_cuda_warning) passed. The close path never raises: if cuStreamSynchronize is skipped (capture) or fails, sync_recorded_stream reports a CUDAWarning and unmaps anyway, so the failure surfaced as a bare event-check failure with no clue why. Match the sibling test: - Mark thread_unsafe so the free-threaded job runs it serially, matching the other close-synchronizes tests that record process-global warnings. - Wrap grown.close() in assert_no_cuda_warning() so a skipped/failed sync surfaces as the warning text instead of a bare assert False. If the test still fails when run serially on that job, it is a real library bug in the VMM sync path rather than a parallel-execution artifact.
Contributor
Contributor
Author
|
/ok to test bb7da36 |
Contributor
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
test_vmm_allocate_zero_size(cuda_core/tests/test_memory.py) asserts that closing a grown zero-size VMM buffer waits on the stream the empty buffer recorded: after queueing a 200 ms nanosleep kernel ons, recordingdone, and callinggrown.close(), it assertsdone.is_done.The close path never raises. If
cuStreamSynchronizeis skipped (because the stream is capturing) or fails,sync_recorded_streamincuda_core/cuda/core/_cpp/rt/virtual_memory.cppreports aCUDAWarningand unmaps the range anyway. So a skipped or failed sync surfaces as a bareassert done.is_done == Falsewith no indication of why.On the free-threaded (py3.14t) CI job this test ran as
PARALLEL(pytest-run-parallel) and failed onlinux-aarch64 / Python 3.14t / CUDA 13.0.2 / GPU a100:Its sibling
test_vmm_close_synchronizes_recorded_streams(test_memory.py:1824), which asserts the same close-waits-on-stream behavior, is marked@pytest.mark.thread_unsafe(reason="records process-global warnings")and wraps the closes inassert_no_cuda_warning(); it passed.test_vmm_allocate_zero_sizehad neither, so it was the only test of this pattern that ran in parallel.This PR makes the test match its sibling:
@pytest.mark.thread_unsafe(reason="records process-global warnings")so the free-threaded job runs it serially, matching the other close-synchronizes tests that record process-global warnings.grown.close()inwith assert_no_cuda_warning():so a skipped or failed sync surfaces as theCUDAWarningtext (which names the cause) instead of a bare event-check failure.The change is purely diagnostic and ordering. If the test still fails when run serially on that aarch64 job, it is a real library bug in the VMM sync path under concurrency rather than a parallel-execution artifact, and the warning text will now point at it. If it passes, the failure was a parallel-execution artifact and
thread_unsafeis the correct fix.Verified locally on RTX 6000 Ada / py3.12 / CUDA 13.4.2: 3/3 passes with
--count=3.Checklist