Skip to content

test: harden test_vmm_allocate_zero_size against parallel-run flakes - #3010

Merged
leofang merged 1 commit into
NVIDIA:mainfrom
juenglin:fix-vmm-zero-size-test-flake
Oct 3, 2026
Merged

leofang merged 1 commit into
NVIDIA:mainfrom
juenglin:fix-vmm-zero-size-test-flake

Conversation

@juenglin

@juenglin juenglin commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Description

test_vmm_allocate_zero_size (cuda_core/tests/test_memory.py) asserts that closing a grown zero-size VMM buffer waits on the stream the empty buffer recorded: after queueing a 200 ms nanosleep kernel on s, recording done, and calling grown.close(), it asserts done.is_done.

The close path never raises. If cuStreamSynchronize is skipped (because the stream is capturing) or fails, sync_recorded_stream in cuda_core/cuda/core/_cpp/rt/virtual_memory.cpp reports a CUDAWarning and unmaps the range anyway. So a skipped or failed sync surfaces as a bare assert done.is_done == False with no indication of why.

On the free-threaded (py3.14t) CI job this test ran as PARALLEL (pytest-run-parallel) and failed on linux-aarch64 / Python 3.14t / CUDA 13.0.2 / GPU a100:

tests/test_memory.py:1787: AssertionError
>           assert done.is_done
E           assert False

Its sibling test_vmm_close_synchronizes_recorded_streams (test_memory.py:1824), which asserts the same close-waits-on-stream behavior, is marked @pytest.mark.thread_unsafe(reason="records process-global warnings") and wraps the closes in assert_no_cuda_warning(); it passed. test_vmm_allocate_zero_size had neither, so it was the only test of this pattern that ran in parallel.

This PR makes the test match its sibling:

  • Mark @pytest.mark.thread_unsafe(reason="records process-global warnings") so the free-threaded job runs it serially, matching the other close-synchronizes tests that record process-global warnings.
  • Wrap grown.close() in with assert_no_cuda_warning(): so a skipped or failed sync surfaces as the CUDAWarning text (which names the cause) instead of a bare event-check failure.

The change is purely diagnostic and ordering. If the test still fails when run serially on that aarch64 job, it is a real library bug in the VMM sync path under concurrency rather than a parallel-execution artifact, and the warning text will now point at it. If it passes, the failure was a parallel-execution artifact and thread_unsafe is the correct fix.

Verified locally on RTX 6000 Ada / py3.12 / CUDA 13.4.2: 3/3 passes with --count=3.

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

test_vmm_allocate_zero_size asserts that grown.close() waits on the
recorded deallocation stream (done.is_done after a 200ms nanosleep). It
ran as PARALLEL under the free-threaded (py3.14t) job and failed on
aarch64 / A100 / CUDA 13.0.2 with assert done.is_done == False, while its
sibling test_vmm_close_synchronizes_recorded_streams (which is marked
thread_unsafe and wraps the close in assert_no_cuda_warning) passed.

The close path never raises: if cuStreamSynchronize is skipped (capture)
or fails, sync_recorded_stream reports a CUDAWarning and unmaps anyway,
so the failure surfaced as a bare event-check failure with no clue why.

Match the sibling test:
- Mark thread_unsafe so the free-threaded job runs it serially, matching
  the other close-synchronizes tests that record process-global warnings.
- Wrap grown.close() in assert_no_cuda_warning() so a skipped/failed sync
  surfaces as the warning text instead of a bare assert False.

If the test still fails when run serially on that job, it is a real
library bug in the VMM sync path rather than a parallel-execution
artifact.
@juenglin juenglin added this to the cuda.core next milestone Oct 2, 2026
@juenglin juenglin added test Improvements or additions to tests cuda.core Everything related to the cuda.core module labels Oct 2, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@juenglin juenglin self-assigned this Oct 2, 2026
@juenglin
juenglin requested a review from Andy-Jost October 2, 2026 20:58
@juenglin

juenglin commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test bb7da36

@juenglin
juenglin requested a review from seberg October 2, 2026 21:00
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor
Doc Preview CI
Preview removed because the pull request was closed or merged.

@leofang leofang left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, Ralf!

-- Leo's bot

@leofang
leofang merged commit ab9e20f into NVIDIA:main Oct 3, 2026
232 of 234 checks passed
@leofang leofang added the bug Something isn't working label Oct 3, 2026
github-actions Bot pushed a commit that referenced this pull request Oct 3, 2026
Removed preview folders for the following PRs:
- PR #2880
- PR #2917
- PR #2947
- PR #2965
- PR #2996
- PR #2998
- PR #3003
- PR #3005
- PR #3008
- PR #3010
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cuda.core Everything related to the cuda.core module test Improvements or additions to tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants