Skip to content

test: handle partial NVML support on Jetson Orin - #2947

Merged
rwgk merged 14 commits into
mainfrom
rwgk/orin-ctk-13-4-tests
Oct 2, 2026
Merged

rwgk merged 14 commits into
mainfrom
rwgk/orin-ctk-13-4-tests

Conversation

@rwgk

@rwgk rwgk commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Description

fixes Orin failures (cuDLA nightly testing)

CUDA Toolkit 13.4 exposes partial NVML support on Jetson Orin: system-level
queries work, while parts of the device API remain unavailable. This caused
tests that were skipped with CUDA Toolkit 13.3 to run and fail.

This change:

  • distinguishes basic NVML availability from device API availability, keeping
    supported bindings and cuda.core coverage enabled;
  • skips only cuda.core tests that require unavailable NVML device APIs;
  • treats the PCI lookup used by device_discover_gpus as an optional device
    capability; and
  • allows bounded driver bookkeeping in the VMM leak regression while still
    detecting the repeated leak that the test covers.

Testing

The full Linux QA build and test workflow was run from scratch on a Jetson AGX
Orin board with CUDA Toolkit 13.4. A clean test rerun completed without
failures:

jetson:~/wrk/forked/cuda-python $ grep_pytest_summary `nlog`
/home/rgrossekunst/wrk/logs/cuda-python_qa_bindings_linux_2026-09-24+005744+0000_tests_log.txt
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_pathfinder
======================= 1618 passed, 2 skipped in 13.23s =======================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 544 passed, 79 skipped, 18 warnings in 63.24s (0:01:03) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 544 passed, 79 skipped, 18 warnings in 62.27s (0:01:02) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
============================== 9 passed in 0.99s ===============================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
= 3960 passed, 464 skipped, 3 xfailed, 7 warnings, 8 subtests passed in 536.62s (0:08:56) =
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
============================== 1 passed in 0.33s ===============================

Log inspection confirms that bindings NVML tests were collected and executed,
unsupported device-dependent cuda.core tests were skipped with the intended
reason, and both test_vmm_allocate_close_does_not_leak[allocate] and
test_vmm_allocate_close_does_not_leak[grow] passed.

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes (no user-facing documentation change).

@rwgk rwgk added this to the cuda.core 1.2.1 milestone Sep 24, 2026
@rwgk rwgk added the test Improvements or additions to tests label Sep 24, 2026
@copy-pr-bot

copy-pr-bot Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@rwgk rwgk self-assigned this Sep 24, 2026
@github-actions github-actions Bot added cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module labels Sep 24, 2026
@rwgk

rwgk commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor
Doc Preview CI
Preview removed because the pull request was closed or merged.

@rwgk
rwgk marked this pull request as ready for review September 30, 2026 19:45
@rwgk
rwgk requested a review from mdboom September 30, 2026 19:45
Comment thread cuda_core/tests/system/test_system_device.py
Comment thread cuda_core/tests/system/test_system_events.py
Comment thread cuda_core/tests/test_memory.py Outdated
assert baseline - free < aligned_size
# The broken path leaks aligned_size per iteration. Allow one allocation's
# worth of driver bookkeeping/caching while still detecting repeated leaks.
assert baseline - free < 2 * aligned_size

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it would be clearer / less brittle to not leak this by making allocate_and_close only run once. Maybe make it a global (non-nested) function and add the @functools.cache decorator?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in f2b7822. I moved the allocation/close operation to a module-level helper and added a cached module-level warm-up wrapper, so the one-time driver bookkeeping happens outside the measurement. The eight measured calls remain uncached so the regression still detects a repeated leak, and the strict < aligned_size assertion is restored. The from-scratch Orin rerun at that commit passed both variants.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

Retained the cached module-level warm-up from Ralf's earlier reply. It is keyed by device ID and growth mode; the eight measured allocation/close calls remain uncached, and the assertion stays baseline - free < aligned_size. Warm-up and measurement now share VMM_LEAK_TEST_REQUESTED_SIZE. Current implementation.

Both variants passed the focused TestVenv run. The grow case emitted a cleanup warning that also reproduced with the pre-change test; this change does not resolve that warning. Ralf's latest from-scratch rerun at the current head also completed without test failures.

@rwgk

rwgk commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Converting this PR back to Draft mode, to avoid triggering the full CI. — I want to push here to retest on Orin first.

@rwgk
rwgk marked this pull request as draft September 30, 2026 21:06
@rwgk

rwgk commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Retesting from scratch on Orin @ f2b7822 was successful:

jetson:~/wrk/forked/cuda-python $ grep_pytest_summary `nlog`
/home/rgrossekunst/wrk/logs/cuda-python_qa_bindings_linux_2026-09-30+212544+0000_tests_log.txt
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_pathfinder
======================= 1618 passed, 2 skipped in 13.86s =======================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 578 passed, 80 skipped, 18 warnings in 68.46s (0:01:08) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 578 passed, 80 skipped, 18 warnings in 65.74s (0:01:05) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
============================== 9 passed in 0.99s ===============================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
= 4003 passed, 465 skipped, 3 xfailed, 7 warnings, 8 subtests passed in 681.52s (0:11:21) =
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
============================== 1 passed in 0.33s ===============================

@rwgk
rwgk marked this pull request as ready for review September 30, 2026 21:44
@rwgk

rwgk commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

@mdboom could you please take another look?

The fixes and the individual replies to your feedback were generated by GPT-5.6-Sol medium.

Comment on lines +71 to +74
skip_if_nvml_device_apis_unsupported = pytest.mark.skipif(
_should_skip_nvml_tests() or not hardware_supports_nvml_device_apis(),
reason="NVML device APIs are incomplete or unavailable on this platform",
)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Instead of skipping if the device API is unsupported, I wonder if it makes sense to instead allow catching the specific nvml.NotSupportedError exception or whatever we would expect the behavior to be when the underlying used APIs aren't supported?

Since users would presumably get that exception if they tried to use the cuda.core.system APIs, it would be good to test that we're delivering the desired / expected user experience?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I gave this as a task to codex GPT-6.1-Sol ultra. It looks like it found an existing bug, which it fixed, and then it made the changes as suggested by @kkraus14. It's now a much bigger PR. @mdboom what do you think about that direction?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The unsupported_before helper function already catches NotSupportedError in order to skip the test. We could use that everywhere.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I asked codex to backtrack, but to preserve genuine improvements. The resulting commits are:

  • 7554cd7 test(cuda.core): restore skips for unsupported NVML APIs
  • 4ed9dd0 fix(cuda.core): remove NVLink state preflight
  • 497203c test(cuda.core): share VMM leak test allocation size

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

The corrective commits Ralf listed restore accurate skip reporting after my broader rewrite accepted unsupported operations without checking their intended result. The conservative device capability gate is back on the affected hardware tests.

Mike is right that unsupported_before already converts NotSupportedError into a skip in its optional-support branches. There is a nuance for Orin: it reports AMPERE, so a call guarded with a KEPLER minimum reaches the branch that propagates errors. Replacing the gate with that existing usage everywhere would therefore still fail here. Helper branches.

Keith's caller-behavior coverage remains in focused, deterministic regressions: CUDA-to-NVML lookup must propagate injected NotFoundError and NotSupportedError, and NVLink count/iteration must propagate injected per-field errors. These tests assert the exceptions rather than treating unsupported hardware queries as successful operation coverage. Lookup regression, NVLink regressions. Supported NVML-to-CUDA UUID mapping is also asserted on this Orin host, with its unsupported optional PCI check reported as a separate skipped subtest.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is a nuance for Orin: it reports AMPERE, so a call guarded with a KEPLER minimum reaches the branch that propagates errors. Replacing the gate with that existing usage everywhere would therefore still fail here.

As I said here, we need to update that logic, so that on Orin we ignore the passed-in architecture.

rwgk added 2 commits October 1, 2026 04:26
Preserve UUIDs reported without a prefix and propagate unsupported NVLink queries before reading unpopulated field results. Keep zero-count behavior for devices without link zero, and cover UUID normalization, lookup errors, and NVLink error propagation with regressions.
Replace the UUID-based device-wide skip with narrow checks of the actual API result. Exercise supported Orin queries, validate precise unsupported errors, preserve CUDA visibility and MIG coverage, and require successful UUID matching when a CUDA counterpart exists.
@rwgk

rwgk commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Retesting from scratch on Orin @ 2ed2554 was successful:

jetson:~/wrk/forked/cuda-python $ grep_pytest_summary `nlog`
/home/rgrossekunst/wrk/logs/cuda-python_qa_bindings_linux_2026-10-01+043322+0000_tests_log.txt
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_pathfinder
======================= 1618 passed, 2 skipped in 13.89s =======================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 578 passed, 80 skipped, 18 warnings in 68.62s (0:01:08) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 578 passed, 80 skipped, 18 warnings in 66.02s (0:01:06) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
============================== 9 passed in 0.98s ===============================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
= 4033 passed, 446 skipped, 3 xfailed, 7 warnings, 32 subtests passed in 630.37s (0:10:30) =
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
============================== 1 passed in 0.34s ===============================

Comment thread cuda_core/tests/system/test_system_nvlink.py Outdated
Comment thread cuda_core/tests/test_device.py Outdated
Comment thread cuda_core/tests/test_device.py Outdated
system_device = device.to_system_device()
except system.NotFoundError:
# Orin enumerates NVML devices but does not support lookup by UUID.
return

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this should only fail on Orin, we should assert that here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

Removed the blanket NotFoundError early return in 7554cd70d37. The integration test now uses the restored capability probe and records a skip when the probe finds incomplete device APIs; once the probe passes, an unexpected conversion error fails the test. Separate deterministic tests explicitly assert propagation of injected NotFoundError and NotSupportedError. Integration test, Error propagation tests.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This solution is so much broader than my suggestion, I'm not sure why it did that and what it was trying to accomplish.

Comment on lines +18 to +19
now raise :class:`system.NotSupportedError` when Orin does not support
NVLink queries, instead of reading an unpopulated NVML field result.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The check isn't actually Orin-specific.

Suggested change
now raise :class:`system.NotSupportedError` when Orin does not support
NVLink queries, instead of reading an unpopulated NVML field result.
now raise :class:`system.NotSupportedError` when the device does not support
NVLink queries, instead of reading an unpopulated NVML field result.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

Removed this release-note entry together with the state preflight in 4ed9dd0350d. Requiring link-zero state support could reject an otherwise valid zero count, so the count again comes from its own field API. The replacement regression asserts zero count and empty iteration with an unsupported state API. Regression.

A fresh probe on this Orin driver still shows the field call returning successfully without changing supplied output sentinels. That driver/binding issue remains unresolved; the restored hardware-test gate avoids claiming a runtime repair.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't understand how your solution flows from my suggestion.

Comment thread cuda_core/docs/source/release/1.2.1-notes.rst Outdated
Comment on lines +111 to +114
try:
affinity = device.get_cpu_affinity(typing.AffinityScope.NODE)
except nvml.NotSupportedError:
continue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This turns a skip (when nvml.NotSupportedError is raised) into a silent pass. This seems like an undesirable test regression to me (and this pattern is repeated ~20 times in this file).

(I wonder if this is due to responding to @kkraus14's comment about checking for NotSupportedError and not realizing that unsupported_before already does that).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We may need to extend unsupported_before with a new argument "on_orin", that when True would happily skip if NotSupportedError is raised on the Orin platform. (Rather than just a too-old architecture as passing typing.DeviceArch.KEPLER currently does).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

Agreed; this was a regression in my rewrite. The affected inline NotSupportedError catches that passed, returned, or continued without validating the operation are rolled back, and the device capability gate is restored. The retained mapping test asserts UUID conversion in one subtest and uses unsupported_before(device, None) for optional PCI validation in another. The focused Orin run reports those as SUBPASSED and SUBSKIPPED, respectively. Mapping test.

The proposed on_orin extension is not included in this correction. Orin's AMPERE architecture does bypass the helper's older-than-KEPLER branch, so finer coverage needs an explicit capability/platform policy. The restored gate provides conservative skip reporting now; that finer policy remains follow-up work. Ralf's fresh current-head Orin rerun also completed without test failures.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, I'm willing to accept this given that the tests passed on Orin, but this contradicts the agents statement here: #2947 (comment)

Comment thread cuda_core/tests/system/test_system_system.py
assert resolved_device.uuid == device.uuid
assert resolved_device.index == device.index

# PCI lookup can work even when the PCI information query is unsupported.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't understand this comment -- it seems to be skipping things when the information query is unsupported.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

The removed fallback distinguished obtaining a handle by PCI bus ID from querying its PCI information. The final continue did omit the information assertions when that query was unsupported. I rolled back this event-test rewrite in 7554cd70d37, including the misleading comment.

Hardware device resolution and event registration again use the capability gate; synthetic event wrapping and packed PCI-ID conversion remain enabled. Restored event tests.

Comment thread cuda_core/tests/system/test_system_events.py Outdated
Comment thread cuda_core/tests/test_memory.py Outdated
device,
config=VirtualMemoryResourceOptions(handle_type="win32_kmt" if IS_WINDOWS else "posix_fd"),
)
return _vmm_allocate_and_close(mr, 8 * 1024 * 1024, grow)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NIT: Put 8 * 1024 * 1024 in a global constant so the two mentions of it stay in sync.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

GPT-6.1-Sol ultra running on Orin:

Done in 497203c78d1: warm-up and measured allocations both use VMM_LEAK_TEST_REQUESTED_SIZE. The cached warm-up, eight uncached measured calls, and strict leak threshold are preserved. Updated VMM test.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why was this file completely removed?

rwgk added 3 commits October 1, 2026 16:42
Restore the conservative device capability gate and the original device and event test bodies instead of accepting unsupported calls as successful tests. Keep positive UUID mapping checks, report unsupported PCI validation as a separate skipped subtest, restore the process-name NotFound skip, and preserve affinity cleanup and fan serialization.
Read the NVLink count from its own field API without requiring support for the link-zero state API. Replace the preflight regressions with valid zero-count and typed per-field error assertions, regenerate the public stub, and retain only the verified UUID release note with the requested wording.
Use one module constant for the cached warm-up and uncached measured allocations. Preserve eight measured iterations and the strict one-allocation leak threshold.
@rwgk
rwgk marked this pull request as draft October 1, 2026 16:56
@rwgk

rwgk commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

/ok to test

@rwgk

rwgk commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Retesting from scratch on Orin @ 497203c was successful:

jetson:~/wrk/forked/cuda-python $ grep_pytest_summary `nlog`
/home/rgrossekunst/wrk/logs/cuda-python_qa_bindings_linux_2026-10-01+170946+0000_tests_log.txt
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_pathfinder
======================= 1618 passed, 2 skipped in 13.87s =======================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 578 passed, 80 skipped, 18 warnings in 67.18s (0:01:07) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 578 passed, 80 skipped, 18 warnings in 65.98s (0:01:05) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
============================== 9 passed in 0.99s ===============================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
= 4016 passed, 464 skipped, 3 xfailed, 7 warnings, 9 subtests passed in 602.18s (0:10:02) =
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
============================== 1 passed in 0.34s ===============================

@rwgk
rwgk marked this pull request as ready for review October 1, 2026 17:56
Preserve the Orin device-API skip while adopting cuda.core's bindings floor, simplify now-obsolete test guards, and move the UUID fix note to the upcoming release.
@github-actions github-actions Bot added the CI/CD CI/CD infrastructure label Oct 1, 2026
@rwgk

rwgk commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Retesting from scratch on Orin @ 66bf6bb was successful:

jetson:~/wrk/forked/cuda-python $ grep_pytest_summary `nlog`
/home/rgrossekunst/wrk/logs/cuda-python_qa_bindings_linux_2026-10-01+214213+0000_tests_log.txt
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_pathfinder
======================= 1618 passed, 2 skipped in 13.92s =======================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 586 passed, 80 skipped, 18 warnings in 68.40s (0:01:08) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 586 passed, 80 skipped, 18 warnings in 66.87s (0:01:06) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
============================== 9 passed in 1.01s ===============================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
= 4076 passed, 463 skipped, 3 xfailed, 7 warnings, 9 subtests passed in 694.83s (0:11:34) =
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
============================== 1 passed in 0.34s ===============================

@rwgk

rwgk commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor Author

How the merge conflicts were resolved in 66bf6bb

Git reported one content conflict, in cuda_core/tests/test_device.py::test_to_system_device. The PR side had changed the test to use the init_cuda fixture and to skip if NVML cannot look up a device by UUID, which is the specific Orin limitation. main removed the older CUDA_BINDINGS_NVML_IS_COMPATIBLE check because its new minimum cuda-bindings requirement makes that flag always true.

The resolution keeps the PR's init_cuda fixture, authorship marker, and hardware_supports_nvml_device_apis() probe. It takes main's removal of the obsolete compatibility check and _system import. A platform with incomplete UUID lookup is recorded as skipped; on a supported platform the test still checks the conversion and UUID match.

Resolve the VMM test conflict by adopting the upstream handle-layer regressions. Their exact mapping and reservation release checks supersede the cached warm-up and device-wide free-memory measurement. Preserve the Orin NVML fixes and UUID release note.
@rwgk

rwgk commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

GPT-6.1-Sol ultra running on Orin:

How the merge conflict in e2ae412 was resolved

The only conflicted file was cuda_core/tests/test_memory.py. I resolved it by adopting main’s version, which accompanies the move of VirtualMemoryResource to the Cython/C++ _rt handle layer in #2917.

The upstream tests replace device-wide free-memory sampling with checks that the exact CUDA mappings and address reservations are released. This supersedes our cached allocator warm-up, shared allocation-size constant, and old free-memory comparison.

The tests also reflect the new ownership behavior: modify_allocation() returns a new buffer while the original remains open. They explicitly close all buffers and cover allocation, ordinary growth, and growth requiring relocation, while checking for cleanup warnings. Updated leak regression.

The Orin NVML changes and UUID release note were preserved.

After rebuilding the merged implementation in TestVenv:

  • VMM tests: 39 passed, 1 skipped, 1 expected failure. All three leak modes passed without cleanup warnings.
  • System/device tests: 66 passed, 41 skipped, 9 subtests passed.
  • Pre-commit checks passed except the existing all-files TruffleHog crash (-11); scanning the resolved file separately passed.

@rwgk

rwgk commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Retesting from scratch on Orin @ e2ae412 was successful:

jetson:~/wrk/forked/cuda-python $ grep_pytest_summary `nlog`
/home/rgrossekunst/wrk/logs/cuda-python_qa_bindings_linux_2026-10-02+154404+0000_tests_log.txt
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_pathfinder
======================= 1618 passed, 2 skipped in 13.64s =======================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 600 passed, 80 skipped, 18 warnings in 67.06s (0:01:07) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
=========== 600 passed, 80 skipped, 18 warnings in 66.66s (0:01:06) ============
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_bindings
============================== 9 passed in 0.99s ===============================
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
= 4117 passed, 463 skipped, 3 xfailed, 5 warnings, 9 subtests passed in 638.17s (0:10:38) =
rootdir: /home/rgrossekunst/wrk/forked/cuda-python/cuda_core
============================== 1 passed in 0.33s ===============================

@mdboom mdboom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks good enough now.

The agent's responses to my comments were really weird, I suspect because it rolled back so many changes from the last version. So on a point-by-point basis it doesn't seem ok, but re-reading it from scratch, I think it's fine.

@rwgk
rwgk enabled auto-merge (squash) October 2, 2026 16:43
@rwgk
rwgk merged commit a52068d into main Oct 2, 2026
304 of 312 checks passed
@rwgk
rwgk deleted the rwgk/orin-ctk-13-4-tests branch October 2, 2026 17:45
@rwgk

rwgk commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

To close the loop here, the scheduled testing on the Orin board passes again:

cuda-python tests on L4T (Orin) [PASSED] - 20261003_003002

cuda-python Test Report
=======================
Board  : local_l4t
Date   : Sat Oct  3 12:52:20 AM IST 2026
Result : PASSED

--- Suite Results ---
cuda_pathfinder  : 1618 passed, 0 failed, 2 skipped, 0 xfailed
cuda_bindings    : 600 passed, 0 failed, 81 skipped, 0 xfailed
cuda_core        : 4106 passed, 0 failed, 469 skipped, 3 xfailed

Total            : 6324 passed, 0 failed, 3 xfailed
Note             : xfailed = expected failures (known issues, marked xfail in tests)

github-actions Bot pushed a commit that referenced this pull request Oct 3, 2026
Removed preview folders for the following PRs:
- PR #2880
- PR #2917
- PR #2947
- PR #2965
- PR #2996
- PR #2998
- PR #3003
- PR #3005
- PR #3008
- PR #3010
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD CI/CD infrastructure cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module test Improvements or additions to tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants