Conversation
Moon migration belongs in PR NVIDIA#2659. Restore the pre-Moon selective-CI planner and workflows from 727ef59.
The v12.9.9 branch shipped two kernelParams pointer updates without updating the generated-file seal. Preserve the released source and record the digest of the imported content after the repository EOF hook normalizes its trailing blank line. This changes no executable code.
|
/ok to test d618c51 |
|
/ok to test bcb537e |
|
/ok to test f287aed |
Major refresh 2026-09-27T15:09:12-07:00 (PDT)OutcomePR #2737 remains open and Draft. I merged current This refresh does not rename the package roots or decide the remaining architecture/policy questions for the reviewers. The temporary At the final check, Baseline and integrationThe previous PR head was f1f5c8e. Public I merged The shipped v12.9.9 Feedback and dispositionReview references: Mike's changes-requested review, Leo's release-root naming comment, Brandon's release-version comment, Brandon's cross-root-check comment, and Brandon's shared-main coupling comment. The individual inline threads contain more detail; this table records the engineering disposition.
The release resolver's historical-tag-tree fallback is retained because removing it would change old release-run behavior; that is a separate compatibility decision. A standalone metapackage sdist cannot assume the repository registry accompanies it, so a small amount of package-local version selection remains. Mike's concern about duplicated tests and Leo's concerns around root naming can be revisited during review without conflating them with the v12.9.9 sync. Local validation
Local full logs are under Commits added
Hosted CIFinal head: f287aed; trigger comment: #issuecomment-5859806421.
The first Docs job built/rendered successfully but its link checker received HTTP 429 from an external Conda documentation page, the sole error among 13,395 unique checked links. I reran only failed jobs. The new Docs job passed rendered-link validation, and the aggregate status job passed. No source change was made for this external rate-limit flake. Remaining reviewer decisions and merge reminders
The older six-layer review branch is an architectural snapshot through |
Pre-merge decisions and short-term follow-onsPR #2737 is large enough that prolonged incubation is becoming a risk of its own. Its current tree imports the maintained CUDA 12.9 source through v12.9.9, has passed fresh local CUDA 12.9 and 13.4 builds and tests, and has a green hosted CI matrix. Yet Status at this assessment
Definitely resolve before merging
The key distinction: items 1 and 2 require an explicit reviewer/owner decision before merge; they do not necessarily require more code in this PR. Items 3 and 4 are the approval and validation gates. The existing CI and dry-run evidence support a prompt merge once those gates are satisfied, but the earlier dry run does not prove the final release payload end to end. Short-term follow-ons to track
Useful lower-priority cleanup includes reconsidering the historical-tag fallback if the team drops old-tag reruns, aligning the Proposed coordinationBefore declaring #2737 ready, record the architecture and emergency-policy decisions in the PR and assign an owner and target milestone to each high-priority follow-on. That makes the trade explicit to Mike, Keith, Leo, and Brandon: merge the well-tested shared-main foundation promptly, then close the narrower gaps in independent, reviewable changes. If any reviewer judges a listed follow-on indispensable to safe first use, move that specific item into the pre-merge set rather than letting the entire list become an open-ended condition. |
…_bindings/pyproject.toml
Revert the cuFile-specific changes from bcb537e. The exceptional dry-run runner failure does not justify selecting a compatibility-enabled config for every cuFile test. Keep host-dependent test behavior while any targeted capability handling is considered separately.
|
/ok to test 5deacef |
|
/ok to test 59ff7bd |
|
This was generated by codex gpt-6-sol ultra @mdboom, following up on your review. Thank you for separating the concrete fixes from the larger design questions. I do not mean to mark the whole review resolved; this is where each point stands after the latest changes.
I also made a fresh six-layer review-only PR, whose final tree exactly matches the current #2737 head. It is a reading aid, not an alternative merge path. CI at head |
|
/ok to test b00b677 |
|
/ok to test 17e19e2 |
Work Report — GPT-6-Astra ultra — Autonomous drive to merge readinessOutcomeThe engineering work is complete, and I am comfortable with merging this implementation after the required renewed reviewer approval. PR #2737 is marked ready for review at The chosen direction is implemented: maintain the two supported CUDA ABI-major source roots together on Expand for detailsDate: September 29, 2026. Requested report start label: 0640 America/Los_Angeles. Starting point and scope
Added product commits
The user explicitly approved retaining the existing fork head as an exception to the standing upstream-branch convention, so all product commits were pushed to Decisions resolvedBranches and maintenance scope
Major-line maintenance can include explicitly selected API, Python, or platform support. Once a patch release is frozen, only approved fixes and release prerequisites enter that release. The checklist records this scope rather than treating the word "maintenance" as a complete release policy. The manual backport workflow validates a merged source PR and one existing non-default target branch. The pinned action creates a reviewable PR. Label-derived targets are disabled; there is no automatic merge. Workflow-file changes need appropriate manual credentials, and generated changes need target-specific regeneration. Branch-first emergency repairs require a linked return to Durable procedure: RELEASE-bindings.md. Layout, registry, and package boundariesKeep the explicit CUDA 12 physical root and separate lifecycle roles. The supported shape is exactly one current and one maintenance major. Retaining the mapping avoids scattering source paths through workflows; it does not implement arbitrary numbers of supported public lines. A future 13/14 rollover must preserve the CUDA 13 source in a separate root before introducing CUDA 14. Keep explicit minor-family SCM selectors so source builds do not accidentally select a prior minor's tags. Keep the standard setuptools-scm parser. The standalone metapackage's small selector/dependency policy is intentional because its sdist cannot import sibling source trees or repository CI metadata. The new guard checks the duplicated selectors and maintenance fallback against the actual package metadata. Release independence and historical compatibilityNormal development CI includes affected dependents and can build both Core ABI variants. A bindings release tag selects one bindings root and its matching metapackage, uses published Pathfinder, and does not require another-major or Core release. Stabilization branches isolate unfinished development by freezing a tested source baseline; they do not waive required checks. The tagged registry is authoritative when present. Historical trees without it retain a bounded compatibility path, with separate source and workflow-control revisions. Expired artifacts and indefinitely retired configurations are not promised to remain releasable. Cross-root review and cybindHandwritten changes require explicit applicability review in both roots. Semantic equivalence cannot generally be established by comparing bytes. Generated imports should record the generator revision, toolkit inputs, command, and any manual adjustments. Seals certify bytes, not generator provenance. Continuous toolkit QA and broader cybind asset/consumer improvements remain independent follow-ups. This change does not need a generator overhaul or QA directory move to establish the new source ownership. The final generated payload remains the reviewed v12.9.9 import plus previously disclosed overlays and seal normalization. Concrete defects fixed during this drive
Removed five redundant snapshot/static-agreement tests. Retained negative artifact selection, matrix completeness, exact-tag identity, historical-source, and mixed-workplan tests because ordinary happy-path CI does not exercise those failures. Local validation
Local logs:
Hosted validationThe product head advanced to
After marking the PR ready, a stable, fully paginated check read confirms 121 successful checks, three intentional skips, and none pending or failing. These 124 checks include auxiliary rehearsal checks; normal CI itself has 108 successful jobs. The separate The final-head CI metadata/helper gate independently confirms the same 191 tests and 55 subtests passed. The original complete attempt and the failed-job retry are preserved separately; the following section records the two intermittent Windows failures and the investigation rather than hiding them behind the final green result. Final-head Windows failure investigationThe first observed normal-CI failure was Windows Python 3.14t / CUDA 13.0.2 wheels / A100 MCDM. The Compared the exact same row against main baseline 36e4d40 and prior PR head 59ff7bd: both had the example pass and reported 3,903 passed, 486 skipped, and two expected failures. All three used published bindings 13.0.3, CuPy 14.2.0, NVRTC 13.0.88, runtime 13.0.96, and nvJitLink 13.4.92. Core source, Pathfinder source, the example, and its test have identical Git object identities between the main baseline and final head. The CUDA 12 imported bindings are not installed in that row. This evidence supports a focused diagnostic rerun; it does not identify the native crash's root cause or justify a speculative code workaround. The original attempt completed with 108 jobs: 105 successful, two failed Windows test rows, and their derived The failed-job retry completed successfully at 16:07:31 UTC. GitHub reports all 108 jobs successful: 105 successful first-attempt jobs retain their original execution timestamps, while only the two failed Windows rows and the aggregate gate execute again. The CUDA 13.0 retry explicitly passes the fractal test and reports 3,903 Core tests passed, 486 skipped, two expected failures, and 15 warnings in 276.51 seconds. The CUDA 13.4 retry explicitly passes the VMM growth case and reports 4,060 Core tests passed, 409 skipped, four expected failures, and 112 warnings in 379.06 seconds. The product source is unchanged throughout; no test suppression, relaxed assertion, or reduced matrix was used. The original failure causes remain unresolved, with the separately confirmed existing ownership defect described below. A second job failed in Windows Python 3.14t / CUDA 13.4.2 wheels / A100 MCDM: Queued x86 L4 rows delayed completion and GitHub rejected a job rerun while the containing workflow remained active. Read-only runner status was unavailable (repository runner API permission denied). Historical successful baseline jobs had already shown 78-162 minute L4 queue delays, so the observed queue alone did not establish an outage. The test matrix was preserved. A separate Windows diagnostic run uses controller The diagnostic completed green on its first attempt: all three jobs passed. The CUDA 13.0 row explicitly reports the fractal example Separate pre-existing Core ownership defectThe failure investigation uncovered and independently confirmed a duplicate deallocation attempt in the existing VMM slow-growth path. Buffer ownership captures the resource and size in a C++ deleter. Slow growth explicitly unmaps and frees the old address, then Buffer._clear() resets the owner, invoking its deallocator again. The deallocator can stop at a failed unmap; this is a duplicate cleanup attempt, not proof of a second successful free. Matching destructor warnings appear in both successful baseline runs as well as the failed run. All implicated source files are byte-identical to the main baseline. The failing temporary-reservation address differs from the warning addresses, so the logs do not establish an address-reuse interleaving or prove this defect caused that failure. No common cause with the JIT subprocess crash is established. A separate Core correction should preserve the old mapping and owner while transactionally constructing the new mapping, transfer new-resource ownership exactly once, and cover rollback and retained-view behavior. Simply resetting an ownership flag or explicitly closing the old Buffer can invalidate DLPack exports that retain it. This ownership/API follow-up is outside the branch migration's source changes and is recorded for team follow-up. Exact-source CUDA 12 rehearsalSource candidate: The controller changes only five workflow files in a separate worktree. It pins source checkouts and artifact identities to the product candidate, creates synthetic The rehearsal deliberately verifies rejection of missing synthetic 12.9.99 notes and independently validates real 12.9.9 notes. It does not manufacture a product release note. The intended payload checks cover 18 release wheels, one matching metapackage wheel, wheel METADATA and the stable metapackage's compatible-release bindings requirements (base and All 48 original prerequisites passed: 24 native wheel builds, two sdists, 17 GPU test rows, documentation, three matrix jobs, and source resolution. The original rehearsal's final custom metadata assertion was too strict: it expected one The first validation-only attempt, 36583210291, exposed another harness-only issue: hosted Evidence artifact: Downloaded and independently inspected the evidence: all 19 wheel hashes match, all wheel metadata versions are 12.9.99, and the binary matrix is exactly CPython 3.10, 3.11, 3.12, 3.13, 3.14, and free-threaded 3.14 across Linux x86-64, Linux AArch64, and Windows x86-64. Experimental 3.15/3.15t builds are excluded from publication inputs by the production validator. The source archive hash is This source rehearsal uses a branch push, so it does not establish tag-event run discovery or publication credentials. The separate production dry runs exercise exact-tag lookup with real retained artifacts. Production release dry runsControl revision:
All four completed successfully, including exact-tag CI discovery, notes validation, docs, wheel-matrix checks, and release-artifact preparation. PyPI/TestPyPI publication was skipped; no GitHub release or docs deployment was created. Exact source/artifact lineage: v13.4.3 = Final-source documentation refreshFinal source: All three jobs passed. Sphinx warnings decreased from 96 to 93; the three authored-doc defects are fixed. Remaining generated-API warning categories match the historical v12.9.9 source. Duplicate object descriptions are more numerous because this tree retains more release-note documents. The full docs build is successful, not warning-free. Review dispositionResolved all 45 threads that were open at the initial audit after checking the implementation, documented decisions, relevant validation, and fresh discussion state. GitHub now shows 52 of 52 threads resolved, including the seven that were already resolved. The PR body contains a grouped disposition map with permanent source links; the appendix below records all 45 individual decisions. New comments and thread state were checked before mutations. No thread reply or unrelated issue/comment was posted; the only new PR comments were the two expressly requested Author-side thread resolution is distinct from human approval. No review was dismissed and no approval was represented as granted. The PR was not merged. Remaining boundaries
Final repository and PR state
Appendix: individual review dispositionsThe PR description contains the grouped review map. These individual decisions cover the 45 threads that were open at the initial audit; they do not imply reviewer approval.
|
@brandon-b-miller, GPT-6-Astra ultra implemented a practical compromise. In its own words: Your concern is real: an ordinary CUDA 12 maintenance PR can still encounter CUDA 13/Core CI failures. The proposal retains that integration testing and adds a release path that can proceed from a tested baseline. The compromise has three parts:
This addresses the release-stage coupling and provides a practical way around unrelated breakage on advancing So the intended guarantee is that a bindings release does not require a Core build, test run, or release. It is not a guarantee that every possible Core problem is irrelevant to the process. |
|
@leofang Following your offline question about whether we should continue #2737, I asked Codex for a deep analysis of the Its conclusion: keep #2737’s direction, combined with explicit release branches for stabilization and urgent fixes. The main reasons:
The proposed division is therefore: develop the two supported majors together on The revised PR implements that direction with focused bindings-release CI, a manual backport workflow, and a short release guide. Normal PRs retain Core integration checks, and generated changes still need explicit applicability review. That gives us a concrete reason to finish #2737 while continuing to improve branch automation and generator coordination. |
Summary
Closes #1199. This continues Keith Kraus's original PR #2675, preserving his CUDA 12 source import and initial build, test, and release integration. The replacement became necessary when that PR's temporary base branch was deleted after #2467 merged.
Maintain both supported CUDA ABI-major source trees on
main, with short release branches available for stabilization and urgent fixes. Release each bindings line independently, together with its matchingcuda-pythonmetapackage.cuda_bindings_12/cuda_bindings/The roots build the same
cuda-bindingsdistribution andcuda.bindingsnamespace in separate environments. The CUDA 12 import is synchronized through v12.9.9. After merge,12.9.xis retained as historical release evidence.Decisions Implemented
cuda_bindings_12, with lifecycle roles represented separately in the registry. Support exactly one current and one maintenance root with distinct majors. The mapping identifies package paths; it does not promise arbitrary numbers of public release lines.main; cut short stabilization branches from a tested commit when needed. Urgent fixes can start on the frozen branch, with a linked follow-up PR to reconcile the fix back intomain. A new manual backport workflow accepts one merged PR and one existing non-default target branch, opens a reviewable PR, and leaves CI approval and merging to the normal process. The historical12.9.xbranch is not resumed.mainis unfinished; it does not waive required validation. — For details, see this comment below.setuptools-scm, with each bindings line selecting its own tag family. A temporary fallback handles CUDA 12 builds before an eligible tag is reachable. The standalone metapackage carries matching build metadata, and pre-commit/CI checks prevent those copies from drifting. See the SCM and packaging review guide below.SCM and packaging review guide
Click to un/expand
setuptools-scmhandles parsing and version progression. CUDA 12 excludes the oldv12.9.0baseline.The bootstrap regression cases cover no eligible tag, the excluded old baseline, a matching tag and its descendants, and an unreachable tag.
The key review question is: Does the custom fallback reliably get out of the way once standard Git-derived versioning can work?
The durable operating procedure is RELEASE-bindings.md, with shared ownership and rollover guidance in ci/README.md. The contributor CI overview now uses a maintained conceptual Mermaid diagram.
Implementation and Review Guide
12.9.xbackport policy. Adds explicit manual automation for selected stabilization branches..lycheeignore. The three new canonicalmain/cuda_bindings_12URLs must be rechecked after merge.The six-layer review aid #2960 remains useful for the large import and integration. Its final tree matches the earlier head
59ff7bd88756bda02164daefb0a71daff47b1d2e; it predates the September 29 completion commits. Read those additional commits here for the final policy, metadata/versioning safeguards, and historical-note fixes. Intermediate review-aid layers are not independently deployable PRs. #2737's final tree is authoritative.Validation of the Final Candidate
Final head:
17e19e25b6b20eee6a5c13c8b39befd1b2b11a38. Its only delta fromb00b677be366fd3638c78e26b99aa5b238602646is four authored documentation files: two explicit section targets, one repaired link, and an orphan marker for the release-note template. A focused production-version Sphinx/MyST render reproduces three warnings before the patch and zero afterward.Local CI tools: 191 passed, 55 subtests passed. Full pre-commit passed, with lychee checked separately: 800 successful checks and only the three expected pre-merge canonical-main 404s. Newly added and changed documentation links pass.
CUDA environment checks report
cuda-13-4.conf; no global toolkit switch was made. Native implementation code is unchanged in these completion commits.Actual v12.9.7/v12.9.9 source archives pass release resolution and note checks for both components.
Final-head CI completed green on attempt 2 at 16:07:31 UTC after
/ok to test 17e19e25b6b. The first attempt passed 105 jobs; two free-threaded Windows A100 rows failed in unchanged Core code, producing a derived status-gate failure. Both exact rows had passed on currentmainand the previous PR head. The isolated diagnostic then passed both full rows, including the originally failing tests, with identical complete package lists, Python, driver, and final-head artifacts. The normal failed-job retry passed both rows and the aggregate gate while preserving the other successful jobs. No source changes, test skips, or weakened checks were used to obtain the retry result.Security and Bandit completed green on the same final head.
After marking the PR ready, the complete check rollup has 121 successful checks, three intentional skips, and none pending or failing; this includes auxiliary rehearsal checks. Normal CI itself has 108 successful jobs. The separate pre-commit.ci commit status also passes.
CUDA 12 package-source rehearsal passed all 48 build, sdist, GPU-test, documentation, and prerequisite jobs with source
b00b677be366fd3638c78e26b99aa5b238602646and disposable controller77cfb553ed6076c5d6d23f0e6d97f2b63f953c7f. Synthetic local-onlyv12.9.99is never pushed. Publishing and docs deployment are disabled; missing synthetic notes are deliberately rejected and real 12.9.9 note routing is checked separately. The final custom audit failed on an incorrect harness expectation for the established stable~=dependency contract. After correcting that expectation and a harness-onlyghinvocation, the validation-only follow-up completed green using those successful artifacts. It verifies original-run provenance, both production wheel validators, actual base and[all]metadata, and exact-source archive identity; controller64d05e466a0f285b1876ffbfe0f281f5a733b6f9retains the evidence. No product dependency policy changed.Production release dry runs use
b00b677be366fd3638c78e26b99aa5b238602646as control revision (release tooling is unchanged in the final docs-only commit) and existing real tags, leaving run ID blank: 13.4.3 bindings, 13.4.3 metapackage, 12.9.9 bindings, 12.9.9 metapackage. All four completed green, including docs and final release-artifact validation; publication jobs were skipped. No production publication or remote release tag is requested.Final-source release-doc validation checks the final four-file docs delta, renders source
17e19e25b6b20eee6a5c13c8b39befd1b2b11a38with unchanged package wheels from the rehearsal, and asserts both repaired HTML links. It completed green, including both HTML link checks. Sphinx warnings decreased from 96 to 93; remaining generated API warning categories match the historical source, alongside duplicate object descriptions in release-note documents.Historical Evidence and Remaining Boundaries
Earlier full local QA on CUDA 13.4 covered Pathfinder, bindings in normal/PTDS modes, bindings Cython tests, and Core. An earlier isolated CUDA 12.9 build covered maintenance bindings in normal/PTDS modes and Cython tests. Those results belong to their recorded earlier revisions; the final-head CI and rehearsal above are the current evidence.
The previous publication-incapable CUDA 12 rehearsal used source
cb95f0d73e10e23d640a9559229d9d94bebde185. It is retained as historical mechanism evidence, not evidence for this final candidate.The v12.9.9 import contained a stale generated-file seal after #2911. Commit 548b2615 corrected that digest and documented a trailing blank-line normalization without changing executable generated content. This does not establish a fresh cybind regeneration. Generator/consumer qualification remains separate follow-up work.
Known inherited boundaries: two NVRTC generated-doc entries refer to functions absent from the maintenance module; raw-digest archive companions do not use conventional filename-bearing
sha256sum -cformat. The missing CUDA 12 support anchor is fixed here. README links intentionally retain published 12.9.7 documentation because the 12.9.8/12.9.9 documentation URLs currently return 404.The Windows CI investigation also confirmed an existing Core VMM slow-growth cleanup issue: explicit cleanup is followed by resetting an owning Buffer handle, attempting deallocation again. The same warnings occur in successful baseline runs, and these Core sources are unchanged by this PR. A separate ownership fix needs retained-view and rollback coverage; the logs do not establish this defect as the cause of either intermittent failure.
Supporting unreleased CTK lines, additional simultaneous public roots, source-tree deduplication, and reorganizing private QA are outside this change.
Merge Status
The implementation decisions, local checks, hosted validation, and author-side thread cleanup are complete. The PR is ready for review: all 45 initially open threads are resolved, and GitHub shows 52 of 52 threads resolved. The existing changes-requested review remains for renewed reviewer consideration and approval; no review was dismissed. After merge, validate the three newly available canonical source links on fresh
main. Production tags and publication require the normal separate release approval.Author-side review disposition map (45 initially open threads)
These dispositions describe the implementation and retained design choices. They do not represent reviewer approval.