Conversation
… aborting The five cuda_safe_call uses in stream_task.cuh's guard bodies are event releases on the way out: the two timing events destroyed in SCOPE(exit) of each operator->*, and the synchronization event destroyed after sync_with_stream's wait. They become cuda_try under ON_THROW(notify), one policy per release: a failing destroy leaks that event and is reported, the guard is noexcept so it cannot become a second exception, and a leaked event is not worth ending the program over. The six calls in the SCOPE(success) timing blocks are left to NVIDIA#11652, which replaces those blocks with task_statistics::record_task_timing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
/ok to test 7fef5c4 |
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 📝 SummarySummary by CodeRabbit
WalkthroughCUDA event cleanup calls now use ChangesCUDA event cleanup
Priority: ⬇️ Low Change: Bug fix Merge Risk: 🔵 Low · up to A failed event destroy still terminates the application in supported no-exceptions builds rather than continuing cleanup. This is a narrow gap in the promised behavior and matches the base failure mode, so it is bounded but should be addressed or explicitly scoped.
Comment ✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
|
There was a problem hiding this comment.
Actionable comments posted: 1
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository: NVIDIA/cccl/.coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 580aaaed-551b-488e-9ef5-f784437ab1e3
📒 Files selected for processing (1)
cudax/include/cuda/experimental/__stf/stream/stream_task.cuh
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.
This comment has been minimized.
This comment has been minimized.
😬 CI Workflow Results🟥 Finished in 1h 39m: Pass: 91%/72 | Total: 1d 09h | Max: 1h 04m | Hits: 22%/182742See results here. AI failure analysis1. Rocky Linux wheel provisioning: CUDA repository metadata returns 404 · 2 jobsExplanation: Both jobs fail before compiling CCCL or STF while installing GCC 13 and ccache in the same CUDA 12.9 Rocky Linux wheel image; DNF consults an enabled CUDA repository whose referenced metadata files consistently return 404. The PR diff only changes STF event cleanup, so the evidence indicates a shared container/repository provisioning failure rather than a regression in the changed source. Evidence: Copy this prompt into a coding agentJobs: |
…ed under a policy, not an abort (#11773) The SCOPE(exit) that destroys the two timing events in each operator->* becomes cuda_try under ON_THROW(notify), one guard per release, as in launch (#11748) and stream_task (#11770): a failing destroy leaks that event and is reported, the guard is noexcept so it cannot become a second exception, and a leaked event is not worth ending the program over. The timing blocks in SCOPE(success) are left to #11652. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Description
Next
cuda_safe_callbatch after #11740 (graph_task) and #11748 (launch): stream_task.cuh.Five of its eleven
cuda_safe_calluses sit inside guard bodies, where a throw would terminate, so each needs an explicit policy. All five are releases of an event on the way out:SCOPE(exit)of eachoperator->*(two copies);SCOPE(exit)after the wait insync_with_stream.Same policy as #11748:
cuda_tryunderON_THROW(notify), one guard per release. A failing destroy leaks that event and is reported to stderr; the guard isnoexcept, so it cannot become a second exception; and a leaked event is not worth ending the program over, which is whatabort()did. One policy per release, so a failure on the first destroy does not skip the second.The other six calls are the
SCOPE(success)timing blocks, which #11652 replaces withtask_statistics::record_task_timing; they are left to that PR to avoid a conflict.A pattern worth noting, now visible across the series: the "destroy the two timing events on exit" guard exists in five copies (launch, parallel_for_scope, host_launch_scope, and twice here). Once #11652 and this PR have landed, folding those into one RAII holder with the release policy in its destructor is the natural follow-up.
Remaining batches: graph_ctx.cuh (7), then the long tail.
Checklist
🤖 Generated with Claude Code