Repository navigation
Flaky test: parallel/test-tls-ticket-cluster #29341
Description
Activity
- addedflaky-testIssues and PRs involving tests that fail intermittently in CI.Issues and PRs involving tests that fail intermittently in CI.tlsIssues and PRs related to the tls subsystem.Issues and PRs related to the tls subsystem.
on Aug 27, 2019 /cc @nodejs/cluster ?
Also on FreeBSD:
https://ci.nodejs.org/job/node-test-commit-freebsd/28120/nodes=freebsd11-x64/console
00:15:03 not ok 1983 parallel/test-tls-ticket-cluster 00:15:03 --- 00:15:03 duration_ms: 0.865 00:15:03 severity: fail 00:15:03 exitcode: 1 00:15:03 stack: |- 00:15:03 [master] got "listening" 00:15:03 [master] connecting 28298 session? false 00:15:03 [worker] connection reused? false 00:15:03 [master] got "not-reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [master] connecting 28298 session? true 00:15:03 [worker] connection reused? true 00:15:03 [master] got "reused" 00:15:03 [worker] got "die" 00:15:03 [worker] server close 00:15:03 [worker] exit 00:15:03 [master] worker died 00:15:03 [worker] got "die" 00:15:03 [worker] server close 00:15:03 [worker] exit 00:15:03 [worker] got "die" 00:15:03 [worker] server close 00:15:03 [worker] exit 00:15:03 [master] worker died 00:15:03 [worker] got "die" 00:15:03 [worker] server close 00:15:03 [worker] exit 00:15:03 events.js:186 00:15:03 throw er; // Unhandled 'error' event 00:15:03 ^ 00:15:03 00:15:03 Error: write EPIPE 00:15:03 at ChildProcess.target._send (internal/child_process.js:813:20) 00:15:03 at ChildProcess.target.send (internal/child_process.js:683:19) 00:15:03 at sendHelper (internal/cluster/utils.js:22:15) 00:15:03 at send (internal/cluster/master.js:351:10) 00:15:03 at internal/cluster/master.js:317:5 00:15:03 at Server.done (internal/cluster/round_robin_handle.js:46:7) 00:15:03 at Object.onceWrapper (events.js:298:28) 00:15:03 at Server.emit (events.js:214:15) 00:15:03 at emitListeningNT (net.js:1332:10) 00:15:03 at processTicksAndRejections (internal/process/task_queues.js:79:21) 00:15:03 Emitted 'error' event on Worker instance at: 00:15:03 at ChildProcess.<anonymous> (internal/cluster/worker.js:27:12) 00:15:03 at ChildProcess.emit (events.js:209:13) 00:15:03 at internal/child_process.js:817:39 00:15:03 at processTicksAndRejections (internal/process/task_queues.js:75:11) { 00:15:03 errno: -32, 00:15:03 code: 'EPIPE', 00:15:03 syscall: 'write' 00:15:03 } 00:15:03 ...
I'm so far unable to replicate this locally, but EPIPE can happen if you try to write to a closed stream, and there have been some changes to streams lately... @nodejs/streams? @ronag?
I can reproduce with
$ ./tools/test.py -J --repeat=500 test/parallel/test-tls-ticket-cluster.jsand bisecting points to 4367732 as first bad commit.
@nodejs/libuv
- addedchild_processIssues and PRs related to the child_process subsystem.Issues and PRs related to the child_process subsystem.macosIssues and PRs related to the macOS platform.Issues and PRs related to the macOS platform.
on Aug 31, 2019 That probably means libuv/libuv@ee24ce9 is responsible (cc @addaleax) but that also means it's probably exposing a latent bug in Node.js (edit: or the test), not introducing a new one.
The sole
uv_try_write()caller is below and thaterr == UV_EAGAINcheck looks highly suspect:
Lines 350 to 354 in 98b7185
err = uv_try_write(stream(), vbufs, vcount); if (err == UV_ENOSYS || err == UV_EAGAIN) return 0; if (err < 0) return err; Reacted by Luigi Pinca and Alex YangIt looks more like issues with the cluster/child process code to me… currently, synchronous write errors from
.send()are reported (but there is an ancient undocumented option to turn that off,swallowErrors), while asynchronous write errors are always discarded silently.It generally also looks like there’s very little error handling in the
clustermodule in general, and in particular, there’s this bit that’s also in the stack trace:node/lib/internal/cluster/round_robin_handle.js
Lines 45 to 48 in 98b7185
// TODO(bnoordhuis) Check err. send(null, { sockname: out }, null); } else { send(null, null, null); // UNIX socket. Judging from the commit message of 0da4e0e (which added
swallowErrors), I’d think it make sense to add that flag for internal cluster messages, too./cc @cjihrig
The sole
uv_try_write()caller is below and thaterr == UV_EAGAINcheck looks highly suspect:I think
UV_ENOSYSandUV_EAGAINare precisely the return values that indicate that it makes sense to useuv_write()after a faileduv_try_write(), correct?Reacted by Luigi PincaI think UV_ENOSYS and UV_EAGAIN are precisely the return values that indicate that it makes sense to use uv_write() after a failed uv_try_write(), correct?
Yes, but before libuv/libuv@ee24ce9 that path was always taken. The
if (err < 0)below was dead code but now it's not.This failure is showing up on v10.x now, specifically on centos7. Did this get figured out on master? Should we mark it as flaky?
This failure is showing up on v10.x now, specifically on centos7. Did this get figured out on master? Should we mark it as flaky?
Using
ncu-ci walk, I can only find one recent example, and it's a node-daily-v10.x-staging daily job.Saw this one failing on OS X today: https://ci.nodejs.org/job/node-test-commit-osx/33696/nodes=osx1015/testReport/junit/(root)/test/parallel_test_tls_ticket_cluster/
github-actions commented
on Jun 27, 2026 on Jun 27, 2026 – with GitHub ActionsContributorMore actionsThis issue has been marked as stale due to 210 days of inactivity.
It will be automatically closed in 30 days if no further activity occurs. If this is still relevant, please leave a comment or update it to keep it open.- addedstaleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.Issues and PRs marked stale due to inactivity and scheduled for automatic closure.
on Jun 27, 2026 github-actions commented
on Jul 28, 2026 on Jul 28, 2026 – with GitHub ActionsContributorMore actionsThis issue has been automatically closed after 30 days of inactivity following its stale status (no activity for a total of 120 days).
If this is still relevant, feel free to reopen it or leave a comment with additional details so we can continue the discussion.
Console output