Repository navigation
Investigate flaky test-child-process-fork-net #5122
Description
Activity
- addedtestIssues and PRs related to Node.js core tests and test infrastructure.Issues and PRs related to Node.js core tests and test infrastructure.armIssues and PRs related to the ARM architecture.Issues and PRs related to the ARM architecture.
on Feb 6, 2016 - addedchild_processIssues and PRs related to the child_process subsystem.Issues and PRs related to the child_process subsystem.
on Feb 6, 2016 Seems like the capacity for the Raspberry Pi 2 devices to connect to localhost has diminished a lot. That doesn't make any sense, but there's also no arguing with the fact that we've probably fixed a half a dozen tests or more by simply reducing their connections to localhost from modest (in the 10-200 range) to tiny (less than 20). Not sure what's up there. Looping in @nodejs/build for ideas... In the meantime, I'll experiment with reducing the number of clients/connections in this test to see if it fixes things...
Stress test against current master for a benchmark on failure frequency: https://ci.nodejs.org/job/node-stress-single-test/588/nodes=pi2-raspbian-wheezy/console
And here's one for reducing the connections from 10 to 8: https://ci.nodejs.org/job/node-stress-single-test/589/console
Argh! So, the test does not (perhaps can not?) clean up after itself if it fails, so once it fails once, a bunch of other things fail with
EADDRINUSE. This includes this test itself if it is run again as in a stress test, so we can't even really get much of a stress test other than "Yeah, it fails once in a while." But we can't really gauge relative flakiness of different scenarios.For what it's worth, the test has not been significantly changed since it was created in 2012.
Reducing to four connections and trying again with the stress test. If that doesn't fix it, then something must be really wrong... https://ci.nodejs.org/job/node-stress-single-test/590/nodes=pi2-raspbian-wheezy/console
Argh! So, the test does not (perhaps can not?) clean up after itself if it fails, so once it fails once, a bunch of other things fail with
EADDRINUSE.Yeah, I've fixed some tests in the past that had this particular cascading effect, but I'm sure there's probably a lot still out there that do not clean up properly on unexpected failure.
Reducing from 10 to 4 worked. Might just go with that. /cc @AndreasMadsen in case there's any additional insight to why 10 in the first place. (I've been assuming the answer is "why not 10?", but I don't actually know.)
Might also take a moment to see if it's feasible to more definitively close the server on failure...
- added a commit that references this issue
on Apr 10, 2016 Proposed fix, or mitigation at least: #6138
@Trott Hmm, it is such a long time ago that I don't remember the specific reason. But it doesn't look like it is a stress test, so I think reducing it to 4 connections is reasonable.
Reacted by Rich Trott- added a commit that references this issue
on Apr 12, 2016 - added a commit that references this issue
on Apr 26, 2016 - added 2 commits that reference this issue
on May 17, 2016
Example failure: