Repository navigation
ACTION REQUIRED: workflow change for merging changes to nodejs/node #2598
Description
Activity
👍
- addedbuildIssues and PRs related to Node.js builds or CI infrastructure.Issues and PRs related to Node.js builds or CI infrastructure.metaIssues and PRs related to the general management of the project.Issues and PRs related to the general management of the project.
on Aug 28, 2015 Do these merges add the committer line of whoever performed the merge?
Do these merges add the committer line of whoever performed the merge?
I was told they will, and they need to for us to use it.
Do these merges add the committer line of whoever performed the merge?
Yes, the information is taken from your full name and email address as set in Jenkins. I will update the wiki to reflect that.
What about documentation-only changes? I believe we don't yet have documentation-related tests, do we?
To ensure you committer info is set correctly, go to https://jenkins-iojs.nodesource.com/, login, click on your username on the top-right, then on the left click on 'Configure'. Double-check 'Full Name' and 'E-mail address'.
What about documentation-only changes? I believe we don't yet have documentation-related tests, do we?
Use node-accept-pull-request, and you can set NODES_SUBSET to
pure_docs_changesto skip the tests. If we'll ever invent a way to test documentation, those tests will run as part of that configuration 😄On the left hand side, you should see "Build with parameters". If not, it probably means that you're not logged in. You need to be logged in to start this job.
Doesn't work for me. I am logged in and there is nothing like «Build with parameters». Can I have a screenshot?
I remeber it being ok in
iojs+any-pr+multi.No, it's definitely not there.
@ChALkeR did you go to https://jenkins-iojs.nodesource.com/job/node-accept-pull-request/ first ?
@orangemocha Yes, of course.
51 remaining items
The high level guess is that due the fast pace of development the quality is slowly drifting down (no offense to anyone, I hope).
Ok this is just incorrect.
They were failing BEFORE THAT. And we cleaned it up besides what was still in our issue tracker. Some of these failures have only manifested in the last few days. I've been keeping an eye on the CI since the beginning of io.js.
there are still a majority of flaky failures that cannot be clearly attributed to issues with the slaves.
Anything with failures relating to reset connections or port binding is a) relatively new, and b) appeared around some CI changes about 2 or so months ago.
Also now the
iojs+any+pr+multijob and all of it's history is gone. Wonderful.We are already making those calls when we say "not related to this change" and try again (or land it manually). Once we make that decision and move forward with the PR, the flakiness is now in master.
We can also parse if the output is or isn't actually flaky.
They were failing BEFORE THAT. And we cleaned it up besides what was still in our issue tracker.
Also for clarification, I am quite sure we have had green runs for over a week, at a couple intervals, on master. (This was after the point where we were accidentally ignoring windows failures.)
I doubt that the failures are related to the shift from iojs+any+pr+multi to node-test-pull-request. They differ in the arguments they take and the way they fetch things, but internally they are just calling the test runner in the same way.
I have to agree that this started in the last few days. Before proceeding with opening this issue, we had several test runs with node-accept-pull-request where things were stably yellow. That means that there might be a real bug in node. If someone has time to start investigating those flaky tests locally, it would be helpful.
OK, after chasing after failures for the last few days, I am convinced that we need to suspend this experiment. 😞
Even though we tried aggressively to mark tests as flaky - this PR is marking 29 new tests as flaky! - we are not even able to land that PR because new tests are failing every time. Given the current state of things, most attempts to land PRs via CI would fail. So please refrain from using node-accept-pull-request / node-merge-commit, and instead merge changes manually.This level of flakiness in the tests is unprecedented, and I believe it started very recently (this week). While there are occasional failures due to misconfigured machines, those are easy to spot. The vast majority of failures seem due reasons outside of the Jenkins/CI realm, and seem to indicate a real problem with the state of the master branch. The recent libuv upgrade (a161594) is high on the list of suspects. What we can do is run some test jobs on recent commits and try to bisect it that way. If you have a chance to investigate some of those failures locally, that would be helpful too.
We can resume this experiment after we have found the underlying cause of all this instability. I guess that at this point, we'll also take the time to fix a few other high priority issues with the CI itself (like making the build faster), and restart only when it looks really solid.
Sorry for the headaches that this has caused.
/cc @nodejs/collaborators @nodejs/tsc
These flaky tests have been plaguing us a lot longer than a week, and the failures were almost never reproduceable locally, so I think we indeed should take a close look on the CI itself. Are other projects using Jenkins also experiencing this? Maybe we're doing something fundamentally wrong?
Jenkins is just running a bunch of scripts on a variety of machines. Sure, it has its own flukes, but I don't see how it could be causing some of the issues that we have seen lately.
Maybe to get a better shot at reproducing locally, we should try running make run-ci. The top failing platforms seems to be armv7-wheezy, RPis, and centos5.
Wild guess, probably wrong, but hey, that's what the Internet is for, so here you go:
Didn't we move a test or three from
sequentialtoparallelrecently? Maybe something in one of those tests slipped by that may be making other tests somehow unstable?That's not to be excluded @Trott . I already started a few runs to try and bisect commits in CI, so that if there was one offending commit, we can find it.
Also, I just checked PRs that landed in v0.12 recently. About 10 PRs landed in the last couple of weeks using node-accept-pull-request from the new Jenkins, using the same exact Jenkins infra, except they don't run on ARM. No records of people having to retry runs. This also seems to confirm that the problem is in the current master branch.
The vast majority of failures seem due reasons outside of the Jenkins/CI realm, and seem to indicate a real problem with the state of the master branch.
Anything that deals with
ECONNRESETor unavailable ports has in the past been said to be configuration, that stuff was already there before the Libuv upgrade. Those were the largest CI issue for a while iirc.Related: I saw two unavailable ports after like 50 test suite runs on my OS X 10.10.5 machine. None on my remote Ubuntu 15 testing box. (It did has a weird fs
EACCESSbut I am quite certain this is config also.)@orangemocha what do you think about delaying the accept-pullrequest project? It looks like it is blocking us from making progress at the moment, and considering the upcoming release - it is very troublesome. We can always return to running the experiments with this at some later point, when things will be more stable.
@indutny yes see above #2598 (comment)
@orangemocha argh, right! sorry, I didn't get it.
Closing for now. I will open a new issue when we have addressed the stability and performance concerns.


Based on previous discussion, feedback, and the approval by the TSC (#2434), we are adopting a new workflow for merging pull request in nodejs/node, and making any changes to the code branches at nodejs/node in general. The new workflow needs to be adopted by all @nodejs/collaborators for all changes to be merged to nodejs/node, starting this coming Monday at 2pm UTC: http://www.timeanddate.com/worldclock/fixedtime.html?msg=Start+merging+changes+to+nodejs%2Fnode+with+Jenkins&iso=20150831T14&p1=1440
From that time on, please refrain from pushing changes to nodejs/node manually, and instead use the workflow documented at https://github.com/nodejs/node/wiki/Merging-pull-requests-with-Jenkins. I also added a wiki page specifically on how to deal with flaky tests.
The workflow for testing (but not landing) pull requests with Jenkins is unchanged, and now documented here: https://github.com/nodejs/node/wiki/Testing-pull-requests-with-Jenkins
I'll be monitoring Jenkins to make sure things keep working smoothly. Let me and @nodejs/jenkins-admins know if you encounter any issues with the CI infrastructure or to share your feedback. Thank you!