Repository navigation
A proposal to add fs.scandir method to FS module #15699
Description
Activity
- addedfeature requestIssues requesting new Node.js features.Issues requesting new Node.js features.fsIssues and PRs related to file-system APIs and the fs module.Issues and PRs related to file-system APIs and the fs module.
on Sep 30, 2017 - changed the title
[-]Ability to return d_name and d_type from fs.readdir[/-][+]A proposal to add fs.scandir method to FS module[/+]on Oct 1, 2017 I think if we're going to introduce
scandir*()methods, they should work more or less exactly the same as the underlying C function of the same name. One benefit of such a function over thereaddir*()methods is that there would be no forced buffering of entry names, which is nice for very large directories. To improve performance, we could allow an option that dictates how many directories to buffer before calling out to the JS callbacks.For context: libuv/libuv#416 - stalled, and itself a continuation of an older, also stalled PR.
- addedlibuvIssues and PRs related to the libuv dependency or the uv binding.Issues and PRs related to the libuv dependency or the uv binding.
on Oct 5, 2017 Speaking from the wild (20+ years of syseng experience): large directories are a recurring ops problem I use as an interview question because I have seen it multiple times, and as recently as 2013. Even at companies like Disney and Amazon, there are developers who don't see a problem with using a directory as a flat key-value store, or creating empty files and never removing them. Eventually an operator like me has to do something with hundreds of millions of files all in the same directory. In the past I've used Perl to deal with it, but I fell out of love with Perl years ago.
The usual tools are usually useless specifically because (as I found out with strace) they stat each entry. While stat itself isn't super expensive, it ends up driving the memory and CPU cost of a scan higher than it needs to be if there's enough information in the file name to make decisions from.
(Though while writing this I discovered 'ls -1' uses mmap and does not perform any stat()s.)
If/when y'all decide to tackle this issue, I would ask that you also offer fs.*dir* so that I can do precisely what I want to:
cleanUp = (problemDir, prefix) -> new Promise (resolve, reject) -> fs.opendir problemDir, (err, handle) -> reject err do next = -> handle.readdir (err, stat) -> switch when err then reject err when not stat then resolve() when stat.isDirectory then next() when (fileName = stat.name).startsWith prefix fullPath = path.resolve problemDir, fileName fs.unlink fullPath, (err) -> if err then reject err else next() else next() cleanUp 'incoming', 'system.system' .then -> console.log 'Completed' .catch (err) -> console.log err
That said, this kind of feature probably has a small audience and third-party
modules exist which addess
he problem, so I woundn't blame you if this never became a high priority.Reacted by Joshua Wiens, Alexander Mills, Marvin Hagemeister, J. S. Choi and coderaiserReacted by Wout MertensAnother not mutually exclusive option is to add a new type of stream that emits
scandir-like entries or full-blownstatsReacted by Wout Mertens, Seth Holladay and Michał WadasNot essential but it would be nice to have an option to produce a recursive scan. I guess it wouldn’t follow symlinks though to avoid falling into a closed loop
Reacted by Alexander MillsThis would be great to have.
I made a very simple parallel find implementation to see if I could use node's async nature to keep a higher queue depth. I'm not sure, but I really think the lack of exposure to dirent- specifically that inability to see that an entry is a directory- is what keeps Node from trouncing find.
I went looking into libuv and found that libuv does expose the dirent type, &c, it's just Node's readdir that is limiting. Then I found this thread.
scandirwould help my immediate problem.But also: having access to things like inodes, extents, &c would open up Node's viability for a wide array of system tools. I'd expect 95% of use cases to be recursing directory trees, but I wanted to call out that there's a lot of other good helpful systems stuff that scandir does.
Proposal: add an option to fs.readdir to get full directory entities. Scandir is useful for recursing, but there's still really good useful things that can be done with readdir(3) that libuv permits, but Node doesn't. I'd love to see dirents results available for Node's readdir!
- addedblockedPRs that are blocked by other issues or PRs.PRs that are blocked by other issues or PRs.
on Jul 12, 2018 - added a commit that references this issue
on Aug 19, 2018 - added a commit that references this issue
on Sep 3, 2018 This was fixed in #22020.
Problem
Now any interaction with files and directories in the File System is as follows:
The problem here is that we are call File System a second time due to the fact that we don't know the directory in front of us, or file (or symlink).
But we can reduce twice File System calls by creating
fs.scandirmethod that can returnd_nameandd_type. This information is returned fromuv_dirent_t(scandir) (libuv). For example, this is implemented in theLuvitandpyuv(also uselibuv).Motivation
String→Object→d_name+d_typefs.readdir: return convertedObjecttoStringfs.scandir: return As isnode-glob(also for each package that usesfs.readdirfor traversing directories) in most cases (needfs.statwhend_typeis aDT_UNKNOWNon the old FS)Proposed solution
Add a methods
fs.scandirandfs.scandirSyncin a standard Fyle System module.Where
entriesis an array of objects:name{String} – filename (d_nameinlibuv)type{Number} –fs.constants.S_*(d_typeinlibuv)Final words
Now I solved this problem by creating C++ Addon but... But I don't speak C++ or speak but very bad (i try 😅 ) and it requires you compile when you install a package that requires additional manipulation to the end user (like https://github.com/nodejs/node-gyp#on-windows).