Plate 85
scandir vs listdir vs iterdir: Localhost Lab
Hands-on os.scandir vs listdir vs Path.iterdir lab: real entries/s for names and is_file on a synthetic tree, measured on Linux localhost (lab) for SREs.
Aditya Challa4 min read
Intro — what this post promises
Listing a directory: os.scandir, os.listdir, or pathlib.Path.iterdir? This lab measures entries/s on a synthetic tree (flat + nested) on Linux localhost, then shows why DirEntry.is_file() beats listdir + os.path.isfile.
It is not a remake of the glob vs rglob vs walk localhost lab (pattern matching / recursive globs). Here the question is directory iterator APIs and when cached DirEntry metadata wins.
Related links:
- glob vs rglob vs walk localhost lab
- mmap vs read scan localhost lab
- bytesio vs spooled tempfile localhost lab
- shutil copyfile vs manual localhost lab
- pathlib vs ospath localhost lab
- tarfile vs zipfile localhost lab
- secrets vs urandom localhost lab
- textwrap fill vs manual localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Fixtures under lab-evidence/90-scandir-vs-listdir/results/fixture_tree/. Affiliates: 0. No Docker.
Verdict up front (2000 entries, is_file path): scandir ~3056.2 k/s vs listdir+isfile ~353.7 k/s (~8.64×) vs Path.iterdir().is_file ~242.4 k/s. For names only at 8000 files, listdir ~4288.1 k/s edged scandir ~3279.9 k/s; iterdir ~1079.8 k/s.
Arms
| Arm | Pattern |
|---|---|
os.listdir | names only |
listdir + join | build full paths |
os.scandir | iterate DirEntry |
Path.iterdir | pathlib objects |
listdir + isfile | stat each path |
DirEntry.is_file() | cached type from scandir |
| nested | os.walk vs scandir stack vs rglob("*") |
Lab topology
Script: lab-evidence/90-scandir-vs-listdir/results/run_lab.py.
Lead table — flat names (p50 entries/s)
| n | listdir | listdir+join | scandir | Path.iterdir |
|---|---|---|---|---|
| 500 | 4187.5k | 1752.6k | 3563.5k | 928.0k |
| 2000 | 4380.1k | 2136.0k | 3374.3k | 1046.3k |
| 8000 | 4288.1k | 2119.9k | 3279.9k | 1079.8k |
When you need is_file (n=2000)
| Arm | entries/s |
|---|---|
scandir DirEntry.is_file() | 3056.2k |
listdir + os.path.isfile | 353.7k |
Path.iterdir + is_file() | 242.4k |
~8.64× — this is the scandir headline. Extra stat per name is the tax DirEntry avoids when the type is already known from the directory read.
Nested walk (2050 entries)
| Arm | entries/s |
|---|---|
| scandir stack walk | 2434.0k |
os.walk | 2112.6k |
Path.rglob("*") | 547.2k |
Reading it
- Names only:
listdiris fine (and slightly fastest here). - Need type/stat: prefer
scandir/DirEntry. Path.iterdiris ergonomic; expect a measurable object-creation tax on huge directories.- Use
os.walkor a scandir stack for trees;rglob("*")was ~4.45× slower on this nest.
Names-only vs metadata
If you only print names, listdir avoids building DirEntry objects and won slightly here. The moment you call isfile/isdir/stat on each name, you pay syscalls that scandir often already satisfied. Rule of thumb from this box: branching on type → scandir; names dump → listdir is acceptable.
pathlib tax
Path.iterdir returns Path objects — great for chaining .suffix / .with_name. On an 8000-file directory that object churn showed up as ~4.0× slower than raw listdir. Use pathlib at the edges; tight loops may want scandir/listdir.
Pitfalls
- Calling
os.path.isfileon everylistdirname in a hot path. - Assuming scandir is always fastest for names-only (not on this box).
- Forgetting
with os.scandir(...)(resource lifetime). - Following symlinks unintentionally in recursive walks.
Reproduce
Evidence: summary.json, fixture_tree/.
Limits
One Linux box, synthetic files on local FS. Not networked filesystems / huge directories with millions of entries. Cold-cache dentries may differ.
Takeaway
For name listing, listdir ~4288.1 k/s ≈ scandir ~3279.9 k/s at 8000 files. For is_file filtering, scandir ~8.64× beat listdir+stat. Default to scandir when you branch on file type; keep listdir for simple name dumps.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Python 3.13.5. Flat 8000 names: listdir 4288.1k/s; scandir 3279.9k/s; iterdir 1079.8k/s. is_file n=2000: scandir 3056.2k/s vs listdir+isfile 353.7k/s (~8.64x). Affiliates: 0. Evidence: lab-evidence/90-scandir-vs-listdir/.
Related links
Plate 18
fnmatch vs re Name Filter: Localhost Lab
Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.
Observability & SRE · 30 Sept 2026
Plate 84
tarfile vs zipfile Create+Extract: Localhost Lab
Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.
Observability & SRE · 30 Sept 2026
Plate 28
glob vs rglob vs os.walk: Listing Lab
Hands-on glob.glob vs Path.rglob vs os.walk lab: real files/s for recursive file listing on a modest fixture tree, measured on Linux localhost for SREs.
Observability & SRE · 30 Sept 2026