ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 85

  1. Blog
  2. /Observability & SRE

scandir vs listdir vs iterdir: Localhost Lab

Hands-on os.scandir vs listdir vs Path.iterdir lab: real entries/s for names and is_file on a synthetic tree, measured on Linux localhost (lab) for SREs.

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — flat names (p50 entries/s)
  5. When you need `is_file` (n=2000)
  6. Nested walk (2050 entries)
  7. Reading it
  8. Names-only vs metadata
  9. pathlib tax
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway

Intro — what this post promises

Listing a directory: os.scandir, os.listdir, or pathlib.Path.iterdir? This lab measures entries/s on a synthetic tree (flat + nested) on Linux localhost, then shows why DirEntry.is_file() beats listdir + os.path.isfile.

It is not a remake of the glob vs rglob vs walk localhost lab (pattern matching / recursive globs). Here the question is directory iterator APIs and when cached DirEntry metadata wins.

Related links:

  • glob vs rglob vs walk localhost lab
  • mmap vs read scan localhost lab
  • bytesio vs spooled tempfile localhost lab
  • shutil copyfile vs manual localhost lab
  • pathlib vs ospath localhost lab
  • tarfile vs zipfile localhost lab
  • secrets vs urandom localhost lab
  • textwrap fill vs manual localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Fixtures under lab-evidence/90-scandir-vs-listdir/results/fixture_tree/. Affiliates: 0. No Docker.

Verdict up front (2000 entries, is_file path): scandir ~3056.2 k/s vs listdir+isfile ~353.7 k/s (~8.64×) vs Path.iterdir().is_file ~242.4 k/s. For names only at 8000 files, listdir ~4288.1 k/s edged scandir ~3279.9 k/s; iterdir ~1079.8 k/s.


Arms

ArmPattern
os.listdirnames only
listdir + joinbuild full paths
os.scandiriterate DirEntry
Path.iterdirpathlib objects
listdir + isfilestat each path
DirEntry.is_file()cached type from scandir
nestedos.walk vs scandir stack vs rglob("*")

Lab topology

flat: 500 / 2000 / 8000 files · 7 rounds · p50
isfile check at n=2000
nested: 50 dirs × 40 files (2050 entries)
metric: entries/s

Script: lab-evidence/90-scandir-vs-listdir/results/run_lab.py.


Lead table — flat names (p50 entries/s)

nlistdirlistdir+joinscandirPath.iterdir
5004187.5k1752.6k3563.5k928.0k
20004380.1k2136.0k3374.3k1046.3k
80004288.1k2119.9k3279.9k1079.8k

When you need is_file (n=2000)

Armentries/s
scandir DirEntry.is_file()3056.2k
listdir + os.path.isfile353.7k
Path.iterdir + is_file()242.4k

~8.64× — this is the scandir headline. Extra stat per name is the tax DirEntry avoids when the type is already known from the directory read.


Nested walk (2050 entries)

Armentries/s
scandir stack walk2434.0k
os.walk2112.6k
Path.rglob("*")547.2k

Reading it

  • Names only: listdir is fine (and slightly fastest here).
  • Need type/stat: prefer scandir / DirEntry.
  • Path.iterdir is ergonomic; expect a measurable object-creation tax on huge directories.
  • Use os.walk or a scandir stack for trees; rglob("*") was ~4.45× slower on this nest.

Names-only vs metadata

If you only print names, listdir avoids building DirEntry objects and won slightly here. The moment you call isfile/isdir/stat on each name, you pay syscalls that scandir often already satisfied. Rule of thumb from this box: branching on type → scandir; names dump → listdir is acceptable.


pathlib tax

Path.iterdir returns Path objects — great for chaining .suffix / .with_name. On an 8000-file directory that object churn showed up as ~4.0× slower than raw listdir. Use pathlib at the edges; tight loops may want scandir/listdir.


Pitfalls

  • Calling os.path.isfile on every listdir name in a hot path.
  • Assuming scandir is always fastest for names-only (not on this box).
  • Forgetting with os.scandir(...) (resource lifetime).
  • Following symlinks unintentionally in recursive walks.

Reproduce

python3 lab-evidence/90-scandir-vs-listdir/results/run_lab.py

Evidence: summary.json, fixture_tree/.


Limits

One Linux box, synthetic files on local FS. Not networked filesystems / huge directories with millions of entries. Cold-cache dentries may differ.


Takeaway

For name listing, listdir ~4288.1 k/s ≈ scandir ~3279.9 k/s at 8000 files. For is_file filtering, scandir ~8.64× beat listdir+stat. Default to scandir when you branch on file type; keep listdir for simple name dumps.

os.scandiros.listdirpath.iterdirdirentrydirectory listinglocalhost labsrepython

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. Flat 8000 names: listdir 4288.1k/s; scandir 3279.9k/s; iterdir 1079.8k/s. is_file n=2000: scandir 3056.2k/s vs listdir+isfile 353.7k/s (~8.64x). Affiliates: 0. Evidence: lab-evidence/90-scandir-vs-listdir/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 18

    fnmatch vs re Name Filter: Localhost Lab

    Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.

    Observability & SRE · 30 Sept 2026

  • Plate 84

    tarfile vs zipfile Create+Extract: Localhost Lab

    Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

  • Plate 28

    glob vs rglob vs os.walk: Listing Lab

    Hands-on glob.glob vs Path.rglob vs os.walk lab: real files/s for recursive file listing on a modest fixture tree, measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — flat names (p50 entries/s)
  5. When you need `is_file` (n=2000)
  6. Nested walk (2050 entries)
  7. Reading it
  8. Names-only vs metadata
  9. pathlib tax
  10. Pitfalls
  11. Reproduce
  12. Limits
  13. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove