ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 18

  1. Blog
  2. /Observability & SRE

fnmatch vs re Name Filter: Localhost Lab

Hands-on fnmatch.filter vs re.compile name-list filtering measured on Linux localhost.

Aditya Challa·30 September 2026·4 min read

Lab
On this page
  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — names/s (p50, millions)
  5. Multi-pattern (any of `*.py` / `*.json` / `test_*`)
  6. Reading it
  7. Convenience vs hot path
  8. translate() note
  9. Batch vs per-name
  10. Match counts sanity
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway

Intro — what this post promises

Filter a large list of basenames with shell globs (fnmatch.filter / fnmatchcase) vs re.compile. This lab reports names/s on ~52700 synthetic names on Linux localhost.

It is not a filesystem walk (glob vs rglob vs walk) and not general string search (re vs str). Patterns run against in-memory names only.

Related links:

  • glob vs rglob vs walk localhost lab
  • python re vs str localhost lab
  • scandir vs listdir localhost lab
  • futures as completed vs wait localhost lab
  • tarfile vs zipfile localhost lab
  • str translate vs replace localhost lab
  • textwrap fill vs manual localhost lab
  • secrets vs urandom localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. No Docker. Regexes are hand-translated equivalents of the globs (\Z-anchored).

Verdict up front (*.py): fnmatch.filter ~9.03 Mnames/s vs re.match ~8.28 M vs re.search ~3.07 M. Multi-glob any-of: re alternation ~1.95 M/s vs fnmatch loop ~1.65 M/s.


Arms

ArmPattern
fnmatch.filter(names, pat)C-accelerated batch filter
fnmatchcase loopper-name case-sensitive
re.compile(...).matchanchored regex
re.compile(...).searchsearch (can be costlier with .*)
multiseveral globs vs one alternation

Lab topology

~52700 basenames · 7 rounds · p50
patterns: *.py, test_*, *.log, app_*.json, pkg.module_*.py
metric: names/s = n / p50_s

Script: lab-evidence/93-fnmatch-vs-re/results/run_lab.py.


Lead table — names/s (p50, millions)

Patternfnmatch.filterfnmatchcasere.matchre.search
*.py9.034.228.283.07
test_*11.624.9210.249.86
app_*.json13.075.3112.6112.11
pkg.module_*.py12.565.2412.6213.96

Multi-pattern (any of *.py / *.json / test_*)

ArmMnames/s
fnmatch any (loop)1.65
re alternation1.95

Reading it

  • fnmatch.filter is the convenience winner for one shell glob over a list — often tied with or ahead of a careful re.match.
  • Per-name fnmatchcase loops pay Python call overhead (~2× slower than filter here).
  • re.search with leading .* can lose to match / fnmatch — compile thoughtfully.
  • Many globs: one compiled alternation beat a Python any-of fnmatch loop (~1.19×).

Convenience vs hot path

Use fnmatch.filter in tools and one-off filters — readable and fast enough. Reach for re.compile when patterns are dynamic, numerous, or already regex; keep the compiled object outside the loop.


translate() note

fnmatch.translate can turn a glob into a regex string for re.compile. This lab used hand-written equivalents for clarity. If you generate regex from globs at runtime, compile once outside the name loop — the compile tax will otherwise dominate.


Batch vs per-name

fnmatch.filter walks the list in C and appends matches — that is why it beat a Python fnmatchcase loop by ~2× on every pattern here. If you already iterate names for other reasons, fnmatchcase inline can still be fine; for “give me all *.py,” call filter.


Match counts sanity

All arms agreed on match counts per pattern (e.g. *.py and app_*.json hits matched across fnmatch and re). Throughput differences are engine cost, not divergent semantics on these fixtures.


Pitfalls

  • Translating globs to regex incorrectly (forgetting to escape dots).
  • Using re.search(r".*\.py") without \Z / fullmatch semantics.
  • Walking the filesystem with fnmatch on every os.walk name without measuring — different problem than this list filter.
  • Case folding: fnmatch vs fnmatchcase differ on case-insensitive platforms.

Reproduce

python3 lab-evidence/93-fnmatch-vs-re/results/run_lab.py

Evidence: summary.json, fixture_names.txt sample.


Limits

One Linux box. Synthetic basenames. Not pathlib.match / recursive glob. Regexes hand-written, not fnmatch.translate microbench.


Takeaway

On ~52700 names, fnmatch.filter ~9.03 Mnames/s for *.py matched re.match ~8.28 M. For multi-glob, compiled alternation ~1.95 M/s edged the fnmatch any-loop. Default to fnmatch for shell-shaped filters; compile re when the hot path is many patterns or regex features.

fnmatch.filterfnmatchcasere.compileshell globname list filterlocalhost labsrepython

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Python 3.13.5. n≈52700. *.py: fnmatch.filter 9.03 Mnames/s; re.match 8.28; re.search 3.07. Multi: re_alt 1.95 vs fnmatch_any 1.65. Not glob walk / re-vs-str string labs. Affiliates: 0. Evidence: lab-evidence/93-fnmatch-vs-re/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 84

    tarfile vs zipfile Create+Extract: Localhost Lab

    Hands-on tarfile vs zipfile create+extract lab: real MB/s on a mixed small-file fixture (uncompressed tar vs zip), measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

  • Plate 85

    scandir vs listdir vs iterdir: Localhost Lab

    Hands-on os.scandir vs listdir vs Path.iterdir lab: real entries/s for names and is_file on a synthetic tree, measured on Linux localhost (lab) for SREs.

    Observability & SRE · 30 Sept 2026

  • Plate 28

    glob vs rglob vs os.walk: Listing Lab

    Hands-on glob.glob vs Path.rglob vs os.walk lab: real files/s for recursive file listing on a modest fixture tree, measured on Linux localhost for SREs.

    Observability & SRE · 30 Sept 2026

On this page

  1. Intro — what this post promises
  2. Arms
  3. Lab topology
  4. Lead table — names/s (p50, millions)
  5. Multi-pattern (any of `*.py` / `*.json` / `test_*`)
  6. Reading it
  7. Convenience vs hot path
  8. translate() note
  9. Batch vs per-name
  10. Match counts sanity
  11. Pitfalls
  12. Reproduce
  13. Limits
  14. Takeaway
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove