Plate 90
Path.read_text vs open().read: Localhost Lab
Aditya Challa3 min read
Intro — what this post promises
Read whole files with Path.read_text vs open(...).read(). This lab reports reads/s on Linux localhost for a small text file and a ~100 KB medium file — plus a binary read_bytes peer.
Related links:
- copy copy vs dict copy localhost lab
- html parser vs regex localhost lab
- xml etree vs json localhost lab
- logging formatter vs fstring localhost lab
- dataclass asdict vs vars localhost lab
- tracemalloc snapshot localhost lab
- gc collect cost localhost lab
- sqlite3 vs shelve localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Affiliates: 0. Differentiates from pathlib-vs-ospath (lab 47) — that post compared path joining/stat style APIs; this one is full-file read throughput.
Verdict up front: small (500 B, n=5000): Path.read_text ~97170 reads/s vs open ~97062. Medium (110890 B, n=1000): Path ~17491 vs open ~15493.
Arms
| Arm | Pattern |
|---|---|
Path.read_text(encoding=) | pathlib helper |
with open(...) as f: f.read() | classic |
| same on medium file | size sensitivity |
read_bytes / open(...,"rb") | binary peer |
Seven rounds, p50. Texts equal across APIs (equal=True).
Lab topology
Script: lab-evidence/132-path-read-text-vs-open/results/run_lab.py.
Lead table (p50 reads/s)
| Arm | reads/s |
|---|---|
| Path.read_text small | 97170 |
| open read text small | 97062 |
| Path.read_text med | 17491 |
| open read text med | 15493 |
| Path.read_bytes small | 180811 |
| open rb small | 183833 |
On small text the APIs were essentially tied. On the medium file, Path.read_text edged open. Binary reads roughly doubled throughput vs UTF-8 decode on the small file.
Reading it for SRE work
- Agent/config loaders → prefer
Path.read_textfor clarity; do not fear a hidden tax on this box. - Hot binary blobs →
read_bytes/"rb"and decode only what you need. - Lab 47 still owns join/
os.pathdebates; this post is whole-file read cost. - Always pass
encoding=explicitly for text — locale surprises are worse than microbench noise.
Document encoding in the runbook so a “fast open() refactor” does not drop UTF-8 assumptions.
Size sensitivity
At 110890 bytes, Path landed ~17491 reads/s vs open ~15493. Decode and syscall dominate; the Path wrapper is not the villain. If you repeatedly reread the same config, cache the string — API choice will not save you.
Binary peer
read_bytes / "rb" hit about ~180811 / ~183833 reads/s on the small file — both skip Unicode decode. Use binary when the payload is opaque; do not read_text a protobuf.
Operational habit
For one-shot config loads at process start, either API is fine — the ~97170 reads/s class of results will not be your bottleneck. Standardize on Path in new code so reviews stop bikeshedding open vs Path, and reserve energy for caching, validation, and encoding errors.
Pitfalls
- Omitting encoding and depending on locale.
- Using read_text in a tight loop over multi-MB files without mmap/chunking.
- Comparing cold-cache first read to warm-cache microbench.
- Conflating with lab 47 path-construction timings.
Reproduce
Evidence: summary.json, summary.txt, sample files beside the script.
Limits
One Linux box, local disk, warm cache. Not NFS, not concurrent readers.
Takeaway
Path.read_text and open().read() tracked closely on small files (~97170 vs ~97062 reads/s); Path led on the medium text file. Prefer Path for readable loaders; use binary reads when you do not need decode.
Lab evidence
What I found running this
Ran lab-evidence/132-path-read-text-vs-open/results/run_lab.py on Linux localhost with Python 3.13.5 on 1 Oct 2026 IST. Seven rounds, p50: small 500 B n=5000 Path.read_text ~97170 reads/s vs open ~97062; medium 110890 B n=1000 Path ~17491 vs open ~15493; binary small read_bytes ~180811 vs open rb ~183833. Text outputs equal=True; warm local cache; affiliates 0.