Plate 98
lru_cache hit vs miss: Cache Cost Lab
Real Linux localhost timings for functools.lru_cache hits, miss-and-insert calls, maxsize choices, and an uncached baseline.
Aditya Challa4 min read
Intro — what this post promises
functools.lru_cache on a cheap function is not free. This lab times hits, miss-and-insert, and a plain uncached call on Linux localhost, and varies maxsize.
Related links:
- copy vs deepcopy localhost lab
- logging vs print localhost lab
- set vs list membership localhost lab
- string concat vs join localhost lab
- dataclass vs slots vs dict localhost lab
- json vs orjson vs msgpack localhost lab
- sorted vs heapq vs bisect localhost lab
- python re vs str methods localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. Function is (n * 17) ^ (n >> 3) — CPU only, no I/O. Affiliates: 0. Hits use a key set that fits in the cache. Misses pass new integers so every call inserts (and may evict).
Verdict up front: uncached is ~11.35M/s (88.1 ns). A hit at maxsize 128 is ~17.12M (58.4 ns, ~1.51×). A miss-insert is slower than uncached: ~6.17M (162 ns). Cache the expensive call, not the cheap one.
Arms
| Arm | Meaning |
|---|---|
| uncached | bare function, 200k calls |
| hit | prime key_mod keys, then call i % key_mod |
| miss_insert | cache_clear, then 200k new keys |
| maxsize=0 | documented “don’t cache”; wrapper still runs |
maxsize=None is unbounded. Hit key space is 8 when maxsize is 8, else 64.
Lab topology
Script: lab-evidence/52-lru-cache-hit-miss/results/run_lab.py.
Lead table — ops/s (p50)
| Arm | ops/s | ns/op | vs uncached |
|---|---|---|---|
| hit maxsize 8 | 18,632,399 | 53.7 | 1.64× |
| hit maxsize None | 17,528,299 | 57.1 | 1.54× |
| hit maxsize 1024 | 17,431,681 | 57.4 | 1.54× |
| hit maxsize 128 | 17,123,732 | 58.4 | 1.51× |
| uncached | 11,354,499 | 88.1 | 1.00× |
| maxsize 0 | 9,630,461 | 103.8 | 0.85× |
| miss maxsize 128 | 6,172,613 | 162.0 | 0.54× |
| miss maxsize 1024 | 6,002,194 | 166.6 | 0.53× |
| miss maxsize None | 5,420,340 | 184.5 | 0.48× |
| miss maxsize 8 | 3,873,678 | 258.2 | 0.34× |
What the curve says
- Hits are only ~1.5–1.6× the bare call because the function is already ~88 ns. The cache lookup itself is tens of nanoseconds. If your real function is milliseconds (SQL, parse, HTTP), that overhead vanishes.
- Miss-and-insert is a tax. At maxsize 128 it is ~0.54× uncached. You pay hash, store, and the function.
- Tiny maxsize + unique keys is the worst: maxsize 8 miss is ~258 ns — every insert evicts.
maxsize=Noneon a miss storm keeps every key (currsize200k). Slightly slower than a bounded cache and it grows RAM.maxsize=0is slower than no decorator (103.8 ns vs 88.1 ns). It does not cache; it still wraps.
Pitfalls
- Caching a 100 ns function “just in case” — a miss makes it slower.
- Unbounded cache on user input — memory leak with a decorator.
- Mutable args / unhashable — lru_cache will throw; this lab uses ints only.
- Assuming hit ratio from folklore — call
cache_info().
When to pick what
| Need | Prefer |
|---|---|
| Expensive pure function, small key set | lru_cache with a real maxsize |
| Cheap arithmetic | no cache |
| Untrusted / unbounded keys | bounded maxsize or no cache |
| “Disable cache in tests” | cache_clear(), not maxsize=0 in prod hot path |
Reproduce
Evidence: /workspace/lab-evidence/52-lru-cache-hit-miss/results/.
Closing
On this cheap function, a hit is ~58 ns vs ~88 ns uncached, and a miss-insert is ~162 ns (worse at maxsize 8: ~258 ns). maxsize does not change hit speed once the key fits; it changes eviction and RAM. Cache work that is actually expensive.
Lab evidence
What I found running this
Ran python3 lab-evidence/52-lru-cache-hit-miss/results/run_lab.py on 1 Oct 2026 IST with Python 3.13.5 on Linux localhost. Repeated 200k calls and compared uncached, cache hits, miss-and-insert, maxsize 0, and bounded/unbounded caches. Measured p50 wall time: uncached 11.35M/s (88.1 ns), hit maxsize 128 17.12M/s (58.4 ns), miss-insert 6.17M/s (162 ns), and maxsize 8 misses 3.87M/s (258 ns). The surprising result was that caching this cheap function makes misses slower and maxsize 0 still adds wrapper overhead.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026