Plate 35
perf_counter vs time vs monotonic Lab
A hands-on Linux localhost lab measuring Python clock-call overhead, observed resolution, and safe choices for elapsed-time measurements.
Aditya Challa4 min read
Intro — what this post promises
Which clock should you call in a hot path: time.perf_counter, time.monotonic, or time.time? This lab measures ns per call on Linux localhost and records get_clock_info resolution plus the smallest positive back-to-back delta we saw.
Related links:
- lru_cache hit vs miss localhost lab
- logging vs print localhost lab
- deque vs list queue localhost lab
- itertools vs python loops localhost lab
- Why your average latency graph is lying (p50 / p95 / p99)
- copy vs deepcopy localhost lab
- struct pack vs to_bytes localhost lab
- set vs list membership localhost lab
Lab honesty (1 Oct 2026 IST): Python 3.13.5. 500,000 calls/batch. Affiliates: 0. Overhead includes the for loop; we also report net vs an empty xor loop. Resolution claims are observed min positive delta, not a kernel guarantee under load.
Verdict up front: perf_counter / monotonic / time are peers at ~58–60 ns/call. process_time is the outlier (~278 ns). Empty loop baseline ~25 ns → net timer cost ~33 ns. Prefer perf_counter for intervals; avoid time() for durations (wall clock can jump).
Clocks under test
| API | Role |
|---|---|
perf_counter / _ns | high-res monotonic intervals |
monotonic / _ns | monotonic (here same CLOCK_MONOTONIC) |
time / _ns | wall clock |
process_time | CPU time |
Lab topology
Script: lab-evidence/57-perf-counter-vs-time/results/run_lab.py.
Clock info (this box)
| Clock | resolution | monotonic | implementation |
|---|---|---|---|
| time | 1e-09 | False | clock_gettime(CLOCK_REALTIME) |
| monotonic | 1e-09 | True | clock_gettime(CLOCK_MONOTONIC) |
| perf_counter | 1e-09 | True | clock_gettime(CLOCK_MONOTONIC) |
| process_time | 1e-09 | True | clock_gettime(CLOCK_PROCESS_CPUTIME_ID) |
Reported resolution is 1 ns for all four via clock_gettime. That is the API’s claimed resolution — not “every call returns a new nanosecond.”
Lead table — call overhead (p50)
| Arm | calls/s | ns/call | net vs empty (ns) |
|---|---|---|---|
| empty xor loop | 40,196,923 | 24.9 | 0 |
| perf_counter | 17,293,537 | 57.8 | 32.9 |
| monotonic | 17,245,707 | 58.0 | 33.1 |
| time | 16,751,351 | 59.7 | 34.8 |
| perf_counter_ns | 16,054,411 | 62.3 | 37.4 |
| monotonic_ns | 16,171,901 | 61.8 | 37.0 |
| time_ns | 16,185,068 | 61.8 | 36.9 |
| process_time | 3,592,895 | 278.3 | 253.4 |
| perf_counter pair Δ | 17,167,531 | 58.2 | 33.4 |
Resolution caveats
Observed minimum positive back-to-back delta (10k trials):
| Clock | min +Δ |
|---|---|
| perf_counter (s) | 4.599860403686762e-08 |
| monotonic (s) | 4.3997715692967176e-08 |
| time (s) | 2.384185791015625e-07 |
| perf_counter_ns | 46 ns |
| time_ns | 48 ns |
So: do not treat sub-100 ns interval measurements as meaningful on this box — call overhead and clock grain are in that band. For microbenchmarks, batch work and divide, or accept that noise floors matter.
time() can step backward/forward with NTP; never use it for elapsed = t1 - t0 in production timers. perf_counter (or monotonic) is the interval API.
Pitfalls
- Timing a single call — overhead ≈ the thing you timed.
- Using
time.timefor durations — wall clock, adjustable. - Trusting
resolution=1e-9as usable precision — see min +Δ. process_timefor wall latency — ignores sleep/IO wait; also slower to call here.
When to pick what
| Need | Prefer |
|---|---|
| Elapsed wall intervals | time.perf_counter |
| Deadlines immune to NTP | time.monotonic |
| Logs / absolute timestamps | time.time / time_ns |
| CPU-only profiling | process_time |
Reproduce
Evidence: /workspace/lab-evidence/57-perf-counter-vs-time/results/.
Closing
Timer calls cost tens of nanoseconds. On this box perf_counter ≈ monotonic ≈ time at ~58 ns/call; process_time ~278 ns. Net after an empty loop ~33 ns. Use perf_counter for intervals; batch when measuring work near that floor.
Lab evidence
What I found running this
Lab 1 Oct 2026 IST. Ran Python 3.13.5 with 500,000 calls per batch. Empty xor loop 24.9 ns; perf_counter 57.8 ns/call, monotonic 58.0, time 59.7, process_time 278. Net perf_counter cost about 33 ns. min positive perf_counter_ns delta 46 ns. Affiliates: 0. Evidence: lab-evidence/57-perf-counter-vs-time/.
Related links
Plate 17
platform vs os.uname Inventory: Localhost Lab
Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 75
uuid.uuid4 vs uuid.uuid1: Localhost Lab
Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026
Plate 50
signal vs threading.Event Wakeup: Localhost Lab
Hands-on signal SIGUSR1 vs threading.Event wakeup lab: real p50 latency in microseconds, measured on Linux localhost today in this hands-on lab for SREs.
1 Oct 2026