ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 35

  1. Blog

perf_counter vs time vs monotonic Lab

A hands-on Linux localhost lab measuring Python clock-call overhead, observed resolution, and safe choices for elapsed-time measurements.

Aditya Challa·30 September 2026·4 min read

Summary
On this page
  1. Intro — what this post promises
  2. Clocks under test
  3. Lab topology
  4. Clock info (this box)
  5. Lead table — call overhead (p50)
  6. Resolution caveats
  7. Pitfalls
  8. When to pick what
  9. Reproduce
  10. Closing

Intro — what this post promises

Which clock should you call in a hot path: time.perf_counter, time.monotonic, or time.time? This lab measures ns per call on Linux localhost and records get_clock_info resolution plus the smallest positive back-to-back delta we saw.

Related links:

  • lru_cache hit vs miss localhost lab
  • logging vs print localhost lab
  • deque vs list queue localhost lab
  • itertools vs python loops localhost lab
  • Why your average latency graph is lying (p50 / p95 / p99)
  • copy vs deepcopy localhost lab
  • struct pack vs to_bytes localhost lab
  • set vs list membership localhost lab

Lab honesty (1 Oct 2026 IST): Python 3.13.5. 500,000 calls/batch. Affiliates: 0. Overhead includes the for loop; we also report net vs an empty xor loop. Resolution claims are observed min positive delta, not a kernel guarantee under load.

Verdict up front: perf_counter / monotonic / time are peers at ~58–60 ns/call. process_time is the outlier (~278 ns). Empty loop baseline ~25 ns → net timer cost ~33 ns. Prefer perf_counter for intervals; avoid time() for durations (wall clock can jump).


Clocks under test

APIRole
perf_counter / _nshigh-res monotonic intervals
monotonic / _nsmonotonic (here same CLOCK_MONOTONIC)
time / _nswall clock
process_timeCPU time

Lab topology

N = 500000 calls / repeat, 7 repeats, p50
Also: empty xor loop; t1-t0 pair pattern
get_clock_info + min positive back-to-back delta

Script: lab-evidence/57-perf-counter-vs-time/results/run_lab.py.


Clock info (this box)

Clockresolutionmonotonicimplementation
time1e-09Falseclock_gettime(CLOCK_REALTIME)
monotonic1e-09Trueclock_gettime(CLOCK_MONOTONIC)
perf_counter1e-09Trueclock_gettime(CLOCK_MONOTONIC)
process_time1e-09Trueclock_gettime(CLOCK_PROCESS_CPUTIME_ID)

Reported resolution is 1 ns for all four via clock_gettime. That is the API’s claimed resolution — not “every call returns a new nanosecond.”


Lead table — call overhead (p50)

Armcalls/sns/callnet vs empty (ns)
empty xor loop40,196,92324.90
perf_counter17,293,53757.832.9
monotonic17,245,70758.033.1
time16,751,35159.734.8
perf_counter_ns16,054,41162.337.4
monotonic_ns16,171,90161.837.0
time_ns16,185,06861.836.9
process_time3,592,895278.3253.4
perf_counter pair Δ17,167,53158.233.4

Resolution caveats

Observed minimum positive back-to-back delta (10k trials):

Clockmin +Δ
perf_counter (s)4.599860403686762e-08
monotonic (s)4.3997715692967176e-08
time (s)2.384185791015625e-07
perf_counter_ns46 ns
time_ns48 ns

So: do not treat sub-100 ns interval measurements as meaningful on this box — call overhead and clock grain are in that band. For microbenchmarks, batch work and divide, or accept that noise floors matter.

time() can step backward/forward with NTP; never use it for elapsed = t1 - t0 in production timers. perf_counter (or monotonic) is the interval API.


Pitfalls

  1. Timing a single call — overhead ≈ the thing you timed.
  2. Using time.time for durations — wall clock, adjustable.
  3. Trusting resolution=1e-9 as usable precision — see min +Δ.
  4. process_time for wall latency — ignores sleep/IO wait; also slower to call here.

When to pick what

NeedPrefer
Elapsed wall intervalstime.perf_counter
Deadlines immune to NTPtime.monotonic
Logs / absolute timestampstime.time / time_ns
CPU-only profilingprocess_time

Reproduce

python3 lab-evidence/57-perf-counter-vs-time/results/run_lab.py

Evidence: /workspace/lab-evidence/57-perf-counter-vs-time/results/.


Closing

Timer calls cost tens of nanoseconds. On this box perf_counter ≈ monotonic ≈ time at ~58 ns/call; process_time ~278 ns. Net after an empty loop ~33 ns. Use perf_counter for intervals; batch when measuring work near that floor.

time.perf_countertime.timetime.monotonictimer overheadpython timinglocalhost labsrebenchmark

Lab evidence

What I found running this

Lab 1 Oct 2026 IST. Ran Python 3.13.5 with 500,000 calls per batch. Empty xor loop 24.9 ns; perf_counter 57.8 ns/call, monotonic 58.0, time 59.7, process_time 278. Net perf_counter cost about 33 ns. min positive perf_counter_ns delta 46 ns. Affiliates: 0. Evidence: lab-evidence/57-perf-counter-vs-time/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 17

    platform vs os.uname Inventory: Localhost Lab

    Hands-on platform.platform vs os.uname host inventory lab: real ops/s plus cache notes, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 75

    uuid.uuid4 vs uuid.uuid1: Localhost Lab

    Hands-on uuid.uuid4 vs uuid.uuid1 ID generation lab: real ops/s plus version/node checks, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

  • Plate 50

    signal vs threading.Event Wakeup: Localhost Lab

    Hands-on signal SIGUSR1 vs threading.Event wakeup lab: real p50 latency in microseconds, measured on Linux localhost today in this hands-on lab for SREs.

    1 Oct 2026

On this page

  1. Intro — what this post promises
  2. Clocks under test
  3. Lab topology
  4. Clock info (this box)
  5. Lead table — call overhead (p50)
  6. Resolution caveats
  7. Pitfalls
  8. When to pick what
  9. Reproduce
  10. Closing
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove