Plate 87
sendfile vs Userspace Copy: Localhost TCP File Serve Lab
Hands-on sendfile vs read+sendall lab: 256 MiB p50 4.41 vs 2.34 GB/s (~1.9×); copy collapses at 4 KiB. Warm localhost TCP measured, no Docker.
Aditya Challa5 min read
Intro — what this post promises
Serving a file over a socket, folklore says sendfile skips a userspace copy that read + write/send pays for. The mmap lab already warned not to confuse page-touch scans with the kernel send path. This lab measures the socket path.
This is a hands-on lab with measured numbers:
- Localhost TCP serve of 16 / 64 / 256 MiB files via
os.sendfilevsread(chunk)+sendall. - Chunk-size sensitivity on the copy arm (4 KiB / 64 KiB / 1 MiB).
- Why sendfile’s “chunk” does not matter the same way.
- Warm-cache / loopback honesty — what these GB/s are not.
Related links:
- mmap vs read (+ O_DIRECT) localhost lab
- Unix Domain Socket vs TCP localhost lab
- HTTP Keep-Alive vs Connection: close lab
- Why your average latency graph is lying (p50 / p95 / p99)
Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Server and client on 127.0.0.1 only. Payloads under /workspace (overlay), warm page cache. Protocol: 8-byte length header + body drain. No Docker. No public bind. No splice/io_uring. Affiliates: 0.
Verdict up front: at 256 MiB / 64 KiB, sendfile 4.41 GB/s beat copy 2.34 GB/s (~1.9×). The copy arm fell to ~1.0 GB/s at 4 KiB chunks; sendfile stayed near ~3 GB/s.
What we compared
| Arm | Shape in this lab |
|---|---|
| sendfile | os.sendfile(out_fd, in_fd, offset, count) until EOF |
| userspace copy | file.read(chunk) then socket.sendall(buf) |
Related links:
Lab topology
Arm A — size sweep @ 64 KiB chunks (p50)
| Size | sendfile GB/s | copy GB/s | Ratio |
|---|---|---|---|
| 16 MiB | 2.60 | 1.83 | ~1.42× |
| 64 MiB | 2.93 | 1.68 | ~1.74× |
| 256 MiB | 4.41 | 2.34 | ~1.88× |
Larger transfers amortize setup and favor the kernel path. Even on loopback — where the NIC is fake and the page cache is hot — skipping the Python bytes round-trip still showed up.
Arm B — chunk sweep @ 64 MiB (p50)
| Chunk | sendfile GB/s | copy GB/s |
|---|---|---|
| 4 KiB | 3.13 | 1.02 |
| 64 KiB | 3.12 | 1.92 |
| 1 MiB | 2.90 | 2.09 |
sendfile does not take a userspace buffer size the way read does — throughput stayed ~3 GB/s across the sweep. The copy arm hated 4 KiB (syscall + object churn). Mid-to-large chunks recovered some, but never caught sendfile here.
When userspace copy still wins the design
- You must transform bytes (encrypt, filter, re-encode) before they hit the wire — sendfile cannot help mid-stream.
- Non-file sources (pipes you already hold as buffers, generators, app-built JSON).
- Portability / simplicity on platforms where sendfile semantics differ; measure before you rewrite.
- Tiny responses where setup dwarfs copy cost (not this lab’s 16–256 MiB regime).
mmap for in-process scans is a different question — see the mmap vs read lab. Keep-alive is about connection reuse, not the sendpath copy.
Related links:
Pitfalls we hit (or avoided)
- Calling loopback GB/s “disk throughput” — warm cache + TCP loopback; cold NVMe ranking may differ.
- Tuning sendfile “chunk size” — the interesting knob is on the copy arm.
- Forgetting the client must drain — undrained sockets back-pressure both arms equally; we measured full receives.
- Confusing with mmap page-touch — that lab scans into the process; this one ships bytes to a peer.
- Assuming zero-copy means zero CPU — the kernel still moves pages on loopback; we still paid less than Python copy.
Practical checklist
- Static / file-backed responses: prefer sendfile /
FileResponse/ nginxsendfile onand measure. - If you must
readinto userspace, avoid tiny chunks on bulk paths. - Report payload size + path (loopback vs NIC, warm vs cold) with every GB/s claim.
- Do not cite this as a WAN or cold-disk result.
- Pair with keep-alive / backlog labs when the bottleneck is connections, not copy.
Verdict
On warm localhost TCP, os.sendfile beat read+sendall ~1.4–1.9× (2.60 vs 1.83 GB/s at 16 MiB; 2.93 vs 1.68 at 64 MiB; 4.41 vs 2.34 at 256 MiB, 64 KiB chunks). The copy arm’s 4 KiB collapse to ~1.0 GB/s is the other headline — chunk size is a real footgun when you stay in userspace.
Evidence path on the lab box: lab-evidence/22-sendfile-vs-copy/results/. Affiliates: 0.
Lab evidence
What I found running this
Lab 30 Sep 2026 IST. Python 3.13.5. Localhost TCP serve: os.sendfile vs read(chunk)+sendall. 16 MiB/64 KiB p50: sendfile 2.599 GB/s vs copy 1.829 (~1.42x). 64 MiB: 2.930 vs 1.680 (~1.74x). 256 MiB: 4.411 vs 2.341 (~1.88x). Chunk sweep 64 MiB: copy 4 KiB 1.018 / 64 KiB 1.919 / 1 MiB 2.091; sendfile ~2.9–3.1 GB/s (chunk unused). Warm page cache + loopback. Affiliates: 0. Evidence: lab-evidence/22-sendfile-vs-copy/.
Related links
Plate 12
islice vs list Slice Windows: Localhost Lab
Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 07
heapq.merge vs sorted(chain): Localhost Lab
Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026
Plate 88
mmap Write vs pwrite Region: Localhost Lab
Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.
Observability & SRE · 1 Oct 2026