ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 87

  1. Blog
  2. /Observability & SRE

sendfile vs Userspace Copy: Localhost TCP File Serve Lab

Hands-on sendfile vs read+sendall lab: 256 MiB p50 4.41 vs 2.34 GB/s (~1.9×); copy collapses at 4 KiB. Warm localhost TCP measured, no Docker.

Aditya Challa·30 September 2026·5 min read

Hands-on
On this page
  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — size sweep @ 64 KiB chunks (p50)
  5. Arm B — chunk sweep @ 64 MiB (p50)
  6. When userspace copy still wins the design
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict

Intro — what this post promises

Serving a file over a socket, folklore says sendfile skips a userspace copy that read + write/send pays for. The mmap lab already warned not to confuse page-touch scans with the kernel send path. This lab measures the socket path.

This is a hands-on lab with measured numbers:

  1. Localhost TCP serve of 16 / 64 / 256 MiB files via os.sendfile vs read(chunk)+sendall.
  2. Chunk-size sensitivity on the copy arm (4 KiB / 64 KiB / 1 MiB).
  3. Why sendfile’s “chunk” does not matter the same way.
  4. Warm-cache / loopback honesty — what these GB/s are not.

Related links:

  • mmap vs read (+ O_DIRECT) localhost lab
  • Unix Domain Socket vs TCP localhost lab
  • HTTP Keep-Alive vs Connection: close lab
  • Why your average latency graph is lying (p50 / p95 / p99)

Lab honesty (30 Sep 2026 IST): Shared Linux lab box (8 cores, kernel 6.12). Python 3.13.5. Server and client on 127.0.0.1 only. Payloads under /workspace (overlay), warm page cache. Protocol: 8-byte length header + body drain. No Docker. No public bind. No splice/io_uring. Affiliates: 0.

Verdict up front: at 256 MiB / 64 KiB, sendfile 4.41 GB/s beat copy 2.34 GB/s (~1.9×). The copy arm fell to ~1.0 GB/s at 4 KiB chunks; sendfile stayed near ~3 GB/s.


What we compared

ArmShape in this lab
sendfileos.sendfile(out_fd, in_fd, offset, count) until EOF
userspace copyfile.read(chunk) then socket.sendall(buf)

Related links:

  • man 2 sendfile

Lab topology

payload_16m.bin / payload_64m.bin / payload_256m.bin
Server threads on 127.0.0.1 · client drains full body
Arms: sendfile | copy (chunk 4 KiB / 64 KiB / 1 MiB)
Metric: wall GB/s (p50 over 7 rounds) for one full transfer

Arm A — size sweep @ 64 KiB chunks (p50)

Sizesendfile GB/scopy GB/sRatio
16 MiB2.601.83~1.42×
64 MiB2.931.68~1.74×
256 MiB4.412.34~1.88×

Larger transfers amortize setup and favor the kernel path. Even on loopback — where the NIC is fake and the page cache is hot — skipping the Python bytes round-trip still showed up.


Arm B — chunk sweep @ 64 MiB (p50)

Chunksendfile GB/scopy GB/s
4 KiB3.131.02
64 KiB3.121.92
1 MiB2.902.09

sendfile does not take a userspace buffer size the way read does — throughput stayed ~3 GB/s across the sweep. The copy arm hated 4 KiB (syscall + object churn). Mid-to-large chunks recovered some, but never caught sendfile here.


When userspace copy still wins the design

  • You must transform bytes (encrypt, filter, re-encode) before they hit the wire — sendfile cannot help mid-stream.
  • Non-file sources (pipes you already hold as buffers, generators, app-built JSON).
  • Portability / simplicity on platforms where sendfile semantics differ; measure before you rewrite.
  • Tiny responses where setup dwarfs copy cost (not this lab’s 16–256 MiB regime).

mmap for in-process scans is a different question — see the mmap vs read lab. Keep-alive is about connection reuse, not the sendpath copy.

Related links:

  • mmap vs read (+ O_DIRECT) localhost lab
  • HTTP Keep-Alive vs Connection: close lab

Pitfalls we hit (or avoided)

  1. Calling loopback GB/s “disk throughput” — warm cache + TCP loopback; cold NVMe ranking may differ.
  2. Tuning sendfile “chunk size” — the interesting knob is on the copy arm.
  3. Forgetting the client must drain — undrained sockets back-pressure both arms equally; we measured full receives.
  4. Confusing with mmap page-touch — that lab scans into the process; this one ships bytes to a peer.
  5. Assuming zero-copy means zero CPU — the kernel still moves pages on loopback; we still paid less than Python copy.

Practical checklist

  • Static / file-backed responses: prefer sendfile / FileResponse / nginx sendfile on and measure.
  • If you must read into userspace, avoid tiny chunks on bulk paths.
  • Report payload size + path (loopback vs NIC, warm vs cold) with every GB/s claim.
  • Do not cite this as a WAN or cold-disk result.
  • Pair with keep-alive / backlog labs when the bottleneck is connections, not copy.

Verdict

On warm localhost TCP, os.sendfile beat read+sendall ~1.4–1.9× (2.60 vs 1.83 GB/s at 16 MiB; 2.93 vs 1.68 at 64 MiB; 4.41 vs 2.34 at 256 MiB, 64 KiB chunks). The copy arm’s 4 KiB collapse to ~1.0 GB/s is the other headline — chunk size is a real footgun when you stay in userspace.

Evidence path on the lab box: lab-evidence/22-sendfile-vs-copy/results/. Affiliates: 0.

sendfile vs writeos.sendfilezero copyuserspace copylinux performancelocalhost labsretcp file serve

Lab evidence

What I found running this

Lab 30 Sep 2026 IST. Python 3.13.5. Localhost TCP serve: os.sendfile vs read(chunk)+sendall. 16 MiB/64 KiB p50: sendfile 2.599 GB/s vs copy 1.829 (~1.42x). 64 MiB: 2.930 vs 1.680 (~1.74x). 256 MiB: 4.411 vs 2.341 (~1.88x). Chunk sweep 64 MiB: copy 4 KiB 1.018 / 64 KiB 1.919 / 1 MiB 2.091; sendfile ~2.9–3.1 GB/s (chunk unused). Warm page cache + loopback. Affiliates: 0. Evidence: lab-evidence/22-sendfile-vs-copy/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Plate 12

    islice vs list Slice Windows: Localhost Lab

    Hands-on itertools.islice vs list slice window lab: real ops/s taking ranges from sequences, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 07

    heapq.merge vs sorted(chain): Localhost Lab

    Hands-on heapq.merge vs sorted(chain) multi-way merge: real records/s on pre-sorted lists, measured on Linux localhost today in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

  • Plate 88

    mmap Write vs pwrite Region: Localhost Lab

    Hands-on mmap MAP_SHARED write+msync vs pwrite region update: real MB/s with durability labels, measured on Linux localhost in this hands-on lab for SREs.

    Observability & SRE · 1 Oct 2026

On this page

  1. Intro — what this post promises
  2. What we compared
  3. Lab topology
  4. Arm A — size sweep @ 64 KiB chunks (p50)
  5. Arm B — chunk sweep @ 64 MiB (p50)
  6. When userspace copy still wins the design
  7. Pitfalls we hit (or avoided)
  8. Practical checklist
  9. Verdict
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove