Plate 63
JVM Heap vs RSS vs OOM on 512 MB: What `-Xmx` Does Not Tell You
Aditya Challa11 min read
On this page
- Intro — what this post promises
- Three numbers people conflate
- What `-Xmx` deliberately ignores
- Metaspace (class metadata)
- Thread stacks
- Direct byte buffers and other off-heap Java APIs
- Code cache, GC data structures, compiler arenas
- Native libraries and allocator behavior
- Lab method — measure the gap on 512 MB
- Suggested lab topology
- Measured budget for a 512 MiB ceiling (this lab)
- Phase table — heap ≠ RSS with real numbers
- Flags that help — and flags that merely rearrange deck chairs
- Reading failure modes correctly
- Advantages / disadvantages of common 512 MB strategies
- Practical checklist (copy into your runbook)
- FAQ
- External citations
- Related ShopperCove posts
- CTAs
Intro — what this post promises
Setting -Xmx256m on a 512 MB machine feels like leaving half the box free. Then dmesg shows Killed process ... java, or the container exits 137, and heap telemetry never hit the max.
The missing piece: the heap is only one arena inside a larger native process.
This post is a measurement lab for that gap. It deliberately does not re-litigate Spring Boot vs Quarkus. For framework footprint context, start here and come back:
Here you will get:
- Vocabulary: heap vs committed vs RSS vs cgroup usage.
- What usually fills the non-heap gap (metaspace, threads, direct buffers, GC/code cache, allocator waste).
- How to use Native Memory Tracking (NMT) without confusing it for RSS.
- A 512 MB-oriented checklist and a measured budget table from a ShopperCove lab run.
Lab honesty: Tables below are from a ShopperCove lab HTTP service (tiny JDK HttpServer — not a production SaaS) under cgroup v2 512 MiB on this Linux box, 29 Sep 2026 (IST): OpenJDK 21.0.12.1, flags -Xmx256m -Xms128m -XX:NativeMemoryTracking=summary -XX:MaxMetaspaceSize=96m -XX:MaxDirectMemorySize=128m -Xss512k -XX:+UseG1GC, load with hey. This was not a dedicated 512 MB VPS — same failure mode, different neighbors. Optional follow-ups: Grafana panels, AlwaysPreTouch idle compare, Serial vs G1 on your app.
Same small-box memory theme on Go (different runtime): Go GC + swap on a tiny VPS (CMS auto-slug until the slug editor ships).
Three numbers people conflate
| Term | Who reports it | What it means |
|---|---|---|
Heap used / -Xmx | JVM | Java object heap ceiling / occupancy |
| NMT “committed” | HotSpot NMT | Memory the JVM thinks it committed via tracked paths |
| RSS / cgroup memory.current | Linux / container runtime | Pages resident (or charged) for the process |
On a tiny VPS, the OOM killer and cgroup OOM care about RSS / cgroup, not your faith in -Xmx.
Container note: modern OpenJDK builds enable container awareness by default (UseContainerSupport), so ergonomics can see cgroup limits — but that still does not make -Xmx equal RSS. See Docker Hub OpenJDK docs on container CPU/RAM detection: https://hub.docker.com/_/openjdk
In this lab, PrintFlagsFinal confirmed MaxHeapSize=256 MiB from -Xmx256m and UseContainerSupport=true.
What -Xmx deliberately ignores
Metaspace (class metadata)
Class metadata lives off-heap (Metaspace since Java 8). Lots of classes (frameworks, generated proxies, Groovy/ByteBuddy) grow metaspace. Caps: -XX:MaxMetaspaceSize=… — set only with measurement; too-low causes OutOfMemoryError: Metaspace while RSS still has room.
Our lab app is tiny (~3 MiB metadata used). A Boot/Quarkus service will show a much larger Class/Metaspace line — measure yours.
Thread stacks
Each platform thread reserves stack space (-Xss; we used 512 KiB). Hundreds of threads → tens of MiB before heap. In this lab, adding 100 parked platform threads moved NMT Thread committed from 2.9 → 13.4 MiB and RSS from 324 → 334 MiB.
Virtual threads (Project Loom) change the trade-offs for application concurrency, but the JVM and carrier threads still have native costs — measure your actual thread model.
Direct byte buffers and other off-heap Java APIs
Netty, NIO, many DB drivers, and compressed caches use direct memory. Caps involve -XX:MaxDirectMemorySize (we set 128 MiB). In this lab, allocating 64 MiB of touched direct buffers raised RSS from 260 → 324 MiB while heap used stayed ~189 MiB. NMT put that under the Other category (~64 MiB committed).
Code cache, GC data structures, compiler arenas
JIT code and GC structures scale with workload and heap. At near-ceiling we saw GC ~42 MiB and Code ~8 MiB committed in NMT.
Native libraries and allocator behavior
NMT does not track all third-party native code or all JDK native allocations outside HotSpot’s accounting. Official Oracle docs: Native Memory Tracking — NMT tracks HotSpot VM usage; it is not a full process RSS auditor.
Additionally, malloc free lists can keep RSS high after Java thinks memory was released. Tools like pmap -X help from the OS side.
Rough mental model (not an equation to tattoo on a runbook):
Lab method — measure the gap on 512 MB
Suggested lab topology
| Piece | Suggestion |
|---|---|
| Host | 512 MB VPS or cgroup / docker run --memory=512m |
| JDK | Record exact build (java -version) |
| App | Same class of service as production — do not turn this into a framework bake-off |
| Load | Steady + spike (define QPS) |
| Flags (start) | -XX:NativeMemoryTracking=summary plus your heap flags |
Enable NMT at startup (cannot fully retrofit without restart):
Query:
OS side:
Measured budget for a 512 MiB ceiling (this lab)
ShopperCove lab API, OpenJDK 21.0.12.1, -Xmx256m, cgroup 512 MiB, 29 Sep 2026 IST.
| Bucket | MiB | Notes |
|---|---|---|
| Target cgroup / VPS | 512 | Hard stop |
-Xmx heap ceiling | 256 | Command-line |
| Heap used (near-ceiling) | ~247 | Retained byte[] live set |
| Direct buffers | ~96 | Touched ByteBuffer.allocateDirect |
| Thread stacks (NMT) | ~13 | 138 threads, -Xss512k |
| GC structures (NMT) | ~42 | G1 |
| Code cache (NMT) | ~8 | |
| Class / metaspace | small (~3 used) | Tiny app — frameworks drift up |
| VmRSS near-ceiling | ~444 | ≈ 1.7× -Xmx |
| Headroom to 512 | ~68 | Thin under churn (oom-probe RSS 457) |
Phase table — heap ≠ RSS with real numbers
| Phase | Heap used | Direct | Threads | NMT committed | VmRSS |
|---|---|---|---|---|---|
| Idle | 5 | 0 | 25 | 200 | 52 |
Light /json load | 21 | 0 | 38 | 211 | 91 |
| Retain ~180 MiB heap | 187 | 0 | 38 | 287 | 260 |
| +64 MiB direct | 189 | 64 | 38 | 351 | 324 |
| +100 threads | 191 | 64 | 138 | 359 | 334 |
Peak load (heap at -Xmx) | 201 | 64 | 138 | 420 | 412 |
| Near ceiling | 247 | 96 | 138 | 436 | 444 |
| OOM probe | 252 | 96 | 138 | 445 | 457 |
Idle teaching point: NMT committed (200 MiB) ≫ RSS (52 MiB) because -Xms128m commits heap pages that are not yet resident (no AlwaysPreTouch). Near-ceiling: NMT (436) ≈ RSS (444) once pages are touched.
Load notes: light /json p99 48 ms; under heap+direct pressure, alloc-heavy /churn slowed to p95 ~420 ms (oom-probe) and slowest 744 ms when the cgroup was squeezed to 384 MiB. Prefer p99/p95, not averages — see why average latency graphs lie (CMS auto-slug).
Flags that help — and flags that merely rearrange deck chairs
| Knob | Role | Caution |
|---|---|---|
-Xmx / -Xms | Heap ceiling / initial | Not RSS |
-XX:MaxRAMPercentage | Heap as % of detected RAM | Needs correct container detection |
-XX:MaxMetaspaceSize | Cap class metadata | Too low → Metaspace OOM |
-XX:MaxDirectMemorySize | Cap direct buffers | App may break if undersized |
-Xss | Stack size per thread | Lower carefully; stack overflow risk |
| GC choice (Serial, G1, …) | Pause vs footprint trade | On 512 MB, fewer GC worker threads often safer — measure |
-XX:+AlwaysPreTouch | Touch heap pages at start | Makes RSS jump early; sometimes desired for predictability |
Do not cargo-cult a giant flag list from a 64 GB server onto a 512 MB VPS.
Reading failure modes correctly
| Failure | Typical signal | First checks |
|---|---|---|
| Java heap OOM | OutOfMemoryError: Java heap space | Heap dump / alloc sites; is -Xmx too low for live set? |
| Metaspace OOM | OutOfMemoryError: Metaspace | Classloader leaks; MaxMetaspaceSize |
| Direct buffer OOM | OutOfMemoryError: Direct buffer memory | Netty pools; MaxDirectMemorySize |
| Native/cgroup OOM | exit 137, Killed process, cgroup oom_kill | RSS vs NMT gap; thread count; native libs |
| Thrash (if swap on) | High si/so, awful p99, process alive | Same lesson as the Go swap post — prefer fail-fast sizing |
What this lab hit: with RSS still under the 512 MiB cgroup, further allocs threw Java heap space and Direct buffer memory OOMs (limits 256 / 128 MiB). Cgroup oom_kill stayed 0. Squeezing memory.max to 384 MiB raised the cgroup max event counter without an immediate kill — pressure, not a clean OOM demo.
Host graphs context: How to read server monitoring graphs.
Advantages / disadvantages of common 512 MB strategies
| Strategy | Advantages | Disadvantages |
|---|---|---|
Tiny -Xmx, ignore RSS | Feels safe | Still OOM from non-heap; false confidence |
MaxRAMPercentage only | Adapts to cgroup size | Still leaves non-heap unbudgeted |
| NMT + RSS dashboards | Debuggable | NMT ≠ RSS; detail mode has overhead |
| Raise VPS RAM | Buys headroom | Cost; can hide leaks |
| Switch framework | May cut baseline RSS | Out of scope here — see existing compare post |
| Disable swap | Clear OOM signal | Requires correct budget |
No testimonials. No “cut memory 60% with one flag” claims without a lab.
Practical checklist (copy into your runbook)
- Record JDK vendor/version and container cgroup limits.
- Confirm
UseContainerSupport/ detected MaxHeap with-XX:+PrintFlagsFinal(filter MaxHeapSize). - Start with NMT=summary; take idle baseline (expect committed ≫ RSS without PreTouch).
- Apply realistic load; NMT summary + VmRSS + cgroup current.
- Attribute top NMT categories; check thread count.
- Inspect direct buffer usage (BufferPool MXBean / app metrics).
- Only then move
-Xmx/ metaspace / direct / stack knobs. - Alert on cgroup OOM events and RSS %, not heap % alone.
- If swap enabled, alert on swap-in (latency killer).
FAQ
Q1. If NMT committed is lower than RSS, is NMT wrong?
Not necessarily. NMT does not track everything; RSS includes untracked native memory and allocator retention. Use both.
Q2. If NMT committed is higher than RSS, is that possible?
Yes — committed virtual memory may not all be resident yet. Our idle phase showed ~200 MiB NMT committed vs ~52 MiB RSS. PreTouch and touch patterns change this.
Q3. Is -Xmx256m correct for 512 MB?
It is a common starting point, not a law. In this lab, RSS reached ~444 MiB with that heap ceiling once direct buffers and threads joined. Measure RSS under load. Some apps need smaller heap and fewer threads; some need more RAM.
Q4. Does G1 vs Serial matter on 512 MB?
GC choice affects native overhead and pause behavior. Compare on your app; do not trust internet defaults blindly. This lab used G1 with 2 parallel / 1 concurrent GC threads.
Q5. Will switching to Quarkus fix RSS automatically?
It might reduce baseline — that is what the existing compare post explores. This article’s job is teaching measurement so any stack stays honest.
Q6. Should I enable NMT detail in production?
Summary is usually enough; detail has higher overhead per Oracle’s NMT notes. Prefer staging for detail hunts.
External citations
- https://docs.oracle.com/en/java/javase/25/vm/native-memory-tracking.html
- https://docs.oracle.com/en/java/javase/21/docs/specs/man/java.html
- https://hub.docker.com/_/openjdk
- https://man7.org/linux/man-pages/man5/proc_pid_status.5.html
Related ShopperCove posts
- Spring Boot vs Quarkus on a 512 MB VPS
- Go GC + swap on a tiny VPS
- Why average latency graphs lie (percentiles)
- How to read server monitoring graphs
- About
CTAs
- Subscribe via RSS: https://www.shoppercove.com/feed.xml
- About the author / method: https://www.shoppercove.com/about
Lab evidence
What I found running this
OpenJDK 21.0.12.1 cgroup 512MiB lab 29 Sep 2026 IST. -Xmx256m. Near-ceiling: heap used 247MiB, direct 96, threads 138, NMT committed 436, VmRSS 444 (RSS≈1.7×Xmx). Idle RSS 52 vs NMT committed 200. Observed Java heap OOM + Direct buffer OOM; cgroup oom_kill=0.
Related links
Plate 27
Go GC + Swap on a Tiny VPS: Stop Thrashing Before the OOM Killer
29 Sept 2026
Spring Boot vs Quarkus on a 512 MB VPS: what the numbers actually showed
A side-by-side JVM-mode comparison on a memory-starved VPS, run through StatLite's new Quarkus support. The results are messier than a simple 'X uses less RAM' headline.
9 Sept 2026
Plate 41
React 19.3 View Transitions & Fragment Refs: Frontend Guide (2026)
2 Oct 2026