2026-07-26
Spring Boot container memory: size the whole JVM and diagnose exit 137
Build a whole-process JVM memory budget, then distinguish Java heap exhaustion from a cgroup OOM kill using JVM, Docker, and cgroup evidence.
An exit 137 after raising -Xmx is not a heap diagnosis. A container limit applies to memory accounted to its cgroup, while -Xmx caps one important part of the JVM. First classify the termination, then make a budget that leaves room for everything else.
I ran two small Java probes, not a Spring Boot application, in Temurin 25.0.3+9 with a 536,870,912-byte cgroup v2 hard limit, or 512 MiB. The exact image reference was:
eclipse-temurin@sha256:201fbb8886b2d273218aa3a192f0afbf7b5ff65ee8cc6ef47f5dce2171f013eaOne MiB is 1,048,576 bytes. The cases are useful contrast, not a benchmark or a safe heap setting for a service.
| Observation | Java heap exhaustion case | Executed container OOM result |
|---|---|---|
| Java error | OutOfMemoryError: Java heap space |
none captured |
| Process or container exit | 1 | 137 |
| Runtime OOM flag | false | true |
| Heap dump | 28,956,444 bytes | absent |
| Primary diagnostic owner | JVM artifacts | Docker runtime state; cgroup delta not captured |
This is the result of the pinned probe, not a universal lookup table. The second probe applied direct-buffer pressure. Docker reported OOMKilled=true, but the probe did not capture a post-failure memory.events oom_kill delta, so it does not meet the cgroup-delta classification later in this article. It does not show that direct buffers cause every cgroup OOM.
Start with the enforced cgroup limit
On cgroup v2, memory.max is the hard memory limit and memory.current reports current memory usage. memory.peak is a persistent high-water mark when available. memory.events exposes cumulative counters including oom and oom_kill; the latter counts processes killed by an OOM killer. The Linux cgroup v2 memory-controller documentation defines those files and counters.
The executed probe read 536870912 from its effective memory.max. Resolve the cgroup v2 mount from /proc/self/mountinfo and the process cgroup from /proc/self/cgroup, then inspect that directory and each visible ancestor through the mount. A local memory.max can be max while a parent, slice, pod, or runtime cgroup supplies the effective constraint. A cgroup namespace can hide host-side ancestors above its root. When that happens, inspect the host or runtime view and its configured limit instead of treating the namespace root as proof that no tighter parent exists.
Cgroup accounting is broader than Java objects. It can include anonymous memory, page cache, socket buffers, and kernel-accounted data. If another process shares the cgroup, it belongs in the same budget. Docker can impose memory constraints and documents that the kernel may kill processes when reclaim cannot satisfy the limit in its resource-constraints documentation.
cgroup_path=$(awk -F: '$1 == "0" && $2 == "" { print $3; exit }' /proc/self/cgroup)
[ -n "$cgroup_path" ] || { printf 'cgroup v2 path not found\n' >&2; exit 1; }
best_root=
best_mount=
best_dir=
while read -r root_field mount_field; do
cgroup_root=$(printf '%b' "$root_field")
cgroup_mount=$(printf '%b' "$mount_field")
[ "$cgroup_mount" = / ] || cgroup_mount=${cgroup_mount%/}
if [ "$cgroup_root" = / ]; then
relative_path=$cgroup_path
elif [ "$cgroup_path" = "$cgroup_root" ]; then
relative_path=/
elif [ "${cgroup_path#"$cgroup_root"/}" != "$cgroup_path" ]; then
relative_path=/${cgroup_path#"$cgroup_root"/}
else
continue
fi
candidate="${cgroup_mount%/}${relative_path}"
[ "$candidate" = / ] || candidate=${candidate%/}
[ -d "$candidate" ] || continue
if [ -z "$best_dir" ] || [ "${#cgroup_root}" -lt "${#best_root}" ]; then
best_root=$cgroup_root
best_mount=$cgroup_mount
best_dir=$candidate
fi
done < <(awk '$0 ~ / - cgroup2 / { print $4, $5 }' /proc/self/mountinfo)
[ -n "$best_dir" ] || { printf 'no cgroup2 mount contains the process cgroup\n' >&2; exit 1; }
cgroup_mount=$best_mount
cgroup_dir=$best_dir
while :; do
printf '%s\n' "$cgroup_dir"
for file in memory.max memory.current memory.peak memory.events memory.events.local; do
[ -r "$cgroup_dir/$file" ] && { printf '%s: ' "$file"; cat "$cgroup_dir/$file"; }
done
[ "$cgroup_dir" = "$cgroup_mount" ] && break
cgroup_dir=${cgroup_dir%/*}
[ -n "$cgroup_dir" ] || cgroup_dir=/
done
docker inspect --format 'Status={{.State.Status}} ExitCode={{.State.ExitCode}} OOMKilled={{.State.OOMKilled}} StartedAt={{.State.StartedAt}} FinishedAt={{.State.FinishedAt}} RestartCount={{.RestartCount}}' <container>Use these as operator templates, not executed probe commands. Capture memory.events and, when present, memory.events.local before and after a controlled reproduction, then calculate deltas for the same cgroup and time window. memory.events is hierarchical, so its delta can come from a descendant cgroup. memory.events.local reports events local to the cgroup being read. A local max, oom, and oom_kill sequence aligned with runtime and host evidence is stronger evidence, but neither counter alone identifies the killed process or the exact limit that caused the event. Runtime metadata is corroboration, not a substitute for cgroup evidence. For repeatable peak evidence, use a fresh cgroup or the kernel-documented reset mechanism supported by the deployed kernel; do not assume a write-based reset is safe or available.
Make a whole-process budget
Do not choose a heap percentage first. Start with the limit, reserve terms that are not heap, then test the candidate against the real image and workload.
cgroup hard limit
= heap committed, bounded by -Xmx
+ HotSpot native committed subtotal
+ thread stacks
+ direct and mapped buffers not already counted above
+ non-JVM native allocations
+ other cgroup-accounted memory
+ measured transient growth
+ safety margin-Xmx is a Java heap ceiling, not heap committed and not memory currently charged to the cgroup. Heap committed is the portion of that heap currently committed. Define the HotSpot native committed subtotal as the sum of NMT committed categories excluding Java Heap and Thread; it includes metaspace and class space, code cache, GC structures, compiler work, and internal allocations. Budget thread stacks separately from NMT's Thread committed value, observed thread count, and stack settings such as -Xss. Category subtotals are already contained in NMT's whole-report committed value and must not be added to that total. Buffer use can come from networking, compression, database drivers, file I/O, or application code; count only the portion not already represented by the chosen NMT categories. JNI, agents, TLS, image libraries, and other native code can allocate outside the part HotSpot observes.
| Budget term | How to measure | Peak or bound | Confidence and caveat |
|---|---|---|---|
| cgroup hard limit | memory.max or runtime configuration |
fixed | verify the effective cgroup |
| Java heap | heap committed from JVM metrics; -Xmx as ceiling |
measured and bounded | committed is not -Xmx |
| HotSpot native subtotal | NMT committed categories excluding Java Heap and Thread | measured | category subtotals are already contained in whole-report NMT committed |
| direct and mapped buffers | JVM buffer metrics plus workload evidence | measured | count only the portion not in the chosen NMT categories; library behavior can burst |
| thread stacks | NMT Thread, thread count, and stack settings such as -Xss |
estimated and measured | reservation differs from commitment |
| non-JVM native | cgroup or RSS gap and OS tools | measured or bounded | NMT does not fully cover it |
| other cgroup use | memory.stat, process inventory |
measured | includes non-process charges |
| transient and safety margin | repeated peak spread and policy | explicit | never an unnamed remainder |
RSS, NMT committed, heap committed, and memory.current are different views. RSS is resident memory for a process. NMT committed describes HotSpot-tracked committed memory. Heap committed is only the committed Java heap. memory.current is the cgroup's current accounted memory. Compare their direction and gaps, but do not add them as independent totals.
A practical sequence is to fix the container limit and swap policy, warm the real Spring Boot application on its production image, then take a post-warmup NMT baseline. Run a production-shaped load window and a failure-oriented stress window separately. Record heap after GC, NMT committed categories and diffs, thread and buffer behavior, process RSS where available, memory.current, memory.peak, and cgroup events. Repeat it. A budget is not credible if it relies on every term peaking at a different time, has no explicit margin, or passes once by chance.
No production workload, thread sweep, framework-specific direct-buffer workload, native-library workload, or GC comparison was run for this article. The probe cannot establish a safe -Xmx for a real service.
Read JVM ergonomics without calling them a budget
The pinned Temurin 25.0.3 probe reported UseContainerSupport=true, MaxRAMPercentage=25.000000, and an ergonomic MaxHeapSize=134217728, or 128 MiB. That is the result for this 512 MiB probe only. It is neither a Spring Boot result nor a production recommendation.
For Oracle JDK 25 on Linux, container support is enabled by default, and the JVM uses detected container memory and processors for selected ergonomic choices. The same documentation specifies the percentage-based heap options, including the default MaxRAMPercentage; inspect the actual immutable image because vendor and release boundaries matter. See the JDK 25 java command reference.
A small ergonomic heap can produce a heap OOM. An aggressively raised heap can crowd out native and cgroup headroom. Container awareness does not prove the whole process fits.
java -XX:+PrintFlagsFinal -version
java -XshowSettings:vm -versionThese commands are templates, not commands run for the probe. Use an explicit -Xmx when the measured nonheap terms and margin justify a fixed ceiling. Percentage-based sizing is another input to inspect, not a claim that one percentage works for all services.
Use NMT for the HotSpot part of the picture
Native Memory Tracking, or NMT, separates reserved address space from committed HotSpot memory. The JDK 25 diagnostic-tools guide describes its summary and detail modes, baselines and diffs, overhead, and the fact that it does not track allocations by non-JVM code.
The executed startup-scale snapshot reported 1,490 MB reserved and 36 MB committed. The 1,490 MB is virtual address-space reservation, not 1,490 MB resident memory inside a 512 MiB cgroup. The committed figure was the relevant NMT physical-commit scale at that instant. The report included Java Heap, Class, Thread, Code, Internal, Symbol, Shared class space, and Metaspace categories.
That snapshot was taken before an application workload. It is not a peak. Category sizes change with the collector, classpath, thread count, agents, native libraries, and workload. A post-warmup baseline followed by a load-window diff is more useful for growth than a startup number.
java -XX:NativeMemoryTracking=summary -jar <application>.jar
jcmd <pid> VM.native_memory summary
jcmd <pid> VM.native_memory baseline
jcmd <pid> VM.native_memory summary.diffThese are configuration and diagnostic templates, not executed probe commands. NMT has allocation-header and CPU overhead. It sees HotSpot internals, not every library or JNI allocation, so reconcile it with process and cgroup observations rather than treating it as a complete bill.
Add Spring Boot telemetry, with a clear boundary
This section is source-reviewed only against the current Spring Boot 4.1 reference, reviewed 2026-07-26. No Spring Boot application was executed for the probe. Spring Boot auto-configures JVM metrics under the jvm. prefix and system, process, and disk metrics under their corresponding names through Micrometer, as documented in the Spring Boot Actuator metrics reference.
Use Actuator for trends and correlation. It does not replace the runtime's termination state or cgroup events. A flat heap line does not exclude native growth.
First enumerate what the chosen application version actually permits and exposes. Endpoint access and transport exposure are separate controls: an endpoint is available only when access permits it and the relevant HTTP or JMX transport exposes it. In the reviewed Spring Boot 4.1 reference, only health is exposed by default over HTTP and JMX. See the Spring Boot Actuator endpoints documentation for access, exposure, and default-health behavior. The metrics endpoint lists available code-defined meter names at /actuator/metrics, and an individual meter can be inspected through /actuator/metrics/{requiredMetricName}. Those details are in the Actuator metrics endpoint documentation.
curl --fail --silent http://<application>/actuator/metrics
curl --fail --silent http://<application>/actuator/metrics/<requiredMetricName>These requests are unexecuted templates; apply the application's access and exposure policy before using them. Use the resulting names to track heap used, committed, and max by area or pool, buffer use when exposed, thread behavior, and process memory where the binder and platform expose it. Chart cgroup working use and OOM events outside the JVM with aligned labels and time windows. Restrict Actuator exposure. Diagnostic endpoints should not be public by accident.
Diagnose the termination before tuning
Preserve the previous exit state and timestamps before a restart, because current Docker State fields describe the latest attempt. Logs generally survive a restart, but remain subject to the log driver, rotation, container replacement or removal, and platform retention. Check for a Java OutOfMemoryError detail message, stack trace, GC failure pattern, fatal-error log, and any successfully written heap dump. Then check the runtime state and the cgroup counters across the same window.
Exit 137 is consistent with a process terminated by SIGKILL under the usual shell and container exit-code convention. It does not identify who sent that signal. Do not call it an OOM kill until runtime or cgroup evidence corroborates it.
| Evidence combination | Classification | Next action |
|---|---|---|
| heap-space OOME, heap dump, no runtime or local cgroup OOM kill evidence | Java heap exhaustion | analyze the dump and live-set behavior before changing heap |
exit 137, runtime OOM flag true, and a time-correlated local oom_kill increment |
cgroup OOM kill evidence, process and limit attribution unresolved | reduce or bound whole-process use, increase a justified limit, or both |
| exit 137 without OOM corroboration | SIGKILL, cause unresolved | investigate runtime stop, operator action, health timeout, and host events |
| no heap dump and no runtime or cgroup evidence | insufficient evidence | repair artifact collection and reproduce safely |
| JVM OOME and local or resolved hierarchical cgroup OOM evidence near the same time | mixed or cascading failure | use timestamps and sequence; do not force one label |
The heap probe matches the first row: it printed java.lang.OutOfMemoryError: Java heap space, exited 1, had oom_killed=false, and created a 28,956,444-byte dump. The direct-buffer case ended with exit 137, OOMKilled=true, and no dump, but no before-and-after memory.events capture exists to establish an oom_kill increment. It is therefore observed Docker runtime evidence, not an executed instance of the second row's recommended cgroup-delta classification. This contrast is diagnostic evidence from a small Java program, not a complete Spring Boot workload.
A missing dump proves little by itself. Heap-dump creation may be disabled, the destination may be unwritable, storage may be full, or the failure may be another exhaustion. When enabling dumps, use a restricted destination, limit access, set retention, and delete the file securely after analysis. Heap dumps can contain credentials, personal data, and request payloads. Oracle documents heap-dump behavior and its operational context in the JDK 25 memory troubleshooting guide.
java -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=<restricted-directory> -jar <application>.jar
docker inspect --format 'Status={{.State.Status}} ExitCode={{.State.ExitCode}} OOMKilled={{.State.OOMKilled}} StartedAt={{.State.StartedAt}} FinishedAt={{.State.FinishedAt}} RestartCount={{.RestartCount}}' <container>This heap-dump configuration and inspection command are unexecuted templates. A memory.events.local oom_kill delta greater than zero is local cgroup evidence that an OOM killer killed a process, while a memory.events delta can reflect a descendant cgroup. A local max, oom, and oom_kill sequence aligned with runtime and host evidence is stronger, but neither counter alone proves which process was killed or which exact limit caused the event. Correlate the cgroup path, container identity, and timestamps before assigning the failure to one process.
Turn the diagnosis into an acceptance gate
Before accepting a memory setting, record the immutable image digest, JDK version, effective JVM flags, container limit, and swap policy. Use the real application and a production-shaped load. Verify that the heap live set and GC behavior fit the explicit ceiling, NMT committed categories and diffs remain bounded, and buffers, threads, agents, and native libraries retain measured headroom.
Keep an explicit margin below cgroup current and peak usage. Repeated runs should show stable peaks, with no oom, oom_kill, heap OOM, failed heap dump, or restart-loop event. Preserve both JVM and cgroup artifacts in dashboards and incident collection.
Size against the cgroup as a whole, then diagnose from the layer that terminated the process.