Identify memory leaks, fragmentation, and OOM errors in ACK clusters with visual diagnostics.
Memory diagnostics covers three areas: memory overview, memory analysis, and OOM analysis. Inspect memory at both the node and pod level.
Diagnostic items shown reflect your actual cluster configuration.
When you run diagnostics, ACK collects data from each node, including system version, load status, Docker and kubelet status, and key error messages in system logs. ACK does not collect business data or sensitive information.
Diagnostic workflow
Use the three diagnostic areas in sequence to narrow down a memory issue:
-
Memory overview — Check for memory risks: leaked memory, fragmentation, unreleased Memcg entries, and THP waste. Use charts to confirm whether abnormal usage is in kernel or application memory.
-
Memory analysis — Drill down to process- and pod-level memory usage to identify which process or container consumes excessive anonymous memory, page cache, or shared memory.
-
OOM analysis — Review OOM event counts and types to determine whether OOM errors occur at the node (Host) or container (cgroup) level, and which containers hit their memory limits.
Memory overview
Identifies memory risks across the following diagnostic items.
| Diagnostic item | Description |
|---|---|
| Leaked Memory | Checks for kernel memory leaks in the Slab, Vmalloc, and buddy system (allocpage). |
| Memory Usage | Displays system memory utilization. |
| Memcg | Checks whether unreleased memory cgroups (Memcg) degrade system performance or cause statistical errors. |
| Memory Fragmentation | Checks for memory fragmentation that degrades system performance. |
| THPZeroPage | Evaluates the ratio of Transparent Huge Page (THP) waste. |
System memory usage is charted in three categories:
-
Kernel memory (kernel): total memory used by the OS kernel.
-
Application memory (app): total memory used by user-mode programs.
-
Free memory (free): total free system memory.
Key concepts
The following terms are used throughout memory diagnostics.
| Term | Description |
|---|---|
| Memory leak | Occurs when dynamically allocated memory is never released, causing memory usage to grow continuously. Unresolved leaks degrade performance and can cause crashes. |
| Memory utilization | Memory utilization = (Total memory - Free memory) x 100 / Total memory. Page cache counts as free memory and does not affect utilization — the kernel can reclaim it at any time. |
| Unreleased Memcg | A memory cgroup not released due to a system exception. Can degrade system performance. |
| Memory fragmentation | Over time, free contiguous memory blocks become too small for large allocation requests. Delays allocation and causes application jitter. |
| Ratio of THP waste | Ratio of THP waste = Number of zero THPs x 100% / Total number of THPs. See THP details below. |
| Buddy system | Linux kernel algorithm for managing memory pages. Divides pages into 11 groups with power-of-two block sizes: 4 KB, 8 KB, 16 KB, 32 KB ... 4 MB. Most pages are 4 KB. |
| Slab | Allocates small memory pieces on top of the buddy system. |
| Vmalloc | Allocates memory with nonlinear mapping on top of the buddy system. |
| Page cache (filecache) | Linux caches file content in memory for faster subsequent access. |
| Anonymous memory | Memory dynamically allocated to a process's heap and stack through new, malloc, or mmap. Not backed by a file system. |
| Shared memory | A memory block shared by two or more processes for inter-process communication. |
| tmpfs | Linux temporary file system backed by memory. All reads and writes are cached in memory. |
| hugetlb | Memory consumed by huge pages in a file system. |
THP details
Transparent Huge Pages (THP) are huge pages sized 2 MiB or 1 GiB in the kernel. Each subpage is 4 KiB, so one 2-MiB THP equals 512 subpages.
With THP enabled, the kernel dynamically allocates THPs to reduce Translation Lookaside Buffer (TLB) misses and improve performance. However, THP can cause memory bloat and overcommitment: when an application requests only 8 KiB (2 subpages), the kernel allocates a full 2-MiB THP — leaving 510 zero subpages that waste resident set size (RSS) and can trigger OOM errors.
Kernel memory metrics
In most cases, memory leaks are indicated by abnormal usage in Sunreclaim or the buddy system. Monitor these metrics closely.
| Metric | Description |
|---|---|
| SReclaimable | Memory that the Slab can reclaim. |
| Sunreclaim | Memory that the Slab cannot reclaim. Abnormal growth strongly indicates a kernel memory leak. |
| PageTables | Memory occupied by kernel page tables. |
| Vmalloc | Memory allocated by Vmalloc. |
| KernelStack | Total kernel stack memory used by processes. |
| AllocPages | Memory allocated from the buddy system by functions such as alloc_pages. Not retrievable through any node file — excessive use creates a memory black hole. |
Application memory metrics
When analyzing user-mode memory usage, focus on anonymous memory, shared memory, and page cache.
| Metric | Description |
|---|---|
| filecache | Page cache that can be reclaimed by running drop caches. |
| anon | Anonymous memory used by a program's heap and stack. High usage suggests a process memory leak or enabled THP. |
| mlock | Memory locked by the system. |
| huge | Memory used by huge pages. |
| buffer | Memory used by block device and file system metadata. |
| shmem | Shared memory (tmpfs). Leaks occur if a tmpfs file is not deleted after the process exits, or if a file is deleted while still open. |
Memory analysis
Provides two views: process memory and pod memory.
Process memory
Lists processes sorted by memory usage and breaks down anonymous memory, page cache, and shared memory.
Pod memory
Shows which files occupy page cache and shared memory in each container and pod, with active and inactive cache ratios.
| Diagnostic item | Description |
|---|---|
| Pod | The name of the pod. |
| Container | The name of the container. |
| File | The full path of the file, including the file name. |
| Cache | The page cache (filecache) occupied by the file. |
| Container Cache | The container-level cache occupied by the file. Multiple processes in the same container may reference the same file. |
| Active Cache | Page cache currently in use. |
| Inactive Cache | Page cache not in use, eligible for reclaim. |
OOM analysis
Diagnoses out-of-memory errors across the following items.
| Diagnostic item | Description |
|---|---|
| OS OOM Count | Total OOM errors from host startup to diagnosis. |
| Available Memory | Current free system memory. |
| Low Watermark | The low memory threshold. When available memory drops below this value, the kernel triggers asynchronous memory reclaim. |
| Container | Pod name, container ID, or cgroup name. |
| limit | The memory limit configured for the container. |
| usage | Current memory used by the container. |
| OOM Count | Total OOM errors in the container. |
| OOM Type | The type of OOM error: Host or cgroup. |