Memory Pressure
Linux kernel memory pressure mitigation involves a combination of proactive memory reclaim strategies, OOM (Out-Of-Memory) killer behavior, and tunable parameters to ensure system stability under heavy workloads. This section explores how the kernel manages memory scarcity and how administrators can optimize these mechanisms.
Kernel Memory Reclaim Strategies¶
The Linux kernel employs a multi-layered approach to reclaim memory when system-wide memory pressure is detected. Key components include:
1. Page Reclaim and kswapd¶
The kernel's page reclaim process is managed by the kswapd daemon, which runs in the background to free memory by:
- Swapping inactive pages to disk.
- Reclaiming memory from caches and unused buffers.
- Forcing applications to release memory via direct reclaim during allocation failures.
The vm.swappiness parameter controls the aggressiveness of swapping:
# Default value (60) balances between reclaiming and swapping
sysctl vm.swappiness=10 # Reduce swapping for server workloads
2. Dirty Page Management¶
The kernel manages dirty (modified but not yet written to disk) pages via:
- vm.dirty_ratio: Maximum percentage of system memory that can be dirty before writeback is forced (default: 20%).
- vm.dirty_background_ratio: Threshold for background writeback (default: 10%).
# Adjust for high-write workloads (e.g., databases)
sysctl vm.dirty_ratio=30
sysctl vm.dirty_background_ratio=15
3. NUMA Awareness¶
On multi-socket systems, the kernel uses NUMA (Non-Uniform Memory Access) to allocate memory closer to CPU cores, reducing latency. Tools like numactl can enforce memory policies for specific processes.
OOM Killer Behavior¶
When memory pressure is critical and reclaim fails, the OOM killer terminates processes to free memory. Its behavior is governed by:
1. OOM Score Calculation¶
The OOM killer assigns a score to each process based on:
- Memory usage (higher usage = higher score).
- Process priority (nice value; lower = higher score).
- oom_score_adj (manual tuning; range: -1000 to +1000).
Example: Prevent a critical service from being killed:
2. OOM Killer Logs¶
Use /var/log/kern.log or dmesg to identify victimized processes:
3. Avoiding OOM Kills¶
- Limit memory allocation for non-critical services.
- Use cgroups to isolate workloads and enforce memory limits.
- Monitor with tools like
smemorslabtop.
Tuning for Large-Scale Workloads¶
For systems handling massive workloads (e.g., databases, HPC clusters), consider these optimizations:
1. Swap Configuration¶
- Use swap files or SSD-based swap for faster I/O.
- Avoid overcommitting memory; adjust
vm.overcommit_memory:
2. Slab and Page Cache Tuning¶
- Optimize slab memory (kernel object caches) with
sysctlparameters likevm.min_free_kbytes. - Use
sysctl vm.vfs_cache_pressureto control page cache eviction (default: 100; higher values increase eviction).
3. Kernel Parameters for Stability¶
# Example tuning for a server workload
sysctl vm.swappiness=10
sysctl vm.min_free_kbytes=32768
sysctl vm.vfs_cache_pressure=200
Key takeaways¶
- Memory reclaim is prioritized over swapping by default; adjust
vm.swappinessfor workload-specific needs. - OOM killer is a last-resort mechanism; use
oom_score_adjto protect critical processes. - NUMA awareness and dirty page tuning optimize performance for large-scale systems.
- Proactive monitoring and cgroup isolation prevent OOM events in production environments.