Skip to content

Memory Pressure

Linux kernel memory pressure mitigation involves a combination of proactive memory reclaim strategies, OOM (Out-Of-Memory) killer behavior, and tunable parameters to ensure system stability under heavy workloads. This section explores how the kernel manages memory scarcity and how administrators can optimize these mechanisms.


Kernel Memory Reclaim Strategies

The Linux kernel employs a multi-layered approach to reclaim memory when system-wide memory pressure is detected. Key components include:

1. Page Reclaim and kswapd

The kernel's page reclaim process is managed by the kswapd daemon, which runs in the background to free memory by: - Swapping inactive pages to disk. - Reclaiming memory from caches and unused buffers. - Forcing applications to release memory via direct reclaim during allocation failures.

The vm.swappiness parameter controls the aggressiveness of swapping:

# Default value (60) balances between reclaiming and swapping
sysctl vm.swappiness=10  # Reduce swapping for server workloads

2. Dirty Page Management

The kernel manages dirty (modified but not yet written to disk) pages via: - vm.dirty_ratio: Maximum percentage of system memory that can be dirty before writeback is forced (default: 20%). - vm.dirty_background_ratio: Threshold for background writeback (default: 10%).

# Adjust for high-write workloads (e.g., databases)
sysctl vm.dirty_ratio=30
sysctl vm.dirty_background_ratio=15

3. NUMA Awareness

On multi-socket systems, the kernel uses NUMA (Non-Uniform Memory Access) to allocate memory closer to CPU cores, reducing latency. Tools like numactl can enforce memory policies for specific processes.


OOM Killer Behavior

When memory pressure is critical and reclaim fails, the OOM killer terminates processes to free memory. Its behavior is governed by:

1. OOM Score Calculation

The OOM killer assigns a score to each process based on: - Memory usage (higher usage = higher score). - Process priority (nice value; lower = higher score). - oom_score_adj (manual tuning; range: -1000 to +1000).

Example: Prevent a critical service from being killed:

# Set oom_score_adj to -500 (lower than default 0)
echo -500 | sudo tee /proc/sys/vm/oom_score_adj

2. OOM Killer Logs

Use /var/log/kern.log or dmesg to identify victimized processes:

dmesg | grep -i 'oom'

3. Avoiding OOM Kills

  • Limit memory allocation for non-critical services.
  • Use cgroups to isolate workloads and enforce memory limits.
  • Monitor with tools like smem or slabtop.

Tuning for Large-Scale Workloads

For systems handling massive workloads (e.g., databases, HPC clusters), consider these optimizations:

1. Swap Configuration

  • Use swap files or SSD-based swap for faster I/O.
  • Avoid overcommitting memory; adjust vm.overcommit_memory:
    sysctl vm.overcommit_memory=2  # Disable overcommit for predictable workloads
    

2. Slab and Page Cache Tuning

  • Optimize slab memory (kernel object caches) with sysctl parameters like vm.min_free_kbytes.
  • Use sysctl vm.vfs_cache_pressure to control page cache eviction (default: 100; higher values increase eviction).

3. Kernel Parameters for Stability

# Example tuning for a server workload
sysctl vm.swappiness=10
sysctl vm.min_free_kbytes=32768
sysctl vm.vfs_cache_pressure=200

Key takeaways

  • Memory reclaim is prioritized over swapping by default; adjust vm.swappiness for workload-specific needs.
  • OOM killer is a last-resort mechanism; use oom_score_adj to protect critical processes.
  • NUMA awareness and dirty page tuning optimize performance for large-scale systems.
  • Proactive monitoring and cgroup isolation prevent OOM events in production environments.