Skip to content

Network Resilience

High availability clusters rely on robust network resilience to ensure continuous operation during failures. Network resilience strategies must address both physical and logical redundancy, ensuring that nodes and virtual IP addresses (VIPs) remain accessible even when individual network components fail. This section outlines best practices for designing resilient network architectures in Pacemaker-based clusters.


Multi-Homed Node Configuration

Multi-homed nodes (nodes with multiple network interfaces) provide redundancy by allowing traffic to route through alternative paths. This is critical for maintaining connectivity during partial network outages.

Example: Configuring Multiple Network Interfaces

# Assign multiple IPs to a single interface (e.g., eth1)
sudo ip addr add 192.168.1.10/24 dev eth1
sudo ip addr add 192.168.2.10/24 dev eth1

# Verify configuration
ip addr show eth1

Bonding/Teaming for Redundancy

Bonding (Linux) or teaming (Windows) combines multiple NICs into a single logical interface, providing failover and load balancing:

# Example bonding configuration (eth0 and eth1)
sudo nmcli connection add type bond con-name bond0 ifname bond0 mode active-backup
sudo nmcli connection add type bond-slave con-name bond0-slave0 ifname eth0 master bond0
sudo nmcli connection add type bond-slave con-name bond0-slave1 ifname eth1 master bond0


Redundant VIP Configuration

Virtual IP addresses (VIPs) must be configured redundantly across cluster nodes to avoid single points of failure. Use multiple VIPs or assign VIPs to different subnets for added resilience.

Example: Assigning VIPs to Multiple Nodes

# Assign VIP to node1
sudo ip addr add 192.168.1.20/24 dev eth1

# Assign VIP to node2
sudo ip addr add 192.168.1.20/24 dev eth1

Pacemaker Resource Configuration

Ensure VIPs are managed by Pacemaker to handle failover automatically:

<resources>
  <ip addr="192.168.1.20" name="vip-1"/>
  <ip addr="192.168.2.20" name="vip-2"/>
</resources>

Cross-Subnet and ISP Diversity

Deploy VIPs across different subnets or ISPs to mitigate risks of localized outages. For example, use a VIP in a public subnet for external access and a private VIP for internal services.


Additional Best Practices

  1. Regular Failover Testing: Simulate network failures using tools like ipfail or keepalived to validate failover behavior.
  2. Network Monitoring: Integrate tools like nagios or Prometheus to monitor interface status, latency, and packet loss.
  3. Redundant Routing Protocols: Use protocols like VRRP (Virtual Router Redundancy Protocol) to manage default gateways across nodes.
  4. Firewall Rules: Ensure firewalls allow traffic between cluster nodes and VIPs, with rules prioritizing failover scenarios.

Key takeaways

  • Multi-homed nodes provide redundancy by using multiple NICs or bonding/teaming.
  • Redundant VIPs across nodes and subnets ensure continuous service availability.
  • Regular testing and monitoring validate network resilience in real-world scenarios.
  • Cross-subnet and ISP diversity reduces risks of localized network failures.
  • Automated failover via Pacemaker ensures VIPs transition seamlessly during outages.