Alert Suppression
Alert Suppression and Deduplication¶
In production environments, alert floods and redundant notifications can overwhelm teams and obscure critical issues. Alert suppression and deduplication strategies in Alertmanager help mitigate this by filtering out unnecessary alerts, ensuring only actionable notifications reach your team. These techniques are essential for maintaining observability hygiene and reducing alert fatigue.
Alert Suppression¶
Alert suppression prevents specific alerts from being sent based on the state of other alerts. This is particularly useful for blocking alerts that are redundant or dependent on other conditions.
1. Prometheus Rule Suppression¶
Use the suppress clause in Prometheus rules to block alerts when another alert is active. For example:
- alert: HighCPUUsage
expr: avg by (instance) (node_cpu_seconds_total{mode="idle"} < 0.1)
for: 5m
labels:
severity: warning
annotations:
summary: "High CPU usage on {{ $labels.instance }}"
description: "CPU usage is above 99% on {{ $labels.instance }}."
- alert: HighMemoryUsage
expr: avg by (instance) (node_memory_usage_bytes{job="node"} > 0.9 * node_memory_capacity_bytes{job="node"})
for: 5m
labels:
severity: critical
annotations:
summary: "High memory usage on {{ $labels.instance }}"
description: "Memory usage is above 90% on {{ $labels.instance }}."
suppress:
- { alertname: HighCPUUsage, equals: { severity: "warning" } }
HighMemoryUsage alerts when HighCPUUsage is active, assuming CPU and memory are correlated.
2. Alertmanager Suppression¶
Use the alertmanager.suppress configuration to block alerts based on labels. For example:
HighMemoryUsage alerts if HighCPUUsage is active.
Deduplication¶
Deduplication avoids sending the same alert multiple times for the same issue. Alertmanager uses label-based grouping to identify duplicates.
1. Label-Based Deduplication¶
Configure deduplicate in Alertmanager to group alerts by labels. For example:
instance and job combination.
2. Time-Based Deduplication¶
Use the for clause in Prometheus rules to delay alerts after a certain duration, reducing redundant notifications:
- alert: HighMemoryUsage
expr: ...
for: 10m
annotations:
summary: "High memory usage on {{ $labels.instance }} (持续 {{ $value }})"
Time-Based Suppression¶
Time-based suppression delays alerts after a specific period, preventing immediate flood of notifications. Use the time suppression option in Alertmanager:
Best Practices¶
- Consistent Labeling: Use standardized labels (e.g.,
job,instance) to enable deduplication and suppression. - Test Rules: Validate suppression and deduplication rules in staging environments to avoid unintended behavior.
- Monitor Metrics: Track suppression metrics (e.g.,
alertmanager_suppressed_alerts_total) to ensure rules are functioning as intended.
Diagram: Alert Suppression and Deduplication Flow¶
[Prometheus] --> [Alert Generation]
|
v
[Prometheus Suppression Rules] --> [Filtered Alerts]
|
v
[Alertmanager] --> [Deduplication]
|
v
[Time-Based Suppression] --> [Final Alerts]
|
v
[Notification Channels]
Key takeaways¶
- Use Prometheus suppression rules to block alerts based on other alert states.
- Configure label-based deduplication in Alertmanager to avoid redundant notifications.
- Implement time-based suppression to delay alerts and prevent floods.
- Prioritize consistent labeling and test suppression rules in staging environments.
- Monitor suppression metrics to ensure rules are functioning as intended.