Skip to content

Measuring SLIs

Defining and Measuring SLIs

Service Level Indicators (SLIs) are the foundational metrics that quantify the reliability and performance of a system. They provide objective, measurable data to evaluate whether a service meets its Service Level Objectives (SLOs) and help teams identify areas for improvement. Selecting, measuring, and monitoring SLIs effectively is critical to maintaining system reliability and aligning engineering efforts with business goals.


Selecting the Right SLIs

SLIs should reflect the user experience and business-critical aspects of your system. Common SLI categories include:

  1. Availability
  2. Measures uptime or the percentage of time a service is operational.
  3. Example: http_2xx_requests_total (number of successful HTTP responses).

  4. Latency

  5. Tracks the time taken to fulfill a request.
  6. Example: http_request_duration_seconds (average or percentile latency).

  7. Error Rate

  8. Quantifies the proportion of failed requests.
  9. Example: http_5xx_requests_total (number of 5xx errors).

  10. Request Volume

  11. Measures the number of requests processed, useful for capacity planning.
  12. Example: http_requests_total (total requests per second).

Key Considerations:
- Prioritize SLIs that align with user expectations and business priorities. For example, a payment gateway might prioritize error rate over latency.
- Avoid overloading with metrics; focus on 2–3 core SLIs that capture the most critical aspects of reliability.
- Use business context to define thresholds. For instance, a healthcare application may require stricter availability guarantees than a public-facing blog.


Measuring SLIs

Accurate measurement requires consistent data collection, aggregation, and normalization.

Tools and Techniques

  • Prometheus: Collect metrics via exporters (e.g., HTTP, MySQL, Kafka). Use query language (PromQL) to calculate SLI values.
    # Example: Error rate as a percentage
    (sum(rate(http_5xx_requests_total[5m])) / sum(rate(http_requests_total[5m]))) * 100
    
  • Datadog/Grafana: Visualize metrics and set up alerts based on SLO thresholds.
  • Sampling and Aggregation: Use appropriate sampling rates (e.g., 1 in 1000 requests) to balance accuracy and resource usage. Aggregate data over time windows (e.g., 1-minute intervals) for stability.

SLI Calculation Examples

  • Availability:
    (Total requests - Failed requests) / Total requests
    
  • Latency:
    Percentile of request duration (e.g., 95th percentile)
    
  • Error Rate:
    (Number of errors / Total requests) * 100
    

Monitoring SLIs

Effective monitoring ensures SLIs are visible, actionable, and aligned with SLOs.

Dashboard Design

  • Use Grafana or Kibana to create dashboards with:
  • Time-series graphs for latency, error rates, and availability.
  • Threshold lines for SLO boundaries (e.g., 99% availability).
  • Alerts for deviations (e.g., error rate exceeding 1% for 5 minutes).

Alerting Strategies

  • Define SLO thresholds (e.g., 99% availability) and trigger alerts when SLIs fall below them.
  • Use contextual alerts to avoid noise. For example, alert only if an error rate spike persists for 10 minutes.
  • Example Prometheus alert rule:
    - alert: HighErrorRate
      expr: (sum(rate(http_5xx_requests_total[5m])) / sum(rate(http_requests_total[5m]))) * 100 > 1
      for: 10m
      labels:
        severity: warning
      annotations:
        summary: "High error rate detected"
        description: "Error rate exceeds 1% for the last 10 minutes."
    

Diagram: SLI Measurement Pipeline

[User Requests] → [Metrics Exporter] → [Prometheus] → [Grafana Dashboard] → [Alerting System]

Key takeaways

  • Select SLIs that directly reflect user experience and business priorities (e.g., error rate for critical systems).
  • Measure SLIs using tools like Prometheus and Grafana, ensuring consistent aggregation and sampling.
  • Monitor SLIs with dashboards and alerts to proactively identify deviations from SLOs.
  • Align SLI definitions with business goals to ensure reliability efforts are impactful and measurable.