← blog · September 1, 2026

Alerting Rules: "Something Broke" vs. "Users Are Affected"

Cause-based alerts (high CPU, full disk, failed job) wear out on-call engineers without answering whether users are actually affected. How to move to SLO-based burn rate alerting, and the traps along the way.

In any system you can build two kinds of alerts. The first says "this component is behaving abnormally": CPU crossed ninety percent, disk filled up, the queue backlog grew, a cron job exited non-zero. The second says "this group of users is having a bad time right now": the request failure rate crossed a defined threshold, the latency distribution moved outside an acceptable range. Most teams pile up alerts in the first category because it is easy to set up and the metrics are already there. The result is a familiar story: the on-call phone rings at 3am over a disk alert, half an hour gets spent, and it turns out disk usage climbed to eighty five percent while not a single user request failed. In the morning nobody can answer whether anything actually happened last night.

The two approaches are not interchangeable, but mixing them up makes both useless. Cause-based alerting, used correctly, gives early warning: you see disk growth before it fills up, you catch a memory leak before it reaches production traffic. Used incorrectly it produces alert fatigue. A team that gets fifty notifications a day stops reading by the forty ninth and loses the ability to tell which one is real. This is the consistent finding across large-scale on-call reports published in recent years: as notification volume rises, average response time does not shrink, it grows.

Making user impact measurable

The precondition for outcome-based alerting is turning the phrase "users are affected" into a number. Three terms from site reliability engineering practice do the job: the Service Level Indicator (SLI), its target value (SLO), and the failure budget the target allows (error budget). Typical SLIs for a web service are the fraction of successful requests, the percentage of requests completing under a given latency threshold, and the rate at which a queued job gets processed within a defined time window. The SLO is the target that indicator must hold over a rolling window, say a thirty day period, for example 99.9 percent success. The remaining 0.1 percent is the error budget the team consciously accepts, and it gets spent on decisions like taking deploy risk or opening a maintenance window.

The critical distinction is this: CPU at ninety percent is not an SLI, because its direct relationship to user experience has never been established. A system can run at ninety percent CPU while every request still returns correctly and on time. An SLI is what the user directly feels: did the request succeed, how long did it take, was the data correct.

Burn rate alerting: when and how severe

Turning an SLO into a single-threshold alert ("warn if monthly error rate exceeds 0.1 percent") does not work in practice, because you would have to wait until the end of the month, and you would miss a fast degradation entirely. The technique used instead is multi-window, multi-burn-rate alerting. The idea is simple: measure how fast your error budget is being consumed. The budget is designed to be spent over thirty days under normal conditions; if the current error rate would exhaust it within one hour, that is a burn rate forty four times the normal pace and demands immediate action. The same calculation is repeated over six hour and one day windows to catch slower, sustained degradation as well.

A sample rule in Prometheus looks like this:

groups:
  - name: slo-burn-rate
    rules:
      - alert: ErrorBudgetBurnFast
        expr: |
          (
            sum(rate(http_requests_total{code=~"5.."}[1h]))
            /
            sum(rate(http_requests_total[1h]))
          ) > (14.4 * 0.001)
          and
          (
            sum(rate(http_requests_total{code=~"5.."}[5m]))
            /
            sum(rate(http_requests_total[5m]))
          ) > (14.4 * 0.001)
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "Error budget burning fast, 2% of the monthly budget spent in the last hour"

Here 0.001 represents the monthly target error rate (a 99.9 percent SLO), and the 14.4 multiplier represents the burn coefficient at which the entire budget would be exhausted in roughly two days at this rate. The short window (5 minutes) is checked alongside the long window (1 hour) so a brief spike does not trigger a false page, while a genuine degradation still cannot hide for a full hour unnoticed. Slower burn rates (6x, 3x, 1x) follow the same pattern with longer windows and are typically routed to ticket-level notification rather than paging.

Do not throw away cause-based alerting

A common mistake once teams move to outcome-based alerting is dropping cause-based monitoring entirely. The two serve different purposes. An outcome-based alert tells you "act now" and carries the authority to wake someone up. Cause-based metrics are used after the page fires, to find the root cause: disk usage, queue depth, connection pool saturation, retry counts. The correct architecture routes cause-based signals to dashboards and low-priority notifications, and reserves paging for signals that prove actual user impact. In Alertmanager this separation is done with a severity label: rules labeled severity=page route to a receiver that reaches the on-call engineer directly, while severity=ticket rules route to the issue tracker.

Traps

The first trap is the "warn and continue" pattern that swallows a failure and keeps going. When a step fails but the rest of the pipeline keeps running, and the only symptom is a log line, that log line gets read by nobody and the dashboard stays green. At every such point two questions need answers: is continuing while this step is broken actually better than not running at all, and if you continue, can you measure what data the next step is now operating on? If both answers are no, the pipeline should stop and make noise.

The second trap is never verifying that the alert actually fires. Computing an SLI from the wrong metric source, querying it with the wrong label set, or reading data written into a different namespace than the one being queried, silently turns the alert into something that will never trigger. The only reliable test is deliberately producing a real degradation: inject errors into test traffic and confirm the alert actually fires, routes to the right channel, and triggers at the correct threshold. This is proving the monitoring chain works end to end by causing the failure in an isolated setting; it is not a one-time setup step, it is something to repeat every time the alerting logic changes.

The third trap is picking windows too narrow or thresholds too loose. A window that is too narrow triggers on ordinary traffic spikes, and the on-call engineer learns to ignore it; a threshold too loose lets a real outage go unnoticed for hours. Window and multiplier choices should be backtested against historical traffic: look at when the new rule would have fired during past real incidents, and eliminate rules that fire either too late or too early.

When not to adopt this approach

This investment is not necessary for every system. An early-stage product with low traffic and few users does not have enough sample size to compute a meaningful SLI; a 99.9 percent target on a service handling a few hundred requests a day cannot be statistically distinguished from noise. In that situation simple threshold alerts (is the service up, do core endpoints respond) are sufficient, and the return on building SLO infrastructure does not justify the engineering time spent. As traffic and user count grow, especially once multiple teams own different parts of the same system, impact-based alert prioritization starts to show its real value.