← Back to blog

2026-08-20 · 8 min read

Prometheus Alerting Rules That Don't Cry Wolf

Most alerting setups page too much and signal too little. Here's how to write Prometheus alert rules that are symptom-based, actionable, and tuned so on-call actually trusts them.

#observability#prometheus#sre#alerting#monitoring
Prometheus Alerting Rules That Don't Cry Wolf

Prometheus Alerting Rules That Don't Cry Wolf

The fastest way to make on-call ignore alerts is to send too many. Once people learn that a page is probably noise, they stop reacting quickly to the one that isn't. Good alerting is less about catching everything and more about earning trust: every page should mean "a human needs to act now."

Alert on symptoms, not causes

The single most useful rule of thumb: alert on what users feel, not on every internal cause.

A high CPU number isn't an incident. Slow responses, elevated errors, or a stalled queue are, they map to user pain. If you alert on causes (CPU, memory, disk), you get paged for things that may never affect anyone, and you miss problems that don't show up as a resource metric.

Map alerts to the four golden signals: latency, traffic, errors, saturation. Page on the first three. Treat saturation mostly as a warning that feeds capacity planning, not a 3am page.

A symptom-based error-rate alert

groups:
  - name: api-availability
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{job="api",status=~"5.."}[5m]))
            / sum(rate(http_requests_total{job="api"}[5m]))
            > 0.05
        for: 10m
        labels:
          severity: page
        annotations:
          summary: "API 5xx error rate above 5% for 10m"
          description: "{{ $value | humanizePercentage }} of requests are failing."
          runbook: "https://github.com/durrello/incident-response-runbooks"

Two things make this an alert you can trust:

  • The for: clause. A 30-second blip won't page anyone: the condition has to hold for 10 minutes. This kills the vast majority of flapping.
  • A ratio, not a count. "More than 100 errors" pages you at 3am during a traffic spike that's perfectly healthy. "More than 5% of requests failing" scales with traffic.

Severity tiers: page vs ticket

Not everything that's wrong needs a human awake. Split severity:

  • page: user-facing and happening now. Wakes someone up. Error rate, latency SLO burn, total outage.
  • ticket / warning: needs attention soon, but not at 3am. Disk filling slowly, a cert expiring in 14 days, a single replica down in a replicated service.

Route them differently in Alertmanager, page to PagerDuty/Opsgenie, warning to Slack or a ticket queue.

route:
  receiver: slack-warnings
  routes:
    - match: { severity: page }
      receiver: pagerduty
      group_wait: 30s
      repeat_interval: 4h

Burn-rate alerts beat static thresholds

If you have an SLO (say 99.9% availability), alert on error-budget burn rate instead of a fixed error percentage. A multi-window burn-rate alert pages fast when you're burning budget catastrophically and slow when it's a gentle leak:

- alert: ErrorBudgetBurnFast
  expr: |
    (
      sum(rate(http_requests_total{status=~"5.."}[5m]))
        / sum(rate(http_requests_total[5m]))
    ) > (14.4 * 0.001)
  for: 2m
  labels: { severity: page }

14.4 is the burn-rate multiplier that exhausts a 30-day budget in ~2 days, the standard fast-burn threshold from the Google SRE workbook. Pair it with a slower window (e.g. 6x over 1h) so you catch both the cliff and the slow leak.

Make every alert actionable

Before you ship an alert, ask: if this fires at 3am, is there something the responder can do? If the answer is "no, just watch it," it's not a page, it's a dashboard. Every paging alert should have:

  • A clear summary that says what is wrong in user terms.
  • A link to a runbook with first steps.
  • Enough labels to know which service/region/tenant.

Tuning is a continuous job

Alerting isn't set-and-forget. Review fired alerts monthly:

  • Which alerts paged but needed no action? Raise the threshold, lengthen for:, or downgrade to a ticket.
  • Which incidents had no alert? Add one.
  • Which alerts fire together? Group or add inhibition rules so one root cause = one page, not twelve.

A healthy on-call rotation has few pages, and almost every one is real. If your team is muting channels, the rules are the problem, not the people.

The short version

  • Alert on symptoms (latency, errors, saturation), not raw causes.
  • Use for: and ratios to kill flapping and false alarms.
  • Split page from ticket and route them differently.
  • Prefer SLO burn-rate alerts over static thresholds.
  • Every page must be actionable, with a runbook link.
  • Review and prune monthly.

Need an alerting setup your team actually trusts? Observability and SRE practice is part of my consulting work, reach out.

Share:LinkedInXWhatsApp

Related articles

Reactions & comments