2026-08-20 · 8 min read
Prometheus Alerting Rules That Don't Cry Wolf
Most alerting setups page too much and signal too little. Here's how to write Prometheus alert rules that are symptom-based, actionable, and tuned so on-call actually trusts them.

Prometheus Alerting Rules That Don't Cry Wolf
The fastest way to make on-call ignore alerts is to send too many. Once people learn that a page is probably noise, they stop reacting quickly to the one that isn't. Good alerting is less about catching everything and more about earning trust: every page should mean "a human needs to act now."
Alert on symptoms, not causes
The single most useful rule of thumb: alert on what users feel, not on every internal cause.
A high CPU number isn't an incident. Slow responses, elevated errors, or a stalled queue are, they map to user pain. If you alert on causes (CPU, memory, disk), you get paged for things that may never affect anyone, and you miss problems that don't show up as a resource metric.
Map alerts to the four golden signals: latency, traffic, errors, saturation. Page on the first three. Treat saturation mostly as a warning that feeds capacity planning, not a 3am page.
A symptom-based error-rate alert
groups:
- name: api-availability
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{job="api",status=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="api"}[5m]))
> 0.05
for: 10m
labels:
severity: page
annotations:
summary: "API 5xx error rate above 5% for 10m"
description: "{{ $value | humanizePercentage }} of requests are failing."
runbook: "https://github.com/durrello/incident-response-runbooks"
Two things make this an alert you can trust:
- The
for:clause. A 30-second blip won't page anyone: the condition has to hold for 10 minutes. This kills the vast majority of flapping. - A ratio, not a count. "More than 100 errors" pages you at 3am during a traffic spike that's perfectly healthy. "More than 5% of requests failing" scales with traffic.
Severity tiers: page vs ticket
Not everything that's wrong needs a human awake. Split severity:
page: user-facing and happening now. Wakes someone up. Error rate, latency SLO burn, total outage.ticket/warning: needs attention soon, but not at 3am. Disk filling slowly, a cert expiring in 14 days, a single replica down in a replicated service.
Route them differently in Alertmanager, page to PagerDuty/Opsgenie, warning to Slack or a
ticket queue.
route:
receiver: slack-warnings
routes:
- match: { severity: page }
receiver: pagerduty
group_wait: 30s
repeat_interval: 4h
Burn-rate alerts beat static thresholds
If you have an SLO (say 99.9% availability), alert on error-budget burn rate instead of a fixed error percentage. A multi-window burn-rate alert pages fast when you're burning budget catastrophically and slow when it's a gentle leak:
- alert: ErrorBudgetBurnFast
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
) > (14.4 * 0.001)
for: 2m
labels: { severity: page }
14.4 is the burn-rate multiplier that exhausts a 30-day budget in ~2 days, the standard fast-burn
threshold from the Google SRE workbook. Pair it with a slower window (e.g. 6x over 1h) so you catch
both the cliff and the slow leak.
Make every alert actionable
Before you ship an alert, ask: if this fires at 3am, is there something the responder can do? If the answer is "no, just watch it," it's not a page, it's a dashboard. Every paging alert should have:
- A clear summary that says what is wrong in user terms.
- A link to a runbook with first steps.
- Enough labels to know which service/region/tenant.
Tuning is a continuous job
Alerting isn't set-and-forget. Review fired alerts monthly:
- Which alerts paged but needed no action? Raise the threshold, lengthen
for:, or downgrade to a ticket. - Which incidents had no alert? Add one.
- Which alerts fire together? Group or add inhibition rules so one root cause = one page, not twelve.
A healthy on-call rotation has few pages, and almost every one is real. If your team is muting channels, the rules are the problem, not the people.
The short version
- Alert on symptoms (latency, errors, saturation), not raw causes.
- Use
for:and ratios to kill flapping and false alarms. - Split
pagefromticketand route them differently. - Prefer SLO burn-rate alerts over static thresholds.
- Every page must be actionable, with a runbook link.
- Review and prune monthly.
Need an alerting setup your team actually trusts? Observability and SRE practice is part of my consulting work, reach out.