π PCA EXAM β DOMAIN 7: ALERTING & ALERTMANAGER (4%)
Complete One-Stop Guide: Theory β Examples β Exam Questions
π This is the FINAL domain! After this, youβll have 100% exam coverage!
TABLE OF CONTENTS
7.1 Alerting Architecture Overview
7.2 Alerting Rules Syntax
7.3 Alert States (Inactive β Pending β Firing)
7.4 Alert Rule Examples (Common Patterns)
7.5 Alertmanager β Overview & Architecture
7.6 Alertmanager Configuration (alertmanager.yml)
7.7 Routing Tree (Deep Dive)
7.8 Grouping (group_by, group_wait, group_interval)
7.9 Receivers (Slack, Email, PagerDuty, Webhook)
7.10 Inhibition Rules
7.11 Silencing
7.12 Notification Templates
7.13 Alertmanager High Availability
7.14 Real-World Alerting Scenarios
7.15 EXAM-STYLE QUESTIONS (20 Questions with Answers)
7.1 π ALERTING ARCHITECTURE OVERVIEW
The Complete Alerting Flow
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PROMETHEUS ALERTING FLOW β
β β
β Step 1: Prometheus evaluates alerting rules β
β (every evaluation_interval, default 1m) β
β β
β ββββββββββββββββββββββββββββββββββββββββββββββββ β
β β PROMETHEUS SERVER β β
β β β β
β β Rule Manager: β β
β β Evaluates: up == 0 β β
β β Result: 3 targets are down β β
β β State: PENDING (for: 5m not yet reached) β β
β β ... 5 minutes later ... β β
β β State: FIRING π₯ β β
β β β β
β β ββββ HTTP POST to Alertmanager βββββΊ β β
β βββββββββββββββββββββββββββββββββββββββββ¬ββββββββ β
β β β
β Step 2: Alertmanager receives alerts β β
β βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββ β
β β ALERTMANAGER β β
β β β β
β β 1. Deduplication (same alert from replicas) β β
β β 2. Grouping (group related alerts) β β
β β 3. Inhibition (suppress dependent alerts) β β
β β 4. Silencing (mute during maintenance) β β
β β 5. Routing (send to correct receiver) β β
β β 6. Notification (Slack, Email, PagerDuty) β β
β β β β
β ββββββββ¬βββββββββββ¬βββββββββββ¬ββββββββββββββββββ β
β β β β β
β βΌ βΌ βΌ β
β ββββββββββββ ββββββββββββ ββββββββββββ β
β β Slack β β Email β β PagerDutyβ β
β β #alerts β β ops@co β β on-call β β
β ββββββββββββ ββββββββββββ ββββββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Key Components
Prometheus Server:
β Evaluates alerting rules (PromQL expressions)
β Tracks alert states (Inactive, Pending, Firing)
β Sends FIRING alerts to Alertmanager via HTTP API
Alertmanager (Separate Binary!):
β Receives alerts from Prometheus
β Groups, deduplicates, routes, and sends notifications
β Manages silences and inhibition rules
β Runs on port 9093 (default)
β Has its own web UI at http://alertmanager:9093
7.2 π ALERTING RULES SYNTAX
Complete Syntax
# alerting_rules.yml
groups:
- name: example_alerts
interval: 30s # Optional: override evaluation_interval for this group
rules:
- alert: <AlertName> # Required: unique alert name
expr: <PromQL Expression> # Required: condition to evaluate
for: <Duration> # Optional: how long condition must be true
labels: # Optional: additional labels
<label_name>: <label_value>
annotations: # Optional: human-readable info
<annotation_name>: <annotation_value>
Detailed Breakdown
- alert: InstanceDown # Alert name (CamelCase convention)
expr: up == 0 # PromQL: fires when target is down
for: 5m # Must be true for 5 minutes
labels:
severity: critical # Added to the alert as a label
team: infrastructure # Used for routing in Alertmanager
annotations:
summary: "Instance {{ $labels.instance }} is down"
description: "{{ $labels.instance }} of job {{ $labels.job }} has been down for more than 5 minutes."
runbook_url: "https://wiki.example.com/runbooks/instance-down"
Template Variables in Annotations
annotations:
# $labels β access label values from the alerting expression
summary: "Instance {{ $labels.instance }} is down"
# If instance="app1:8080", result: "Instance app1:8080 is down"
# $value β the value of the PromQL expression
description: "CPU usage is {{ $value }}%"
# If expr evaluates to 95.5, result: "CPU usage is 95.5%"
# $externalLabels β access external_labels from prometheus.yml
info: "Cluster: {{ $externalLabels.cluster }}"
# Humanize functions
description: "Memory usage: {{ $value | humanizePercentage }}"
# Converts 0.85 to "85%"
description: "Duration: {{ $value | humanizeDuration }}"
# Converts 3661 to "1h 1m 1s"
Configuration in prometheus.yml
# prometheus.yml
rule_files:
- "alerting_rules.yml"
- "recording_rules.yml"
- "/etc/prometheus/rules/*.yml"
alerting:
alertmanagers:
- static_configs:
- targets:
- 'alertmanager1:9093'
- 'alertmanager2:9093'
scheme: http
path_prefix: /
timeout: 10s
api_version: v2
7.3 π ALERT STATES (Inactive β Pending β Firing)
The Three States
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β β
β INACTIVE ββ(expr becomes true)βββΊ PENDING ββ(for: X)βββΊ FIRING
β β β β
β β β (expr becomes false) β
β β βΌ β
β ββββββββββββββββββββββββββββ INACTIVE β
β β β
β β (expr becomes false while FIRING) β
β ββββββββββββββββββββββββββββββββββββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
State Details
1. INACTIVE (Default)
β The alert condition (expr) is FALSE
β No alert is generated
β This is the normal state for most alerts most of the time
β Example: up == 1 β InstanceDown alert is INACTIVE
2. PENDING
β The alert condition (expr) has become TRUE
β BUT the "for" duration has NOT yet elapsed
β The alert is "waiting" to confirm the issue is persistent
β If the condition becomes FALSE before "for" elapses β back to INACTIVE
β If no "for" is specified β skips PENDING, goes directly to FIRING
Example:
expr: up == 0
for: 5m
t=0: up == 0 β TRUE β Alert enters PENDING
t=2m: up == 0 β still TRUE β still PENDING
t=3m: up == 1 β FALSE β Alert returns to INACTIVE (false alarm!)
OR:
t=0: up == 0 β TRUE β PENDING
t=5m: up == 0 β still TRUE β Alert transitions to FIRING π₯
3. FIRING
β The alert condition has been TRUE for the entire "for" duration
β Prometheus sends the alert to Alertmanager
β Alertmanager processes and routes the notification
β The alert remains FIRING as long as the condition is TRUE
β When the condition becomes FALSE β alert resolves β INACTIVE
Example:
t=0: up == 0 β PENDING
t=5m: up == 0 β FIRING π₯ (sent to Alertmanager)
t=10m: up == 0 β still FIRING (Alertmanager may re-notify)
t=12m: up == 1 β RESOLVED β INACTIVE (resolution sent to Alertmanager)
The for Duration
# With "for" β prevents flapping alerts
- alert: HighCPU
expr: cpu_usage > 90
for: 10m # CPU must be > 90% for 10 continuous minutes
# Prevents alerts from brief CPU spikes
# Without "for" β fires immediately
- alert: ServiceDown
expr: up == 0
# No "for" β fires immediately when condition is true
# Use for critical alerts that need instant notification
π‘ Key Takeaway for Exam
Inactive = condition is false (normal state) Pending = condition is true but
forduration not yet elapsed Firing = condition true for entireforduration β sent to Alertmanager Nofor= skips Pending, fires immediately Resolved = condition becomes false while Firing β back to Inactive
7.4 π ALERT RULE EXAMPLES (Common Patterns)
1. Instance Down (Most Common)
- alert: InstanceDown
expr: up == 0
for: 5m
labels:
severity: critical
annotations:
summary: "Instance {{ $labels.instance }} is down"
description: "{{ $labels.instance }} of job {{ $labels.job }} has been down for more than 5 minutes."
2. High CPU Usage
- alert: HighCPUUsage
expr: 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 10m
labels:
severity: warning
annotations:
summary: "High CPU usage on {{ $labels.instance }}"
description: "CPU usage is {{ $value | humanize }}% on {{ $labels.instance }}."
3. High Memory Usage
- alert: HighMemoryUsage
expr: (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100 > 90
for: 5m
labels:
severity: critical
annotations:
summary: "High memory usage on {{ $labels.instance }}"
description: "Memory usage is {{ $value | humanize }}% on {{ $labels.instance }}."
4. Disk Space Running Out
- alert: DiskSpaceLow
expr: (1 - (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"})) * 100 > 85
for: 15m
labels:
severity: warning
annotations:
summary: "Low disk space on {{ $labels.instance }}"
description: "Disk usage is {{ $value | humanize }}% on {{ $labels.instance }}."
- alert: DiskWillFillIn24Hours
expr: predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 24 * 3600) < 0
for: 30m
labels:
severity: critical
annotations:
summary: "Disk will fill within 24 hours on {{ $labels.instance }}"
5. High HTTP Error Rate
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
/
sum(rate(http_requests_total[5m])) by (job)
* 100 > 5
for: 5m
labels:
severity: critical
annotations:
summary: "High error rate for {{ $labels.job }}"
description: "Error rate is {{ $value | humanize }}% for job {{ $labels.job }}."
6. High Latency
- alert: HighLatency
expr: |
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le, job)
) > 1
for: 10m
labels:
severity: warning
annotations:
summary: "High p99 latency for {{ $labels.job }}"
description: "p99 latency is {{ $value | humanizeDuration }} for {{ $labels.job }}."
7. SSL Certificate Expiring
- alert: SSLCertExpiringSoon
expr: (probe_ssl_earliest_cert_expiry - time()) / 86400 < 30
for: 1h
labels:
severity: warning
annotations:
summary: "SSL certificate expiring in {{ $value | humanize }} days"
description: "SSL certificate for {{ $labels.instance }} expires in less than 30 days."
8. Dead Manβs Switch (Watchdog)
- alert: Watchdog
expr: vector(1)
labels:
severity: none
annotations:
summary: "This alert should always be firing"
description: "If this alert stops firing, the alerting pipeline is broken!"
# This alert is ALWAYS firing (vector(1) = 1)
# If you stop receiving this alert, Alertmanager or Prometheus is broken
7.5 π ALERTMANAGER β OVERVIEW & ARCHITECTURE
What is Alertmanager?
Alertmanager is a separate component (not part of Prometheus server) that handles alerts sent by Prometheus. It deduplicates, groups, routes, and sends notifications.
Key Facts
Binary: alertmanager
Port: 9093 (default)
UI: http://alertmanager:9093
Config: alertmanager.yml
Maintainer: Prometheus Team
Alertmanager Responsibilities
1. DEDUPLICATION:
β Multiple Prometheus replicas may send the same alert
β Alertmanager deduplicates based on alert fingerprint
β Only one notification per unique alert
2. GROUPING:
β Groups related alerts into a single notification
β Example: 50 pods down β one notification "50 pods down in namespace X"
β Prevents notification storms
3. INHIBITION:
β Suppresses alerts when a "parent" alert is already firing
β Example: If "Cluster Down" fires, inhibit "Node Down" alerts
4. SILENCING:
β Temporarily mutes alerts matching specific matchers
β Used during planned maintenance
β Configured via Alertmanager UI or API
5. ROUTING:
β Routes alerts to the correct receiver based on labels
β Example: severity=critical β PagerDuty, severity=warning β Slack
6. NOTIFICATION:
β Sends alerts via various channels
β Supports: Email, Slack, PagerDuty, OpsGenie, Webhook, etc.
7.6 π ALERTMANAGER CONFIGURATION (alertmanager.yml)
Complete Configuration Structure
# alertmanager.yml
# ββββββββββββββββββββββββββββββββββββββββββββββ
# GLOBAL CONFIGURATION
# ββββββββββββββββββββββββββββββββββββββββββββββ
global:
resolve_timeout: 5m # Time to wait before declaring alert resolved
smtp_smarthost: 'smtp.example.com:587'
smtp_from: 'alertmanager@example.com'
smtp_auth_username: 'alertmanager@example.com'
smtp_auth_password: 'password'
smtp_require_tls: true
slack_api_url: 'https://hooks.slack.com/services/T00/B00/XXX'
pagerduty_url: 'https://events.pagerduty.com/v2/enqueue'
# ββββββββββββββββββββββββββββββββββββββββββββββ
# TEMPLATES
# ββββββββββββββββββββββββββββββββββββββββββββββ
templates:
- '/etc/alertmanager/templates/*.tmpl'
# ββββββββββββββββββββββββββββββββββββββββββββββ
# ROUTING TREE
# ββββββββββββββββββββββββββββββββββββββββββββββ
route:
receiver: 'default-receiver' # Default receiver (catch-all)
group_by: ['alertname', 'cluster'] # Group alerts by these labels
group_wait: 30s # Wait before sending first notification
group_interval: 5m # Wait before sending updated notification
repeat_interval: 4h # Wait before re-sending same alert
routes:
# Critical alerts β PagerDuty
- match:
severity: critical
receiver: 'pagerduty-critical'
group_wait: 10s # Faster for critical
repeat_interval: 1h # Re-notify every hour
# Warning alerts β Slack
- match:
severity: warning
receiver: 'slack-warnings'
repeat_interval: 4h
# Database alerts β DB team
- match_re:
job: 'mysql|postgres|redis'
receiver: 'db-team-slack'
# Watchdog β dead man's switch (no notification)
- match:
alertname: Watchdog
receiver: 'null' # Drop/sink receiver
# ββββββββββββββββββββββββββββββββββββββββββββββ
# RECEIVERS
# ββββββββββββββββββββββββββββββββββββββββββββββ
receivers:
- name: 'default-receiver'
email_configs:
- to: 'ops-team@example.com'
send_resolved: true
- name: 'pagerduty-critical'
pagerduty_configs:
- service_key: 'your-pagerduty-integration-key'
severity: '{{ .CommonLabels.severity }}'
description: '{{ .CommonAnnotations.summary }}'
- name: 'slack-warnings'
slack_configs:
- channel: '#alerts-warning'
title: '{{ .CommonLabels.alertname }}'
text: '{{ .CommonAnnotations.description }}'
send_resolved: true
- name: 'db-team-slack'
slack_configs:
- channel: '#db-alerts'
title: 'Database Alert: {{ .CommonLabels.alertname }}'
- name: 'null'
# Empty receiver β discards alerts (for Watchdog)
# ββββββββββββββββββββββββββββββββββββββββββββββ
# INHIBITION RULES
# ββββββββββββββββββββββββββββββββββββββββββββββ
inhibit_rules:
# If "ClusterDown" fires, inhibit "NodeDown" alerts
- source_match:
alertname: ClusterDown
target_match:
alertname: NodeDown
equal: ['cluster']
# If "critical" fires, inhibit "warning" for same alertname
- source_match:
severity: critical
target_match:
severity: warning
equal: ['alertname', 'instance']
7.7 π ROUTING TREE (Deep Dive)
How Routing Works
Alert arrives at Alertmanager
β
βΌ
βββββββββββββββββββββββββββ
β ROOT ROUTE β
β receiver: default β
β group_by: [alertname] β
β β
β Does alert match any β
β child route? β
β βββββ¬ββββ¬ββββ β
β β β β β β
β βΌ βΌ βΌ βΌ β
β R1 R2 R3 R4 β
β sev= sev= job= alert= β
β crit warn db Watch β
β β β β β β
β βΌ βΌ βΌ βΌ β
β PD Slack DB null β
βββββββββββββββββββββββββββ
If NO child route matches β use ROOT route's receiver (default)
If a child route matches β use that route's receiver
If continue: true β also check sibling routes
Route Matching
route:
receiver: 'default'
routes:
# Exact match
- match:
severity: critical
team: backend
receiver: 'backend-critical'
# Alert must have BOTH severity=critical AND team=backend
# Regex match
- match_re:
job: 'mysql|postgres|redis'
receiver: 'db-team'
# Alert job must match the regex
# Continue to next route even if this matches
- match:
severity: critical
receiver: 'pagerduty'
continue: true # β Also check next routes!
- match:
team: backend
receiver: 'backend-slack'
# With continue: true above, a critical backend alert
# goes to BOTH pagerduty AND backend-slack
The continue Flag
# Without continue (default: false):
# First matching route wins, stops checking
routes:
- match: {severity: critical}
receiver: pagerduty # β Alert goes here, stops
- match: {team: backend}
receiver: backend-slack # β Never reached for critical alerts
# With continue: true:
# Keeps checking sibling routes after a match
routes:
- match: {severity: critical}
receiver: pagerduty
continue: true # β Keep checking!
- match: {team: backend}
receiver: backend-slack # β Also reached if team=backend
# Result: critical backend alert β BOTH pagerduty AND backend-slack
7.8 π GROUPING
Why Grouping?
Problem: 100 pods go down simultaneously
β Without grouping: 100 separate Slack messages! πππ...
β With grouping: 1 message "100 pods down in namespace production"
Grouping reduces notification noise by combining related alerts.
Grouping Parameters
route:
group_by: ['alertname', 'cluster', 'namespace']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
group_by β What to Group On
group_by: ['alertname', 'cluster']
# Alerts with the SAME alertname AND cluster are grouped together
# Example:
# Alert 1: {alertname="InstanceDown", cluster="prod", instance="app1"}
# Alert 2: {alertname="InstanceDown", cluster="prod", instance="app2"}
# Alert 3: {alertname="InstanceDown", cluster="staging", instance="app3"}
#
# Group 1: InstanceDown + prod β [app1, app2] (one notification)
# Group 2: InstanceDown + staging β [app3] (separate notification)
group_by: ['...']
# Special: group ALL alerts into a single group (use with caution!)
group_wait β Initial Wait
group_wait: 30s
# When a new group of alerts is created, wait 30s before sending
# the first notification. This allows more alerts to accumulate
# in the group before notifying.
# Use case: 50 pods crash at the same time
# β Wait 30s β collect all 50 alerts β send ONE notification
# β Instead of sending 50 separate notifications
group_interval β Update Wait
group_interval: 5m
# After the first notification, wait 5m before sending an UPDATE
# if new alerts have been added to the group.
# Use case:
# t=0: 10 pods down β wait group_wait β notify "10 pods down"
# t=2m: 5 more pods down β added to group
# t=5m: group_interval elapsed β notify "15 pods down" (update)
repeat_interval β Re-notification
repeat_interval: 4h
# If an alert is still firing after 4 hours, re-send the notification.
# This ensures on-call engineers don't forget about ongoing issues.
# Use case:
# t=0: Alert fires β notify
# t=4h: Alert still firing β re-notify (reminder)
# t=8h: Alert still firing β re-notify again
Timing Diagram
Alert fires at t=0
β
βΌ
t=0 ββββ group_wait (30s) βββββΊ First notification sent at t=30s
β
New alerts added to group β
β β
βΌ βΌ
t=5m ββ group_interval (5m) βββΊ Update notification at t=5m30s
β
Alert still firing β
β β
βΌ βΌ
t=4h ββ repeat_interval (4h) βββΊ Re-notification at t=4h
β
Alert resolves β
β β
βΌ βΌ
t=4h30m βββββββββββββββββββββββΊ Resolution notification
7.9 π RECEIVERS
Supported Receiver Types
| Receiver | Use Case | Config Key |
|---|---|---|
| General notifications | email_configs |
|
| Slack | Team chat notifications | slack_configs |
| PagerDuty | On-call escalation | pagerduty_configs |
| OpsGenie | On-call management | opsgenie_configs |
| Webhook | Custom integrations | webhook_configs |
| VictorOps | Incident management | victorops_configs |
| Pushover | Mobile push notifications | pushover_configs |
| Telegram | Telegram bot messages | telegram_configs |
| Microsoft Teams | Teams channel (via webhook) | msteams_configs |
| SNS | AWS SNS (SMS/Email) | sns_configs |
Email Receiver
receivers:
- name: 'email-team'
email_configs:
- to: 'ops-team@example.com'
from: 'alertmanager@example.com'
smarthost: 'smtp.example.com:587'
auth_username: 'alertmanager@example.com'
auth_password: 'password'
require_tls: true
send_resolved: true
headers:
subject: '[{{ .Status | toUpper }}] {{ .CommonLabels.alertname }}'
Slack Receiver
receivers:
- name: 'slack-alerts'
slack_configs:
- api_url: 'https://hooks.slack.com/services/T00/B00/XXX'
channel: '#alerts'
username: 'Prometheus Alertmanager'
icon_emoji: ':fire:'
title: '[{{ .Status | toUpper }}] {{ .CommonLabels.alertname }}'
text: >-
*Description:* {{ .CommonAnnotations.description }}
*Severity:* {{ .CommonLabels.severity }}
*Instance:* {{ .CommonLabels.instance }}
send_resolved: true
color: '{{ if eq .Status "firing" }}danger{{ else }}good{{ end }}'
PagerDuty Receiver
receivers:
- name: 'pagerduty-critical'
pagerduty_configs:
- routing_key: 'your-pagerduty-integration-key'
severity: '{{ .CommonLabels.severity }}'
description: '{{ .CommonAnnotations.summary }}'
details:
firing: '{{ template "pagerduty.default.description" . }}'
send_resolved: true
Webhook Receiver
receivers:
- name: 'custom-webhook'
webhook_configs:
- url: 'http://my-service:5001/alerts'
send_resolved: true
http_config:
basic_auth:
username: 'admin'
password: 'secret'
The send_resolved Flag
send_resolved: true # Send notification when alert resolves (back to normal)
send_resolved: false # Only notify when alert fires (default for some receivers)
# Best practice: Enable for critical alerts so you know when the issue is fixed
7.10 π INHIBITION RULES
What is Inhibition?
Inhibition suppresses (mutes) certain alerts when other related alerts are already firing. This prevents notification noise from cascading failures.
Syntax
inhibit_rules:
- source_match: # The "parent" alert (must be firing)
alertname: ClusterDown
severity: critical
target_match: # The "child" alert (will be suppressed)
alertname: NodeDown
equal: ['cluster'] # Labels that must match between source and target
How It Works
Example:
Source alert: {alertname="ClusterDown", cluster="prod-us-east"}
Target alert: {alertname="NodeDown", cluster="prod-us-east", instance="node1"}
Rule:
source_match: {alertname: ClusterDown}
target_match: {alertname: NodeDown}
equal: [cluster]
Result:
β ClusterDown is FIRING for cluster="prod-us-east"
β NodeDown for cluster="prod-us-east" is INHIBITED (suppressed)
β NodeDown for cluster="prod-eu-west" is NOT inhibited (different cluster)
β You only get ONE notification: "Cluster prod-us-east is down"
β Instead of 50 notifications: "Cluster down" + 49Γ "Node down"
Common Inhibition Patterns
inhibit_rules:
# 1. Critical suppresses Warning (same alert)
- source_match:
severity: critical
target_match:
severity: warning
equal: ['alertname', 'instance']
# 2. Cluster down suppresses Node down
- source_match:
alertname: ClusterDown
target_match:
alertname: NodeDown
equal: ['cluster']
# 3. Node down suppresses Pod down
- source_match:
alertname: NodeDown
target_match:
alertname: PodDown
equal: ['instance']
# 4. Network partition suppresses service unreachable
- source_match:
alertname: NetworkPartition
target_match_re:
alertname: 'ServiceUnreachable|HighLatency'
equal: ['datacenter']
π‘ Key Takeaway for Exam
Inhibition = Suppress child alerts when parent alert is firing
source_match= the alert that must be firing (parent)target_match= the alert to suppress (child)equal= labels that must match between source and target Prevents notification cascades (cluster down β node down β pod down)
7.11 π SILENCING
What is Silencing?
Silencing temporarily mutes alerts matching specific matchers. Unlike inhibition (which is config-based and permanent), silences are temporary and managed via the Alertmanager UI or API.
Use Cases
β
Planned maintenance windows
β "We're upgrading the database tonight, mute DB alerts for 4 hours"
β
Known issues being worked on
β "We know about the disk space issue, muting until fix is deployed"
β
Testing
β "Muting alerts while testing new alerting rules"
Creating a Silence (via Alertmanager UI)
1. Go to http://alertmanager:9093
2. Click "Silences" β "New Silence"
3. Configure:
β Matchers: alertname="DiskSpaceLow", instance="db1:9100"
β Starts at: 2024-11-15 22:00 UTC
β Ends at: 2024-11-16 02:00 UTC
β Created by: admin
β Comment: "Planned disk expansion maintenance"
4. Click "Create"
Creating a Silence (via API)
curl -X POST http://alertmanager:9093/api/v2/silences \
-H 'Content-Type: application/json' \
-d '{
"matchers": [
{"name": "alertname", "value": "DiskSpaceLow", "isRegex": false},
{"name": "instance", "value": "db1:9100", "isRegex": false}
],
"startsAt": "2024-11-15T22:00:00Z",
"endsAt": "2024-11-16T02:00:00Z",
"createdBy": "admin",
"comment": "Planned maintenance"
}'
Silence vs Inhibition
| Feature | Silence | Inhibition |
|---|---|---|
| Configuration | UI/API (temporary) | alertmanager.yml (permanent) |
| Duration | Time-limited (start/end) | Always active |
| Trigger | Manual | Automatic (based on parent alert) |
| Use case | Maintenance windows | Cascading failure suppression |
| Persistence | Stored in Alertmanager | In config file |
7.12 π NOTIFICATION TEMPLATES
Template Variables
.Status β "firing" or "resolved"
.Alerts β List of all alerts in the group
.Alerts.Firing β Only firing alerts
.Alerts.Resolved β Only resolved alerts
.CommonLabels β Labels common to ALL alerts in the group
.CommonAnnotations β Annotations common to ALL alerts
.ExternalURL β Alertmanager URL
.GroupLabels β Labels used for grouping
.Receiver β Name of the receiver
Example Template
receivers:
- name: 'slack-detailed'
slack_configs:
- channel: '#alerts'
title: >-
[{{ .Status | toUpper }}{{ if eq .Status "firing" }}:{{ .Alerts.Firing | len }}{{ end }}]
{{ .CommonLabels.alertname }}
text: >-
{{ range .Alerts }}
*Alert:* {{ .Labels.alertname }}
*Severity:* {{ .Labels.severity }}
*Instance:* {{ .Labels.instance }}
*Description:* {{ .Annotations.description }}
*Started:* {{ .StartsAt }}
{{ end }}
send_resolved: true
7.13 π ALERTMANAGER HIGH AVAILABILITY
HA Setup
βββββββββββββββββββ βββββββββββββββββββ
β Prometheus 1 β β Prometheus 2 β
β (replica) β β (replica) β
ββββββββββ¬βββββββββ ββββββββββ¬βββββββββ
β β
β Same alerts β Same alerts
βΌ βΌ
βββββββββββββββββββ βββββββββββββββββββ
β Alertmanager 1 ββββββΊβ Alertmanager 2 β β Gossip protocol
β :9093 β β :9093 β (cluster peers)
ββββββββββ¬βββββββββ ββββββββββ¬βββββββββ
β β
βββββββββ¬ββββββββββββββββ
βΌ
Single notification (deduplicated!)
Configuration
# Alertmanager 1
alertmanager \
--cluster.listen-address=0.0.0.0:9094 \
--cluster.peer=alertmanager2:9094
# Alertmanager 2
alertmanager \
--cluster.listen-address=0.0.0.0:9094 \
--cluster.peer=alertmanager1:9094
Key Points
β Alertmanager instances communicate via gossip protocol (port 9094)
β They deduplicate alerts across replicas
β Both Prometheus replicas send to BOTH Alertmanager instances
β Only ONE notification is sent per unique alert
β Silences are also synchronized across the cluster
7.14 π REAL-WORLD ALERTING SCENARIOS
Scenario 1: Complete Alerting Pipeline
# prometheus.yml
rule_files:
- "alerts.yml"
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093']
# alerts.yml
groups:
- name: infrastructure
rules:
- alert: InstanceDown
expr: up == 0
for: 5m
labels:
severity: critical
annotations:
summary: "{{ $labels.instance }} is down"
- alert: HighCPU
expr: 100 - avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance) * 100 > 80
for: 10m
labels:
severity: warning
annotations:
summary: "High CPU on {{ $labels.instance }}"
# alertmanager.yml
route:
receiver: 'default'
group_by: ['alertname', 'instance']
group_wait: 30s
routes:
- match: {severity: critical}
receiver: 'pagerduty'
- match: {severity: warning}
receiver: 'slack'
receivers:
- name: 'default'
email_configs:
- to: 'ops@example.com'
- name: 'pagerduty'
pagerduty_configs:
- routing_key: 'xxx'
- name: 'slack'
slack_configs:
- channel: '#warnings'
7.15 π EXAM-STYLE QUESTIONS (20 Questions)
Question 1
What are the three states of a Prometheus alert?
A) Active, Inactive, Resolved B) Inactive, Pending, Firing C) Open, Closed, Acknowledged D) Warning, Critical, Resolved
β Answer & Explanation
Correct Answer: B
Prometheus alerts have three states: Inactive (condition is false, normal state), Pending (condition is true but the for duration hasnβt elapsed yet), and Firing (condition has been true for the entire for duration, alert is sent to Alertmanager). When a Firing alertβs condition becomes false, it transitions back to Inactive (resolved).
Question 2
What is the purpose of the for field in an alerting rule?
A) It specifies how long the alert notification should be displayed B) It defines how long the PromQL condition must be continuously true before the alert transitions from Pending to Firing C) It sets the interval between alert evaluations D) It defines the retention period for the alert
β Answer & Explanation
Correct Answer: B
The for field specifies the duration that the alert condition must be continuously true before the alert transitions from Pending to Firing. This prevents flapping alerts caused by brief spikes. For example, for: 5m means the condition must be true for 5 continuous minutes. If the condition becomes false before the for duration elapses, the alert returns to Inactive. If for is omitted, the alert fires immediately.
Question 3
What is the default port for Alertmanager?
A) 9090 B) 9093 C) 3000 D) 9100
β Answer & Explanation
Correct Answer: B
Alertmanager runs on port 9093 by default. Port 9090 is the Prometheus server, 3000 is Grafana, and 9100 is the Node Exporter.
Question 4
Is Alertmanager part of the Prometheus server?
A) Yes, it is built into the Prometheus server binary B) No, it is a separate component/binary that receives alerts from Prometheus C) Yes, but it must be explicitly enabled in prometheus.yml D) No, it is a Grafana plugin
β Answer & Explanation
Correct Answer: B
Alertmanager is a separate component with its own binary, configuration file (alertmanager.yml), and web UI (port 9093). It is NOT part of the Prometheus server. Prometheus evaluates alerting rules and sends firing alerts to Alertmanager via HTTP API. The connection is configured in the alerting section of prometheus.yml.
Question 5
What does the group_by configuration in Alertmanager do?
A) It groups Prometheus scrape targets together B) It groups alerts with the same label values into a single notification to reduce noise C) It groups recording rules for faster evaluation D) It groups Grafana dashboards
β Answer & Explanation
Correct Answer: B
group_by specifies which labels to use for grouping alerts. Alerts with the same values for the specified labels are combined into a single notification. For example, group_by: ['alertname', 'cluster'] groups all alerts with the same alertname and cluster into one notification. This prevents notification storms β instead of 100 separate βPodDownβ alerts, you get one β100 pods down in cluster Xβ notification.
Question 6
What is the difference between group_wait and group_interval?
A) They are the same thing
B) group_wait is the initial wait before the first notification for a new group; group_interval is the wait before sending updates when new alerts are added to an existing group
C) group_wait is for critical alerts; group_interval is for warnings
D) group_wait controls evaluation frequency; group_interval controls notification frequency
β Answer & Explanation
Correct Answer: B
group_wait (default 30s) is the initial delay before sending the first notification for a newly created alert group. This allows time for related alerts to accumulate. group_interval (default 5m) is the delay before sending subsequent notifications when new alerts are added to an already-notified group. repeat_interval (default 4h) controls how often to re-send notifications for alerts that are still firing.
Question 7
What does the repeat_interval in Alertmanager control?
A) How often Prometheus evaluates alerting rules B) How often Alertmanager re-sends notifications for alerts that are still firing C) How often Alertmanager checks for new silences D) How often the Alertmanager cluster syncs state
β Answer & Explanation
Correct Answer: B
repeat_interval controls how often Alertmanager re-sends notifications for alerts that remain in the Firing state. For example, with repeat_interval: 4h, if an alert has been firing for 8 hours, the on-call engineer receives notifications at t=0, t=4h, and t=8h. This ensures ongoing issues arenβt forgotten. The default is 4 hours.
Question 8
What is the purpose of inhibition rules in Alertmanager?
A) To permanently delete alerts from the system B) To suppress (mute) certain alerts when a related βparentβ alert is already firing, preventing notification cascades C) To delay alert notifications by a specified duration D) To route alerts to different receivers
β Answer & Explanation
Correct Answer: B
Inhibition rules suppress βchildβ alerts when a βparentβ alert is already firing. This prevents notification cascades during widespread failures. For example, if a βClusterDownβ alert fires, you can inhibit all βNodeDownβ and βPodDownβ alerts for that cluster, since theyβre all consequences of the same root cause. The source_match defines the parent alert, target_match defines the child alerts to suppress, and equal specifies which labels must match.
Question 9
What is the difference between silencing and inhibition in Alertmanager?
A) They are the same thing with different names B) Silencing is temporary and managed via UI/API (e.g., for maintenance); inhibition is permanent and configured in alertmanager.yml based on parent-child alert relationships C) Silencing is for critical alerts; inhibition is for warnings D) Silencing works in Prometheus; inhibition works in Grafana
β Answer & Explanation
Correct Answer: B
Silencing is a temporary, manually created mute for alerts matching specific matchers. It has a start and end time and is typically used during planned maintenance. Silences are created via the Alertmanager UI or API. Inhibition is a permanent, config-based rule that automatically suppresses child alerts when a parent alert is firing. Inhibition is defined in alertmanager.yml and is always active.
Question 10
Which template variable would you use in an Alertmanager annotation to access the value of the PromQL expression that triggered the alert?
A) {{ .Labels.value }}
B) {{ $value }}
C) {{ .CommonLabels.value }}
D) {{ .Alerts.Value }}
β Answer & Explanation
Correct Answer: B
In Prometheus alerting rule annotations, {{ $value }} gives you the value of the PromQL expression that triggered the alert. For example, if expr: cpu_usage > 90 evaluates to 95.5, then {{ $value }} in the annotation would be β95.5β. {{ $labels.xxx }} accesses label values. Note: In Alertmanager notification templates (not Prometheus annotations), you use .CommonLabels, .Alerts, etc.
Question 11
What happens when an alerting rule has no for field?
A) The alert never fires B) The alert is evaluated but stays in Pending state forever C) The alert fires immediately when the condition becomes true, skipping the Pending state D) Prometheus returns an error
β Answer & Explanation
Correct Answer: C
When the for field is omitted from an alerting rule, the alert transitions directly from Inactive to Firing as soon as the PromQL condition becomes true. It skips the Pending state entirely. This is useful for critical alerts that need immediate notification (e.g., service completely down). However, it can cause flapping alerts if the condition oscillates rapidly.
Question 12
Which section of prometheus.yml configures the connection to Alertmanager?
A) rule_files
B) scrape_configs
C) alerting
D) remote_write
β Answer & Explanation
Correct Answer: C
The alerting section in prometheus.yml configures how Prometheus connects to Alertmanager. It specifies the Alertmanager target addresses, scheme, timeout, and API version. Example:
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093']
rule_files loads alerting/recording rules, scrape_configs defines scrape targets, and remote_write sends data to external storage.
Question 13
What does the continue: true flag do in an Alertmanager routing rule?
A) It continues sending notifications even after the alert resolves B) It tells Alertmanager to continue checking sibling routes after a match, instead of stopping at the first match C) It continues evaluating the alerting rule even if Prometheus restarts D) It continues the silence after it expires
β Answer & Explanation
Correct Answer: B
By default, when an alert matches a route in the Alertmanager routing tree, it stops checking further sibling routes. Setting continue: true tells Alertmanager to continue checking subsequent sibling routes even after a match. This allows an alert to be sent to multiple receivers. For example, a critical alert could match both a PagerDuty route and a Slack route if continue: true is set on the first match.
Question 14
What is the resolve_timeout in Alertmanagerβs global configuration?
A) How long to wait before sending the first notification B) The time Alertmanager waits after the last alert notification before declaring the alert as resolved C) How long silences last by default D) The timeout for connecting to receivers
β Answer & Explanation
Correct Answer: B
resolve_timeout (default 5m) is the time Alertmanager waits after receiving the last alert notification from Prometheus before declaring the alert as resolved. If Prometheus stops sending an alert (because the condition became false), Alertmanager waits for resolve_timeout before sending a βresolvedβ notification. This prevents premature resolution notifications due to temporary network issues between Prometheus and Alertmanager.
Question 15
Which of the following is NOT a supported Alertmanager receiver type?
A) Slack B) PagerDuty C) Grafana Dashboard D) Webhook
β Answer & Explanation
Correct Answer: C
Grafana Dashboard is NOT a supported Alertmanager receiver type. Supported receivers include: Email, Slack, PagerDuty, OpsGenie, Webhook, VictorOps, Pushover, Telegram, Microsoft Teams, SNS, and others. While Grafana has its own alerting system that can integrate with Prometheus, βGrafana Dashboardβ is not a receiver type in Alertmanager.
Question 16
In an Alertmanager inhibition rule, what does the equal field specify?
A) The labels that must have the same value in both the source and target alerts for the inhibition to apply B) The labels that must be different between source and target C) The exact alert names that should be inhibited D) The severity levels that are considered equal
β Answer & Explanation
Correct Answer: A
The equal field in an inhibition rule specifies a list of labels that must have identical values in both the source (parent) and target (child) alerts for the inhibition to take effect. For example, equal: ['cluster'] means the inhibition only applies when the source and target alerts have the same cluster label value. This ensures that a βClusterDownβ alert for cluster-A only inhibits βNodeDownβ alerts for cluster-A, not cluster-B.
Question 17
What is the purpose of the send_resolved flag in an Alertmanager receiver?
A) It determines whether to send a notification when an alert transitions from Firing back to Inactive (resolved) B) It determines whether to resolve DNS names in the receiver URL C) It determines whether to resolve template variables in the notification D) It determines whether to automatically resolve silences
β Answer & Explanation
Correct Answer: A
send_resolved: true tells the receiver to send a notification when an alert is resolved (transitions from Firing back to Inactive). This lets the team know the issue has been fixed. For example, a Slack notification might say β[RESOLVED] InstanceDown - app1:8080 is back up.β By default, send_resolved is false for some receivers and true for others. Itβs recommended to enable it for critical alerts.
Question 18
Which PromQL expression is commonly used for a βdead manβs switchβ (Watchdog) alert?
A) up == 0
B) absent(up)
C) vector(1)
D) rate(errors_total[5m]) > 0
β Answer & Explanation
Correct Answer: C
vector(1) always evaluates to 1, meaning the alert is always firing. This is used as a βdead manβs switchβ or Watchdog alert. The alert is sent to a receiver that expects to receive it regularly (e.g., every 5 minutes). If the Watchdog alert stops arriving, it means the entire alerting pipeline is broken (Prometheus is down, Alertmanager is down, or the notification channel is broken). This ensures youβre alerted when the alerting system itself fails.
Question 19
How does Alertmanager deduplicate alerts in a high-availability setup?
A) By comparing alert names only B) By using a gossip protocol between Alertmanager instances to synchronize state and deduplicate based on alert fingerprints C) By relying on Prometheus to send alerts to only one Alertmanager instance D) By using a shared database
β Answer & Explanation
Correct Answer: B
In an HA setup, multiple Prometheus replicas send the same alerts to multiple Alertmanager instances. Alertmanager instances communicate via a gossip protocol (default port 9094) to synchronize their state. They deduplicate alerts based on alert fingerprints (hash of labels). This ensures that even though multiple copies of the same alert arrive, only one notification is sent to the receiver. Silences are also synchronized across the cluster.
Question 20
What is the correct order of the alerting pipeline?
A) Alertmanager evaluates rules β Prometheus sends notifications β Grafana displays alerts B) Prometheus evaluates alerting rules β Firing alerts sent to Alertmanager β Alertmanager groups, routes, and sends notifications to receivers C) Grafana evaluates rules β Sends to Prometheus β Prometheus sends to Slack D) Exporters generate alerts β Prometheus forwards to Grafana β Grafana notifies users
β Answer & Explanation
Correct Answer: B
The correct alerting pipeline is:
- Prometheus evaluates alerting rules (PromQL expressions) at
evaluation_interval - When an alert transitions to Firing (condition true for
forduration), Prometheus sends it to Alertmanager via HTTP API - Alertmanager deduplicates, groups, applies inhibition/silencing, routes based on labels, and sends notifications to receivers (Slack, Email, PagerDuty, etc.)
Grafana is not part of the core alerting pipeline (though it has its own alerting feature). Exporters donβt generate alerts β they only expose metrics.
β DOMAIN 7 REVISION SUMMARY CHEAT SHEET
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DOMAIN 7: ALERTING & ALERTMANAGER CHEAT SHEET β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β ALERTING PIPELINE: β
β Prometheus evaluates rules β Firing alerts β Alertmanager β Notify β
β β
β ALERT STATES: β
β Inactive β (expr true) β Pending β (for: X elapsed) β Firing π₯ β
β Firing β (expr false) β Resolved β Inactive β
β No "for" field β Skips Pending, fires immediately β
β β
β ALERTING RULE SYNTAX: β
β - alert: AlertName β
β expr: <PromQL> β
β for: <duration> (optional, prevents flapping) β
β labels: {severity: critical} β
β annotations: {summary: "...", description: "..."} β
β Templates: {{ $labels.xxx }}, {{ $value }}, {{ $externalLabels }} β
β β
β ALERTMANAGER: β
β Separate binary, port 9093, config: alertmanager.yml β
β Responsibilities: Dedup, Group, Inhibit, Silence, Route, Notify β
β β
β ROUTING: β
β route: β
β receiver: default β
β group_by: [alertname, cluster] β
β group_wait: 30s (initial wait for new group) β
β group_interval: 5m (wait for updates to existing group) β
β repeat_interval: 4h (re-notify if still firing) β
β routes: β
β - match: {severity: critical} β receiver: pagerduty β
β - match: {severity: warning} β receiver: slack β
β continue: true β Check sibling routes after match β
β match_re: β Regex matching β
β β
β RECEIVERS: β
β Email, Slack, PagerDuty, OpsGenie, Webhook, VictorOps, etc. β
β send_resolved: true β Notify when alert resolves β
β β
β INHIBITION: β
β Suppress child alerts when parent is firing β
β source_match (parent) β target_match (child) β
β equal: [labels that must match] β
β Example: ClusterDown inhibits NodeDown β
β β
β SILENCING: β
β Temporary mute via UI/API (start/end time) β
β Use for: Maintenance windows, known issues β
β Different from inhibition (permanent, config-based) β
β β
β HA SETUP: β
β Multiple Alertmanager instances with gossip protocol (port 9094) β
β Deduplicates alerts across replicas β
β Silences synchronized across cluster β
β β
β WATCHDOG: β
β expr: vector(1) β Always firing β
β If this alert stops β alerting pipeline is broken! β
β β
β CONFIG IN prometheus.yml: β
β rule_files: ["alerts.yml"] β
β alerting: β
β alertmanagers: β
β - static_configs: β
β - targets: ['alertmanager:9093'] β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
π CONGRATULATIONS! YOUβVE COMPLETED ALL 7 DOMAINS!
Complete Exam Coverage Summary
| Domain | Topic | Weight | Status |
|---|---|---|---|
| 1 | Observability Concepts | 18% | β Complete |
| 2 | Prometheus Fundamentals | 20% | β Complete |
| 3 | PromQL | 28% | β Complete |
| 4 | Instrumentation & Exporters | 16% | β Complete |
| 5 | Dashboarding & Visualization | 8% | β Complete |
| 6 | Service Discovery | 6% | β Complete |
| 7 | Alerting & Alertmanager | 4% | β Complete |
| TOTAL | 100% | π DONE! |
Final Exam Tips
1. PromQL is 28% β practice queries hands-on!
2. Know the 4 metric types cold (Counter, Gauge, Histogram, Summary)
3. Understand rate() vs irate() vs increase()
4. Know histogram_quantile() syntax perfectly
5. Remember: Histogram aggregatable β
, Summary NOT β
6. Pull model advantages (failure detection, debugging, centralized config)
7. SLI = measure, SLO = target, SLA = contract
8. Golden Signals vs RED vs USE β know when to use each
9. relabel_configs (before) vs metric_relabel_configs (after)
10. Alert states: Inactive β Pending β Firing
π The exam is OPEN BOOK (official Prometheus docs allowed)
But you won't have time to look up everything!
Aim to know 80% from memory.
β° 90 minutes, 60 questions = 1.5 min per question
Don't spend too long on any single question.
π― PASSING SCORE: 75% (45 out of 60 correct)
Best of luck with your PCA exam in December! Youβve got this! ππ