π PCA EXAM β DOMAIN 2: PROMETHEUS FUNDAMENTALS (20%)
Complete One-Stop Guide: Theory β Examples β Exam Questions
TABLE OF CONTENTS
2.1 Prometheus Architecture & Components
2.2 Prometheus Data Model
2.3 Metric Naming Conventions
2.4 The Four Metric Types (Counter, Gauge, Histogram, Summary)
2.5 Prometheus Configuration (prometheus.yml)
2.6 Storage & TSDB Internals
2.7 Exposition Format
2.8 Staleness & Timestamps
2.9 Federation
2.10 Remote Read / Remote Write
2.11 Real-World Scenarios
2.12 EXAM-STYLE QUESTIONS (30 Questions with Answers)
2.1 π PROMETHEUS ARCHITECTURE & COMPONENTS
The Big Picture Architecture Diagram
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PROMETHEUS ECOSYSTEM β
β β
β ββββββββββββ ββββββββββββ ββββββββββββ ββββββββββββ β
β β App 1 β β App 2 β β App 3 β β Batch β β
β β (Client β β (Client β β (Client β β Job β β
β β Library)β β Library)β β Library)β β (Short) β β
β ββββββ¬ββββββ ββββββ¬ββββββ ββββββ¬ββββββ ββββββ¬ββββββ β
β β β β β β
β β /metrics β /metrics β /metrics β push β
β βΌ βΌ βΌ βΌ β
β ββββββββββββββββββββββββββββββββββββββββ ββββββββββββββββ β
β β EXPORTERS β β Pushgateway β β
β β ββββββββββββ ββββββββββββ β β β β
β β β Node β β MySQL β ... β ββββββββ¬ββββββββ β
β β β Exporter β β Exporter β β β β
β β ββββββββββββ ββββββββββββ β β β
β ββββββββββββββββ¬ββββββββββββββββββββββββ β β
β β β β
β β βββββ PULL (scrape) βββββΊ β β
β βΌ βΌ β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β PROMETHEUS SERVER β β
β β β β
β β βββββββββββββββ βββββββββββββββ βββββββββββββββ β β
β β β Retrieval β β Storage β β HTTP β β β
β β β (Scrape β β (TSDB) β β Server β β β
β β β Manager) β β β β (API/UI) β β β
β β ββββββββ¬βββββββ ββββββββ¬βββββββ ββββββββ¬βββββββ β β
β β β β β β β
β β ββββββββ΄βββββββββββββββββββββββββββββββββββ΄βββββββ β β
β β β Service Discovery (SD) β β β
β β β Kubernetes β Consul β EC2 β DNS β File β Static β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β β
β β ββββββββββββββββββββββββββββββββββββββββββββββββββ β β
β β β Rule Manager (Alerting + Recording) β β β
β β ββββββββββββββββββββββ¬ββββββββββββββββββββββββββββ β β
β βββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββ β
β β β
β βΌ β
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββ β
β β Alertmanager β β Grafana β β Remote β β
β β (Routing, β β (Dashboards, β β Storage β β
β β Grouping, β β Visualization)β (Thanos, β β
β β Silencing) β β β β Cortex, β β
β ββββββββ¬ββββββββ ββββββββββββββββ β VictoriaM.) β β
β β ββββββββββββββββ β
β βΌ β
β ββββββββββββββββββββββββββββββββββββ β
β β Notification Receivers β β
β β Slack β Email β PagerDuty β WH β β
β ββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Component Breakdown
Component 1: Prometheus Server (The Core)
The Prometheus server is the central brain of the entire system. It has three main internal components:
A) Retrieval Module (Scrape Manager)
Responsibility: Pulls metrics from targets at configured intervals
How it works:
1. Reads target list from Service Discovery or static config
2. Sends HTTP GET request to each target's /metrics endpoint
3. Parses the response (exposition format)
4. Stores the data in TSDB
Configuration:
scrape_interval: 15s # How often to scrape
scrape_timeout: 10s # Max time to wait for response
Example scrape flow:
Prometheus β HTTP GET http://app:8080/metrics
Response:
http_requests_total{method="GET"} 1027
http_requests_total{method="POST"} 342
Prometheus β Parses β Stores in TSDB with timestamp
B) Storage Module (TSDB β Time Series Database)
Responsibility: Stores all scraped metrics efficiently on local disk
Key features:
β Optimized for time-series data (append-only)
β Data stored in blocks (typically 2 hours each)
β Uses compression (very efficient: ~1-2 bytes per sample)
β Local storage by default (not distributed!)
β Retention controlled by flags:
--storage.tsdb.retention.time=15d
--storage.tsdb.retention.size=50GB
Storage path: /prometheus/data/ (default)
C) HTTP Server (API & UI)
Responsibility: Serves the web UI and API endpoints
Key endpoints:
/graph β Expression browser (PromQL queries)
/targets β Shows all scrape targets and their status
/status β Runtime info, config, flags, rules
/alerts β Active alerting rules
/rules β All loaded rules
/metrics β Prometheus's OWN metrics (self-monitoring!)
/api/v1/* β REST API for programmatic access
Examples:
GET /api/v1/query?query=up
GET /api/v1/query_range?query=up&start=...&end=...&step=15s
GET /api/v1/targets
GET /api/v1/labels
Component 2: Service Discovery (SD)
Responsibility: Automatically finds targets to scrape
Why needed:
β In modern environments (Kubernetes, cloud), targets are dynamic
β Pods are created/destroyed constantly
β Manually maintaining a target list is impossible at scale
Supported SD mechanisms:
β static_configs (manual list)
β file_sd_configs (read from JSON/YAML files)
β kubernetes_sd_configs (K8s API: pods, services, nodes, endpoints)
β ec2_sd_configs (AWS EC2 instances)
β consul_sd_configs (HashiCorp Consul)
β dns_sd_configs (DNS records)
β azure_sd_configs (Azure VMs)
β gce_sd_configs (Google Cloud instances)
β dockerswarm_sd_configs
β ... and many more
Example (Kubernetes SD):
kubernetes_sd_configs:
- role: pod
namespaces:
names: ['production']
# Prometheus automatically discovers all pods in 'production' namespace
Component 3: Client Libraries
Responsibility: Instrument application code to expose custom metrics
Available for:
β Go (github.com/prometheus/client_golang) β Most mature
β Python (prometheus_client)
β Java (io.prometheus.simpleclient)
β Ruby (prometheus-client)
β .NET (prometheus-net)
β Rust, C++, etc. (community maintained)
What they do:
1. Provide APIs to create Counters, Gauges, Histograms, Summaries
2. Automatically expose a /metrics HTTP endpoint
3. Handle metric registration and exposition format
Example (Go):
var httpRequests = prometheus.NewCounterVec(
prometheus.CounterOpts{
Name: "http_requests_total",
Help: "Total HTTP requests",
},
[]string{"method", "status"},
)
func handler(w http.ResponseWriter, r *http.Request) {
httpRequests.WithLabelValues(r.Method, "200").Inc()
// ... handle request
}
Example (Python):
from prometheus_client import Counter, start_http_server
REQUESTS = Counter('http_requests_total', 'Total HTTP requests', ['method'])
def handle_request(method):
REQUESTS.labels(method=method).inc()
start_http_server(8000) # Exposes /metrics on port 8000
Component 4: Exporters
Responsibility: Expose metrics from third-party systems that don't
natively support Prometheus format
How they work:
β An exporter is a standalone process
β It connects to a third-party system (MySQL, Redis, Linux, etc.)
β It translates that system's metrics into Prometheus format
β It exposes a /metrics endpoint for Prometheus to scrape
Common Exporters:
βββββββββββββββββββββββ¬βββββββββ¬βββββββββββββββββββββββββββ
β Exporter β Port β What it monitors β
βββββββββββββββββββββββΌβββββββββΌβββββββββββββββββββββββββββ€
β Node Exporter β 9100 β Linux system metrics β
β MySQL Exporter β 9104 β MySQL/MariaDB β
β Redis Exporter β 9121 β Redis β
β PostgreSQL Exporter β 9187 β PostgreSQL β
β Blackbox Exporter β 9115 β HTTP/TCP/ICMP probing β
β cAdvisor β 8080 β Docker containers β
β SNMP Exporter β 9116 β Network devices (SNMP) β
β JMX Exporter β varies β Java JMX metrics β
β Kafka Exporter β 9308 β Apache Kafka β
βββββββββββββββββββββββ΄βββββββββ΄βββββββββββββββββββββββββββ
Example (Node Exporter):
$ ./node_exporter
# Now available at http://localhost:9100/metrics
# Exposes: node_cpu_seconds_total, node_memory_MemTotal_bytes,
# node_disk_io_time_seconds_total, etc.
Component 5: Pushgateway
Responsibility: Accept metrics pushed by short-lived batch jobs
When to use:
β Cron jobs that run for 5 seconds (Prometheus can't scrape in time)
β CI/CD pipeline jobs
β Serverless functions
β Any job shorter than the scrape interval
When NOT to use:
β Long-running services (use direct scraping!)
β As a general replacement for pull model
β To push metrics from behind a firewall (use federation instead)
Flow:
Batch Job β HTTP POST β Pushgateway β Prometheus scrapes Pushgateway
Example:
# In a bash cron job:
echo "backup_duration_seconds 45.2" | curl --data-binary @- \
http://pushgateway:9091/metrics/job/backup/instance/server1
# Prometheus config:
scrape_configs:
- job_name: 'pushgateway'
honor_labels: true # Important! Preserves job/instance labels
static_configs:
- targets: ['pushgateway:9091']
Component 6: Alertmanager
Responsibility: Handles alerts sent by the Prometheus server
What it does:
1. Receives alerts from Prometheus (via HTTP API)
2. Deduplicates alerts (same alert from multiple sources)
3. Groups related alerts together
4. Routes alerts to the correct receiver
5. Handles silencing (suppressing alerts during maintenance)
6. Handles inhibition (suppressing alerts when a parent alert fires)
Flow:
Prometheus Rule Engine β Alert fires β Sends to Alertmanager
Alertmanager β Groups β Routes β Sends to Slack/Email/PagerDuty
Key concepts:
β Grouping: Group by cluster, alertname
β Inhibition: If "cluster down" fires, inhibit "node down" alerts
β Silencing: Mute alerts during planned maintenance
β Receivers: Where to send (email, Slack, webhook, PagerDuty)
Note: Alertmanager is a SEPARATE binary, not part of Prometheus server!
Component 7: Grafana (Visualization)
Responsibility: Create dashboards and visualize Prometheus metrics
Key features:
β Add Prometheus as a data source
β Write PromQL queries in panels
β Create beautiful dashboards
β Set up Grafana-managed alerts (alternative to Alertmanager)
β Import community dashboards (grafana.com/dashboards)
Note: Grafana is NOT part of Prometheus. It's a separate open-source tool
that integrates with Prometheus as a data source.
π‘ Key Takeaway for Exam
Prometheus Server = Retrieval + TSDB + HTTP Server Service Discovery = Finds targets automatically Client Libraries = Instrument your code Exporters = Bridge to third-party systems Pushgateway = ONLY for short-lived batch jobs Alertmanager = Separate component for alert routing Grafana = Separate tool for visualization
2.2 π PROMETHEUS DATA MODEL
The Fundamental Concept: Time Series
Prometheus stores all data as time series. A time series is a stream of timestamped values belonging to the same metric and set of labels.
Data Model Structure
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β β
β Time Series = Metric Name + Labels + Samples β
β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Metric Name: http_requests_total β β
β β Labels: {method="GET", status="200"} β β
β β β β
β β Samples (timestamp + value): β β
β β @1699000000 β 1027 β β
β β @1699000015 β 1042 β β
β β @1699000030 β 1058 β β
β β @1699000045 β 1071 β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
β A DIFFERENT combination of labels = DIFFERENT series! β
β β
β http_requests_total{method="GET", status="200"} β S1 β
β http_requests_total{method="GET", status="404"} β S2 β
β http_requests_total{method="POST", status="200"} β S3 β
β http_requests_total{method="POST", status="500"} β S4 β
β β
β These are 4 SEPARATE time series! β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The Three Components in Detail
1. Metric Name
Definition: Identifies the general feature of a system being measured
Rules:
β Must match regex: [a-zA-Z_:][a-zA-Z0-9_:]*
β Must start with a letter, underscore, or colon
β Can contain letters, digits, underscores, colons
β Colons are reserved for recording rules (don't use in raw metrics!)
β Convention: snake_case
Examples:
β
http_requests_total
β
node_cpu_seconds_total
β
process_resident_memory_bytes
β
go_goroutines
β http-requests-total (hyphens not allowed)
β 1http_requests (can't start with digit)
β http.requests.total (dots not allowed)
2. Labels (Key-Value Pairs)
Definition: Labels add dimensions to a metric, allowing you to
differentiate between time series
Rules:
β Label names must match: [a-zA-Z_][a-zA-Z0-9_]*
β Label names starting with __ are reserved for internal use
β Label values can be any Unicode string
β Labels are what make Prometheus powerful (high cardinality!)
Special Labels (automatically added by Prometheus):
β job: The job name from scrape config (e.g., "web-app")
β instance: The target that was scraped (e.g., "app1:8080")
Examples:
http_requests_total{job="api", instance="app1:8080", method="GET", status="200"}
node_cpu_seconds_total{job="nodes", instance="server1:9100", cpu="0", mode="idle"}
β οΈ Cardinality Warning:
High cardinality = too many unique label combinations = memory explosion!
BAD: http_requests_total{user_id="user_12345"}
β Millions of unique user IDs = millions of time series!
GOOD: http_requests_total{method="GET", status="200"}
β Limited combinations = manageable number of series
3. Samples (Timestamp + Value)
Definition: Each data point in a time series is a sample consisting of:
β A float64 value
β A millisecond-precision timestamp
Characteristics:
β Values are always 64-bit floating point numbers
β Timestamps are Unix epoch in milliseconds
β Samples are appended in chronological order
β Prometheus stores samples efficiently (~1-2 bytes per sample)
Example:
Metric: http_requests_total{method="GET"}
Sample 1: (timestamp=1699000000000, value=1027.0)
Sample 2: (timestamp=1699000015000, value=1042.0)
Sample 3: (timestamp=1699000030000, value=1058.0)
Note: Even though the counter is an integer, Prometheus stores it as float64
Notation Format
Full notation:
<metric_name>{<label_name>=<label_value>, ...}
Examples:
http_requests_total{method="GET", status="200"}
node_memory_MemAvailable_bytes{instance="server1:9100"}
up{job="prometheus", instance="localhost:9090"}
In PromQL:
# Instant vector (single point in time)
http_requests_total{method="GET"}
# Range vector (over a time window)
http_requests_total{method="GET"}[5m]
π‘ Key Takeaway for Exam
Time Series = Metric Name + Labels Different label values = different time series Labels
jobandinstanceare added automatically Labels starting with__are internal/reserved High cardinality labels (like user_id) = BAD (memory explosion) Values are always float64, timestamps are millisecond precision
2.3 π METRIC NAMING CONVENTIONS
Prometheus has strict naming conventions that you MUST know for the exam.
General Rules
1. Use snake_case (underscores, not hyphens or camelCase)
β
http_requests_total
β httpRequestsTotal
β http-requests-total
2. Use base units (seconds, bytes, not milliseconds, megabytes)
β
http_request_duration_seconds
β http_request_duration_milliseconds
β
process_resident_memory_bytes
β process_resident_memory_megabytes
3. Include the unit in the metric name as a suffix
β
node_cpu_seconds_total (unit: seconds)
β
node_memory_MemTotal_bytes (unit: bytes)
β
http_request_duration_seconds (unit: seconds)
4. Use _total suffix for counters
β
http_requests_total
β
node_cpu_seconds_total
β http_requests (missing _total for a counter)
5. Use _ratio or _fraction for ratios (0 to 1)
β
go_memstats_alloc_ratio
6. Don't put the metric type in the name
β http_requests_counter_total (redundant!)
β
http_requests_total
7. Use _bucket, _sum, _count suffixes for histograms/summaries
(These are auto-generated, don't create them manually)
Common Metric Name Patterns
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Pattern β Example β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β <namespace>_<subsystem>_<name> β http_requests_total β
β β node_cpu_seconds_total β
β β process_open_fds β
β β go_goroutines β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Namespaces: β β
β http, node, process, go, β β
β jvm, mysql, redis, kafka β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Units: β β
β _seconds, _bytes, _total, β β
β _ratio, _info, _created β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
π‘ Key Takeaway for Exam
snake_case, base units (seconds, bytes), _total for counters Include unit in name, donβt include type in name Colons (:) reserved for recording rules only
2.4 π THE FOUR METRIC TYPES
This is one of the most heavily tested topics in the PCA exam!
Type 1: COUNTER π’
Definition: A cumulative metric that only goes up (or resets to zero on restart). It represents a count of events.
Key Characteristics:
- Monotonically increasing (never decreases)
- Resets to 0 when the process restarts
- Used with
rate(),increase(),irate()in PromQL - Always has
_totalsuffix (by convention)
When to use:
- Total HTTP requests served
- Total errors occurred
- Total bytes sent/received
- Total tasks completed
- Total CPU seconds consumed
Examples:
# Total HTTP requests (always increasing)
http_requests_total{method="GET", status="200"} 15432
http_requests_total{method="GET", status="404"} 234
http_requests_total{method="POST", status="500"} 12
# CPU time consumed (always increasing)
node_cpu_seconds_total{cpu="0", mode="user"} 78345.23
node_cpu_seconds_total{cpu="0", mode="idle"} 234567.89
# Bytes transmitted (always increasing)
node_network_transmit_bytes_total{device="eth0"} 9876543210
PromQL Usage:
# NEVER use raw counter values! Always use rate() or increase()
# Per-second rate of requests (last 5 minutes)
rate(http_requests_total[5m])
# Total increase in requests over last hour
increase(http_requests_total[1h])
# Instant rate (last 2 data points) β for volatile graphs
irate(http_requests_total[5m])
# Total requests per second across all instances
sum(rate(http_requests_total[5m]))
Counter Reset Handling:
Timeline:
t=0: counter = 100
t=1: counter = 150
t=2: counter = 200
t=3: counter = 0 β PROCESS RESTARTED!
t=4: counter = 30
t=5: counter = 75
rate() and increase() AUTOMATICALLY handle this reset!
They detect the drop from 200 β 0 and adjust the calculation.
rate() over [t=0 to t=5]:
= (75 + 200) / 5 seconds = 55 per second
(It adds the pre-reset value to the post-reset value)
Type 2: GAUGE π‘οΈ
Definition: A metric that represents a single numerical value that can go up AND down. Itβs a snapshot of the current state.
Key Characteristics:
- Can increase, decrease, or stay the same
- Represents a point-in-time value
- Can be used directly in PromQL (no need for
rate()) - No
_totalsuffix
When to use:
- Current temperature
- Current memory usage
- Current number of active connections
- Current queue size
- Current CPU utilization percentage
- Number of goroutines currently running
Examples:
# Current memory usage (goes up and down)
process_resident_memory_bytes 256000000
node_memory_MemAvailable_bytes 4294967296
# Current temperature
node_hwmon_temp_celsius{chip="coretemp", sensor="temp1"} 65.0
# Current active connections
mysql_global_status_threads_connected 45
# Current queue depth
rabbitmq_queue_messages 1200
# Number of goroutines
go_goroutines 150
# Up/down status (1 = up, 0 = down)
up{job="web-app", instance="app1:8080"} 1
PromQL Usage:
# Use gauges directly (no rate needed!)
process_resident_memory_bytes
# Average memory across all instances
avg(process_resident_memory_bytes)
# Instances with memory > 1GB
process_resident_memory_bytes > 1073741824
# Current CPU idle percentage
avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100
# Note: node_cpu_seconds_total is a COUNTER, so we use rate()
# But the RESULT is a gauge-like value (percentage)
Type 3: HISTOGRAM π
Definition: A histogram samples observations (usually request durations or response sizes) and counts them in configurable buckets. It also provides a sum and count of all observations.
Key Characteristics:
- Divides observations into buckets (ranges)
- Each bucket is a cumulative counter (le = βless than or equalβ)
- Automatically creates
_bucket,_sum, and_countmetrics - Allows server-side quantile calculation using
histogram_quantile() - Aggregatable across instances (unlike Summary!)
Auto-generated Metrics:
# If you create a histogram called http_request_duration_seconds
# with buckets [0.1, 0.5, 1.0, 2.5, 5.0, 10.0]
# You get these metrics automatically:
# 1. Bucket counters (cumulative!)
http_request_duration_seconds_bucket{le="0.1"} 500 # β€ 0.1s
http_request_duration_seconds_bucket{le="0.5"} 800 # β€ 0.5s
http_request_duration_seconds_bucket{le="1.0"} 950 # β€ 1.0s
http_request_duration_seconds_bucket{le="2.5"} 990 # β€ 2.5s
http_request_duration_seconds_bucket{le="5.0"} 998 # β€ 5.0s
http_request_duration_seconds_bucket{le="10.0"} 999 # β€ 10.0s
http_request_duration_seconds_bucket{le="+Inf"} 1000 # ALL requests
# 2. Sum of all observed values
http_request_duration_seconds_sum 245.7
# 3. Count of all observations
http_request_duration_seconds_count 1000
Understanding Cumulative Buckets:
le="0.1" β 500 requests took β€ 0.1 seconds
le="0.5" β 800 requests took β€ 0.5 seconds (INCLUDES the 500 above!)
le="1.0" β 950 requests took β€ 1.0 seconds (INCLUDES the 800 above!)
le="+Inf" β 1000 total requests (MUST always equal _count)
So the actual distribution is:
0.0 - 0.1s: 500 requests
0.1 - 0.5s: 300 requests (800 - 500)
0.5 - 1.0s: 150 requests (950 - 800)
1.0 - 2.5s: 40 requests (990 - 950)
2.5 - 5.0s: 8 requests (998 - 990)
5.0 - 10.0s: 1 request (999 - 998)
> 10.0s: 1 request (1000 - 999)
PromQL Usage:
# Calculate the 99th percentile latency
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# Calculate the 95th percentile latency
histogram_quantile(0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# Calculate the 50th percentile (median)
histogram_quantile(0.5,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# Average request duration
rate(http_request_duration_seconds_sum[5m])
/
rate(http_request_duration_seconds_count[5m])
# Percentage of requests under 200ms
sum(rate(http_request_duration_seconds_bucket{le="0.2"}[5m]))
/
sum(rate(http_request_duration_seconds_bucket{le="+Inf"}[5m]))
When to use Histogram:
- Request durations (latency)
- Response sizes
- Processing times
- When you need to calculate quantiles across multiple instances
- When you want to aggregate data from multiple servers
Type 4: SUMMARY π
Definition: A summary also samples observations but calculates quantiles on the client side (in the application). It provides pre-calculated quantile values.
Key Characteristics:
- Calculates quantiles (e.g., p50, p90, p99) in the application
- Also provides
_sumand_countmetrics - Quantiles are NOT aggregatable across instances!
- More expensive on the client side (calculation overhead)
- Quantile values are exact (not approximations like histogram)
Auto-generated Metrics:
# If you create a summary called http_request_duration_seconds
# with quantiles [0.5, 0.9, 0.99]
# You get these metrics automatically:
# 1. Pre-calculated quantiles
http_request_duration_seconds{quantile="0.5"} 0.12 # Median: 120ms
http_request_duration_seconds{quantile="0.9"} 0.45 # p90: 450ms
http_request_duration_seconds{quantile="0.99"} 1.23 # p99: 1.23s
# 2. Sum of all observed values
http_request_duration_seconds_sum 245.7
# 3. Count of all observations
http_request_duration_seconds_count 1000
PromQL Usage:
# Read pre-calculated quantile directly
http_request_duration_seconds{quantile="0.99"}
# Average duration (same as histogram)
rate(http_request_duration_seconds_sum[5m])
/
rate(http_request_duration_seconds_count[5m])
# β οΈ You CANNOT aggregate quantiles across instances!
# This is WRONG:
# sum(http_request_duration_seconds{quantile="0.99"})
# β Averaging percentiles is mathematically invalid!
HISTOGRAM vs SUMMARY β The Big Comparison (Exam Critical!)
| Aspect | Histogram | Summary |
|---|---|---|
| Quantile calculation | Server-side (PromQL) | Client-side (application) |
| Aggregatable? | β YES (across instances) | β NO (mathematically invalid) |
| Bucket/Quantile config | Buckets configured in code | Quantiles configured in code |
| Accuracy | Approximation (depends on buckets) | Exact (for configured quantiles) |
| Client overhead | Low (just counting) | Higher (calculating quantiles) |
| Flexibility | Can calculate ANY quantile in PromQL | Only pre-configured quantiles |
| Use case | Multi-instance, distributed systems | Single instance, exact quantiles |
| Auto metrics | _bucket{le="..."}, _sum, _count |
{quantile="..."}, _sum, _count |
π Real-World Example: Choosing Between Histogram and Summary
Scenario: You have 10 instances of your web service behind a load balancer.
You want to know the p99 latency across ALL instances.
Using HISTOGRAM β
:
β Each instance exposes bucket counts
β In PromQL: sum(rate(..._bucket[5m])) by (le)
β Then: histogram_quantile(0.99, ...)
β Result: Accurate p99 across all 10 instances!
Using SUMMARY β:
β Each instance exposes its own p99
β Instance 1: p99 = 200ms
β Instance 2: p99 = 500ms
β Instance 3: p99 = 100ms
β What's the overall p99? You CAN'T just average them!
β avg(p99) = 267ms β This is WRONG!
β The actual p99 could be 800ms (if Instance 2 handles most traffic)
Conclusion: Use HISTOGRAM for distributed systems!
π‘ Key Takeaway for Exam
Counter = Only goes up (use
rate(),increase()) Gauge = Goes up and down (use directly) Histogram = Buckets, server-side quantiles, AGGREGATABLE β Summary = Pre-calculated quantiles, NOT aggregatable βFor distributed systems β Always prefer Histogram over Summary Histogram buckets are CUMULATIVE (
le= less than or equal)_sumand_countare available for both Histogram and Summary
2.5 π PROMETHEUS CONFIGURATION (prometheus.yml)
Complete Configuration Structure
# prometheus.yml β Complete annotated example
# ββββββββββββββββββββββββββββββββββββββββββββββ
# GLOBAL CONFIGURATION
# ββββββββββββββββββββββββββββββββββββββββββββββ
global:
scrape_interval: 15s # How often to scrape targets (default: 1m)
evaluation_interval: 15s # How often to evaluate rules (default: 1m)
scrape_timeout: 10s # Timeout for each scrape (default: 10s)
# Labels added to all time series and alerts
external_labels:
cluster: 'production'
region: 'us-east-1'
environment: 'prod'
# ββββββββββββββββββββββββββββββββββββββββββββββ
# RULE FILES
# ββββββββββββββββββββββββββββββββββββββββββββββ
rule_files:
- "alerting_rules.yml" # Alerting rules
- "recording_rules.yml" # Recording rules
- "/etc/prometheus/rules/*.yml" # Glob patterns supported
# ββββββββββββββββββββββββββββββββββββββββββββββ
# ALERTING (Alertmanager configuration)
# ββββββββββββββββββββββββββββββββββββββββββββββ
alerting:
alertmanagers:
- static_configs:
- targets:
- 'alertmanager1:9093'
- 'alertmanager2:9093'
scheme: http
path_prefix: /
timeout: 10s
# ββββββββββββββββββββββββββββββββββββββββββββββ
# SCRAPE CONFIGURATIONS
# ββββββββββββββββββββββββββββββββββββββββββββββ
scrape_configs:
# Job 1: Prometheus self-monitoring
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
# Job 2: Application servers (static)
- job_name: 'web-app'
scrape_interval: 10s # Override global for this job
scrape_timeout: 5s # Override global for this job
metrics_path: '/metrics' # Default: /metrics
scheme: http # Default: http
static_configs:
- targets:
- 'app1.example.com:8080'
- 'app2.example.com:8080'
labels:
team: 'backend'
env: 'production'
# Job 3: Node Exporter (file-based SD)
- job_name: 'node-exporter'
file_sd_configs:
- files:
- '/etc/prometheus/targets/nodes.json'
refresh_interval: 30s
# Job 4: Kubernetes pods
- job_name: 'kubernetes-pods'
kubernetes_sd_configs:
- role: pod
namespaces:
names: ['default', 'monitoring']
relabel_configs:
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_port]
action: replace
target_label: __address__
regex: (.+)
replacement: ${1}:${2}
# Job 5: Blackbox Exporter (probing)
- job_name: 'blackbox'
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets:
- https://example.com
- https://prometheus.io
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: blackbox-exporter:9115
# Job 6: Pushgateway
- job_name: 'pushgateway'
honor_labels: true # Important for Pushgateway!
static_configs:
- targets: ['pushgateway:9091']
Key Configuration Concepts
1. Job vs Instance vs Target
Job: A logical group of targets (e.g., "web-app", "database")
β Defined by job_name in scrape_configs
β Added as the "job" label to all metrics
Instance: A specific endpoint being scraped (e.g., "app1:8080")
β Added as the "instance" label automatically
β Usually host:port
Target: The actual URL being scraped
β e.g., http://app1:8080/metrics
Example:
job="web-app", instance="app1:8080" β target: http://app1:8080/metrics
job="web-app", instance="app2:8080" β target: http://app2:8080/metrics
job="database", instance="db1:9104" β target: http://db1:9104/metrics
2. honor_labels vs honor_timestamps
honor_labels: true
β Keeps the labels from the scraped target as-is
β Prometheus won't override "job" and "instance" labels
β ESSENTIAL for Pushgateway and federation
β Default: false
honor_timestamps: true
β Uses the timestamps from the scraped target
β Default: true
β Set to false if you want Prometheus to use its own scrape time
3. metrics_path and scheme
metrics_path: '/metrics' # Default path to scrape
scheme: 'http' # Default scheme (http or https)
# Custom example:
metrics_path: '/actuator/prometheus' # Spring Boot apps
scheme: 'https'
4. Relabeling (Preview β covered more in Domain 6)
relabel_configs: # Applied BEFORE scraping (on targets)
- source_labels: [__address__]
target_label: instance
metric_relabel_configs: # Applied AFTER scraping (on metrics)
- source_labels: [__name__]
regex: 'go_.*'
action: drop # Drop all Go runtime metrics
π‘ Key Takeaway for Exam
global section sets defaults for all scrape configs scrape_interval default is 1m, evaluation_interval default is 1m job_name becomes the
joblabel, target address becomesinstancelabel honor_labels: true is critical for Pushgateway rule_files loads alerting and recording rules alerting section configures Alertmanager connection Each job can override global settings
2.6 π STORAGE & TSDB INTERNALS
Prometheus Local Storage (TSDB)
Prometheus uses its own Time Series Database (TSDB) for local storage.
Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PROMETHEUS TSDB β
β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β HEAD BLOCK (In-Memory) β β
β β β Receives all new incoming samples β β
β β β Keeps last ~2 hours of data in memory β β
β β β Uses WAL (Write-Ahead Log) for crash recovery β β
β β β When full β Compacted to disk as a block β β
β ββββββββββββββββββββββββ¬βββββββββββββββββββββββββββ β
β β compaction (every 2 hours) β
β βΌ β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β PERSISTED BLOCKS (On Disk) β β
β β β β
β β Block 1: [0h - 2h] (compressed, immutable) β β
β β Block 2: [2h - 4h] (compressed, immutable) β β
β β Block 3: [4h - 6h] (compressed, immutable) β β
β β ... β β
β β Block N: [Xh - Yh] (compressed, immutable) β β
β β β β
β β Larger blocks created by merging smaller ones: β β
β β [0h-2h] + [2h-4h] + [4h-6h] β [0h-6h] β β
β βββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β
β Storage path: /prometheus/data/ (default) β
β Each block is a directory with: β
β β meta.json (block metadata) β
β β index (series index) β
β β chunks/ (compressed sample data) β
β β tombstones (deleted series) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Key Storage Concepts
1. Write-Ahead Log (WAL)
Purpose: Prevents data loss if Prometheus crashes
How it works:
1. Every incoming sample is first written to the WAL (on disk)
2. Then it's stored in the Head block (in memory)
3. If Prometheus crashes, it replays the WAL on restart
4. WAL is truncated after the Head block is compacted to disk
Location: /prometheus/data/wal/
2. Compaction
Purpose: Merge small blocks into larger ones for efficiency
Process:
β Head block (2h in memory) β Compacted to 2h disk block
β Three 2h blocks β Merged into one 6h block
β Three 6h blocks β Merged into one 18h block
β And so on...
Benefits:
β Reduces the number of files
β Improves query performance
β Better compression ratios
3. Retention
Two ways to control how long data is kept:
By Time (default):
--storage.tsdb.retention.time=15d # Keep 15 days (default)
--storage.tsdb.retention.time=6h # Keep 6 hours
--storage.tsdb.retention.time=1y # Keep 1 year
By Size:
--storage.tsdb.retention.size=50GB # Keep max 50GB of data
--storage.tsdb.retention.size=100GB # Keep max 100GB
Both can be combined:
--storage.tsdb.retention.time=30d
--storage.tsdb.retention.size=100GB
β Data is deleted when EITHER limit is reached
β οΈ Important: Retention applies to BLOCKS, not individual samples.
A block is only deleted when its ENTIRE time range is outside
the retention window.
4. Storage Efficiency
Prometheus TSDB is very efficient:
β ~1-2 bytes per sample (after compression)
β A single Prometheus server can handle:
β’ Millions of active time series
β’ Hundreds of thousands of samples per second
β Typical storage: ~10GB per day for 1M active series at 15s scrape
Example calculation:
1,000,000 series Γ 4 scrapes/minute Γ 60 min Γ 24 hours
= 5,760,000,000 samples/day
Γ 1.5 bytes/sample
β 8.6 GB/day
Important Storage Flags
# Data directory
--storage.tsdb.path=/prometheus/data
# Retention
--storage.tsdb.retention.time=15d
--storage.tsdb.retention.size=50GB
# Minimum block duration
--storage.tsdb.min-block-duration=2h # Default: 2h
# Maximum block duration
--storage.tsdb.max-block-duration=36h # Default: 36h
# Disable compaction (not recommended!)
--storage.tsdb.no-lockfile
# WAL compression
--storage.tsdb.wal-compression # Enable WAL compression
π‘ Key Takeaway for Exam
TSDB = Prometheusβs built-in local time series database Head block = Last ~2 hours in memory, uses WAL for crash safety Persisted blocks = Compressed, immutable blocks on disk Compaction = Merges small blocks into larger ones Retention = Controlled by time (
--storage.tsdb.retention.time) or size Default retention = 15 days Storage efficiency = ~1-2 bytes per sample
2.7 π EXPOSITION FORMAT
What is the Exposition Format?
The exposition format is the text-based format that targets use to expose metrics on the /metrics endpoint. Prometheus scrapes this format.
Format Structure
# Each metric has:
# 1. HELP line (description) β optional but recommended
# 2. TYPE line (metric type) β required
# 3. Data lines (actual metric values)
# βββ COUNTER Example βββ
# HELP http_requests_total The total number of HTTP requests.
# TYPE http_requests_total counter
http_requests_total{method="post",code="200"} 1027
http_requests_total{method="post",code="400"} 3
http_requests_total{method="get",code="200"} 15432
# βββ GAUGE Example βββ
# HELP node_memory_MemAvailable_bytes Available memory in bytes.
# TYPE node_memory_MemAvailable_bytes gauge
node_memory_MemAvailable_bytes 4294967296
# βββ HISTOGRAM Example βββ
# HELP http_request_duration_seconds Request duration in seconds.
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{le="0.05"} 24054
http_request_duration_seconds_bucket{le="0.1"} 33444
http_request_duration_seconds_bucket{le="0.2"} 100392
http_request_duration_seconds_bucket{le="0.5"} 129389
http_request_duration_seconds_bucket{le="1"} 133988
http_request_duration_seconds_bucket{le="+Inf"} 144320
http_request_duration_seconds_sum 53423
http_request_duration_seconds_count 144320
# βββ SUMMARY Example βββ
# HELP rpc_duration_seconds RPC duration in seconds.
# TYPE rpc_duration_seconds summary
rpc_duration_seconds{quantile="0.01"} 3102
rpc_duration_seconds{quantile="0.05"} 3272
rpc_duration_seconds{quantile="0.5"} 4773
rpc_duration_seconds{quantile="0.9"} 9001
rpc_duration_seconds{quantile="0.99"} 76656
rpc_duration_seconds_sum 1.7560473e+07
rpc_duration_seconds_count 2693
# βββ Special: "up" metric βββ
# Automatically generated by Prometheus for each scrape
# up{job="web", instance="app1:8080"} 1 β Target is healthy
# up{job="web", instance="app2:8080"} 0 β Target is DOWN!
# βββ Special: "scrape_duration_seconds" βββ
# How long the scrape took
# scrape_duration_seconds{job="web"} 0.023
# βββ Special: "scrape_samples_scraped" βββ
# How many samples were scraped
# scrape_samples_scraped{job="web"} 342
Format Rules
1. Lines starting with # are comments (except HELP and TYPE)
2. # HELP <metric_name> <description>
3. # TYPE <metric_name> <type> (counter|gauge|histogram|summary|untyped)
4. Metric lines: <metric_name>{<labels>} <value> [<timestamp>]
5. Labels are comma-separated key="value" pairs inside {}
6. Label values must be in double quotes
7. Timestamp is optional (Unix epoch in milliseconds)
8. Empty lines are allowed
9. Encoding must be UTF-8
10. Content-Type header should be: text/plain; version=0.0.4
(or application/openmetrics-text for OpenMetrics)
π‘ Key Takeaway for Exam
Exposition format is text-based, exposed on
/metricsEach metric should have# HELPand# TYPElines Types: counter, gauge, histogram, summary, untyped Special auto-generated metrics:up,scrape_duration_seconds,scrape_samples_scrapedup = 1means target is healthy,up = 0means target is DOWN
2.8 π STALENESS & TIMESTAMPS
Staleness
Definition: Staleness is Prometheusβs mechanism for handling time series that stop receiving new samples (e.g., when a target goes down or a label combination disappears).
How it works:
1. Prometheus scrapes a target every 15 seconds
2. Each scrape produces a set of time series
3. If a time series is NOT present in the latest scrape,
it is marked as "stale" after 5 minutes
4. Stale series return "StaleNaN" in queries
5. This prevents showing outdated data as if it were current
Example:
t=0: http_requests_total{status="200"} = 100 β
scraped
t=15s: http_requests_total{status="200"} = 115 β
scraped
t=30s: http_requests_total{status="200"} = 130 β
scraped
t=45s: (target goes down, scrape fails)
t=60s: (still down)
...
t=5m: Series marked as STALE β
Queries will no longer return this series
Why 5 minutes?
β Default stale timeout = 5 minutes (lookback delta)
β Configurable via --query.lookback-delta flag
β Should be at least 2Γ the scrape interval
The up Metric and Staleness
The "up" metric is the most important staleness indicator:
up{job="web", instance="app1:8080"} 1 β Target was scraped successfully
up{job="web", instance="app2:8080"} 0 β Scrape failed (target down)
Common alerting pattern:
- alert: InstanceDown
expr: up == 0
for: 5m
labels:
severity: critical
annotations:
summary: "Instance {{ $labels.instance }} is down"
Timestamps
Prometheus timestamps:
β Millisecond precision (Unix epoch)
β Added automatically at scrape time
β Targets CAN provide their own timestamps (honor_timestamps: true)
β Recording rules use evaluation time as timestamp
Example:
Scraped at: 2024-11-15T10:30:00.000Z
Timestamp: 1731667800000 (milliseconds since epoch)
Value: 1027.0
π‘ Key Takeaway for Exam
Staleness = Series not seen in 5 minutes are marked stale Lookback delta = 5 minutes default (
--query.lookback-delta) up = 0 means target is down (scrape failed) up = 1 means target is healthy (scrape succeeded) Stale series return StaleNaN, not the last known value
2.9 π FEDERATION
What is Federation?
Federation allows one Prometheus server to scrape metrics from another Prometheus server. This is used for hierarchical monitoring setups.
Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββ
β GLOBAL PROMETHEUS (Tier 2) β
β Scrapes aggregated metrics from regional serversβ
β β
β scrape_configs: β
β - job_name: 'federate' β
β honor_labels: true β
β metrics_path: '/federate' β
β params: β
β 'match[]': β
β - '{job="prometheus"}' β
β - 'up' β
β - 'sum:http_requests:rate5m' β
β static_configs: β
β - targets: β
β - 'prometheus-us-east:9090' β
β - 'prometheus-eu-west:9090' β
ββββββββββββββββββββ¬βββββββββββββββ¬ββββββββββββββββ
β β
/federateβ β/federate
βΌ βΌ
ββββββββββββββββββββββββ ββββββββββββββββββββββββ
β REGIONAL PROMETHEUS β β REGIONAL PROMETHEUS β
β (US-East, Tier 1) β β (EU-West, Tier 1) β
β Scrapes all local β β Scrapes all local β
β targets directly β β targets directly β
ββββββββββββββββββββββββ ββββββββββββββββββββββββ
Key Points
1. The /federate endpoint returns metrics matching the given selectors
2. honor_labels: true is ESSENTIAL (preserves original job/instance)
3. Use recording rules on Tier 1 to pre-aggregate data
4. Federation is for HIERARCHICAL setups, not for HA
5. Don't federate raw data β federate aggregated/recording rule data
π‘ Key Takeaway for Exam
Federation = Prometheus scraping another Prometheus Uses
/federateendpoint withmatch[]parameters honor_labels: true is required Used for hierarchical monitoring (regional β global) Federate aggregated data, not raw metrics
2.10 π REMOTE READ / REMOTE WRITE
Why Remote Storage?
Problem: Prometheus local storage is:
β Single-node (not distributed)
β Limited by local disk
β Limited retention (typically days/weeks)
β No built-in replication
Solution: Remote Read/Write to external long-term storage
Architecture
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PROMETHEUS SERVER β
β β
β Scrape β TSDB (local) βββ¬βββ Remote Write βββ β
β β β
β β ββββββββββββββββββββ β
β ββββ β Remote Storage β β
β β (Thanos, Cortex, β β
β Query β TSDB (local) βββ¬βββ β VictoriaMetrics, β β
β β β InfluxDB, etc.) β β
β ββββ β β β
β ββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Configuration
# prometheus.yml
remote_write:
- url: "http://thanos-receive:19291/api/v1/receive"
queue_config:
max_samples_per_send: 1000
batch_send_deadline: 5s
write_relabel_configs:
- source_labels: [__name__]
regex: 'go_.*'
action: drop # Don't send Go metrics to remote
remote_read:
- url: "http://thanos-query:10902/api/v1/read"
read_recent: true # Also read recent data from remote
Remote Storage Solutions
βββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββ
β Solution β Description β
βββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββ€
β Thanos β CNCF project, adds HA + long-term β
β β storage + global query to Prometheus β
βββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββ€
β Cortex β Horizontally scalable Prometheus β
β β as a service (now Grafana Mimir) β
βββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββ€
β VictoriaMetrics β High-performance TSDB, drop-in β
β β replacement for Prometheus storage β
βββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββ€
β InfluxDB β Time-series database with remote β
β β write support β
βββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββ€
β Grafana Mimir β Successor to Cortex, scalable β
β β long-term storage β
βββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββ
π‘ Key Takeaway for Exam
Remote Write = Send data to external storage (Thanos, Cortex, etc.) Remote Read = Query data from external storage Prometheus local storage is single-node, not distributed Remote storage solves: long-term retention, HA, global queries
remote_writeandremote_readconfigured in prometheus.yml
2.11 π REAL-WORLD SCENARIOS
Scenario 1: Choosing the Right Metric Type
You're instrumenting a new payment service. Which metric type for each?
1. Total number of payments processed
β COUNTER β
(only goes up, cumulative count)
β Name: payment_processed_total
2. Current number of active payment sessions
β GAUGE β
(goes up and down)
β Name: payment_active_sessions
3. Distribution of payment processing times
β HISTOGRAM β
(need to aggregate across 10 instances)
β Name: payment_processing_duration_seconds
4. Current CPU temperature of the server
β GAUGE β
(fluctuates up and down)
β Name: node_hwmon_temp_celsius
5. Total bytes sent to the payment gateway
β COUNTER β
(cumulative, only increases)
β Name: payment_gateway_bytes_sent_total
Scenario 2: Debugging a Storage Issue
Problem: Prometheus is running out of disk space.
Investigation:
$ du -sh /prometheus/data/
450G /prometheus/data/
$ promtool tsdb list /prometheus/data/
BLOCK ULID MIN TIME MAX TIME DURATION NUM SAMPLES SIZE
01HXYZ... 1699000000 1699007200 2h0m0s 50000000 2.1GB
... (hundreds of blocks)
Solution:
1. Reduce retention:
--storage.tsdb.retention.time=7d (was 30d)
--storage.tsdb.retention.size=100GB
2. Drop unnecessary metrics:
metric_relabel_configs:
- source_labels: [__name__]
regex: 'go_.*|process_.*'
action: drop
3. Reduce scrape frequency for non-critical targets:
scrape_interval: 60s (was 15s)
2.12 π EXAM-STYLE QUESTIONS (30 Questions)
Question 1
Which of the following is NOT a component of the Prometheus server?
A) Retrieval Module (Scrape Manager) B) Storage Module (TSDB) C) Alertmanager D) HTTP Server (API & UI)
β Answer & Explanation
Correct Answer: C
Alertmanager is a separate component in the Prometheus ecosystem, NOT part of the Prometheus server itself. The Prometheus server consists of three internal modules: Retrieval (scraping), Storage (TSDB), and HTTP Server (API/UI). Alertmanager receives alerts from Prometheus but runs as its own binary.
Question 2
What is the default scrape interval in Prometheus?
A) 10 seconds B) 15 seconds C) 30 seconds D) 1 minute
β Answer & Explanation
Correct Answer: D
The default scrape_interval in Prometheus is 1 minute (60s). This can be overridden globally in the global section or per-job in scrape_configs. Most production setups override this to 15s or 30s for more granular data.
Question 3
Which metric type should you use to track the total number of HTTP requests served by your application?
A) Gauge B) Counter C) Histogram D) Summary
β Answer & Explanation
Correct Answer: B
A Counter is the correct type for tracking cumulative counts that only increase (like total HTTP requests). Counters are monotonically increasing and reset to zero only on process restart. You would then use rate() or increase() in PromQL to calculate the per-second rate or total increase over a time window.
Question 4
What is the key difference between a Histogram and a Summary in Prometheus?
A) Histograms track counts while Summaries track durations B) Histogram quantiles are calculated server-side and are aggregatable; Summary quantiles are calculated client-side and are NOT aggregatable C) Summaries are more accurate than Histograms in all cases D) Histograms cannot calculate percentiles
β Answer & Explanation
Correct Answer: B
The fundamental difference is WHERE quantiles are calculated and whether they can be aggregated. Histograms store bucket counts, and quantiles are calculated at query time using histogram_quantile() β this allows aggregation across instances. Summaries pre-calculate quantiles in the application code, and averaging percentiles across instances is mathematically invalid.
Question 5
What does the up metric indicate in Prometheus?
A) The uptime of the Prometheus server B) The CPU utilization of the target C) Whether the last scrape of a target was successful (1) or failed (0) D) The number of active targets
β Answer & Explanation
Correct Answer: C
The up metric is automatically generated by Prometheus for every scrape target. up = 1 means the last scrape was successful (target is healthy). up = 0 means the scrape failed (target is down or unreachable). Itβs the most fundamental health check metric in Prometheus.
Question 6
Which of the following is a valid Prometheus metric name?
A) http-requests-total B) http_requests_total C) 2http_requests D) http.requests.total
β Answer & Explanation
Correct Answer: B
Prometheus metric names must match the regex [a-zA-Z_:][a-zA-Z0-9_:]*. They must start with a letter, underscore, or colon, and can only contain letters, digits, underscores, and colons. Hyphens (A), leading digits (C), and dots (D) are not allowed.
Question 7
What happens to a counter metric when the application restarts?
A) It continues from the last value B) It resets to zero C) It becomes a gauge D) It is deleted from Prometheus
β Answer & Explanation
Correct Answer: B
When an application restarts, all in-memory counters reset to zero. Prometheus functions like rate() and increase() automatically detect and handle counter resets. They see the value drop from a high number to zero and adjust the calculation accordingly.
Question 8
Which PromQL function should you use with a Counter metric to get the per-second rate?
A) avg()
B) sum()
C) rate()
D) histogram_quantile()
β Answer & Explanation
Correct Answer: C
The rate() function calculates the per-second average rate of increase of a counter over a specified time window. For example, rate(http_requests_total[5m]) gives the per-second rate of HTTP requests averaged over the last 5 minutes. You should NEVER use raw counter values directly β always use rate(), increase(), or irate().
Question 9
What is the default data retention period in Prometheus?
A) 7 days B) 15 days C) 30 days D) 90 days
β Answer & Explanation
Correct Answer: B
The default retention period in Prometheus is 15 days. This can be changed using the --storage.tsdb.retention.time flag (e.g., --storage.tsdb.retention.time=30d). You can also set a size-based retention with --storage.tsdb.retention.size.
Question 10
What is the purpose of the Write-Ahead Log (WAL) in Prometheus?
A) To store long-term historical data B) To compress old data blocks C) To prevent data loss
π PCA EXAM β DOMAIN 2: PROMETHEUS FUNDAMENTALS (Continued)
Questions 10β30 + Summary Cheat Sheet
Question 10 (Complete Answer)
What is the purpose of the Write-Ahead Log (WAL) in Prometheus?
A) To store long-term historical data B) To compress old data blocks C) To prevent data loss if Prometheus crashes before the Head block is compacted to disk D) To replicate data to remote storage
β Answer & Explanation
Correct Answer: C
The WAL (Write-Ahead Log) ensures crash recovery. Every incoming sample is first written to the WAL on disk before being stored in the in-memory Head block. If Prometheus crashes, it replays the WAL on restart to recover any data that hadnβt been compacted to a persistent block yet. The WAL is NOT for long-term storage (A), compression (B), or remote replication (D).
Question 11
In the Prometheus data model, what makes two time series different from each other?
A) Different metric names only B) Different timestamps only C) Different metric names OR different label combinations D) Different values
β Answer & Explanation
Correct Answer: C
A time series in Prometheus is uniquely identified by its metric name + the complete set of label key-value pairs. If either the metric name differs OR any label value differs, itβs a completely separate time series. For example, http_requests_total{method="GET"} and http_requests_total{method="POST"} are two distinct time series. Timestamps and values are the data WITHIN a time series, not identifiers.
Question 12
Which labels are automatically added by Prometheus to every scraped time series?
A) cluster and region
B) job and instance
C) host and port
D) service and namespace
β Answer & Explanation
Correct Answer: B
Prometheus automatically adds two labels to every scraped time series:
job: Thejob_namefrom the scrape configurationinstance: The<host>:<port>of the target being scraped
Labels like cluster, region, service, and namespace can be added via relabeling or external_labels, but they are NOT automatic. host and port are not standard Prometheus labels.
Question 13
What does the honor_labels: true configuration do in a scrape job?
A) It renames all labels to lowercase
B) It prevents Prometheus from overriding labels that already exist in the scraped data (like job and instance)
C) It drops all labels from the scraped metrics
D) It adds additional labels from the Prometheus server
β Answer & Explanation
Correct Answer: B
When honor_labels: true is set, Prometheus keeps the job and instance labels (and any other conflicting labels) as they appear in the scraped data, instead of overwriting them with the values from the scrape configuration. This is critical for Pushgateway and Federation, where the original labels must be preserved. By default, honor_labels is false, meaning Prometheus overwrites conflicting labels.
Question 14
Which of the following is the correct way to calculate average request duration from a Histogram?
A) avg(http_request_duration_seconds)
B) histogram_quantile(0.5, rate(http_request_duration_seconds_bucket[5m]))
C) rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m])
D) sum(http_request_duration_seconds_sum) / sum(http_request_duration_seconds_count)
β Answer & Explanation
Correct Answer: C
To calculate the average request duration from a Histogram, you divide the rate of the _sum (total duration) by the rate of the _count (total number of observations). Both _sum and _count are counters, so you must use rate(). Option B calculates the median (p50), not the average. Option A is invalid syntax for histograms. Option D uses raw counter values without rate(), which is incorrect.
Question 15
What is the Prometheus Pushgateway used for?
A) Replacing the pull model for all monitoring targets B) Pushing metrics from behind a firewall C) Accepting metrics from short-lived batch jobs that cannot be scraped D) Sending alerts to external systems
β Answer & Explanation
Correct Answer: C
The Pushgateway is specifically designed for short-lived batch jobs (cron jobs, CI/CD pipelines, serverless functions) that may complete before Prometheus has a chance to scrape them. The job pushes its metrics to the Pushgateway, and Prometheus scrapes the Pushgateway. It should NOT be used as a general replacement for the pull model (A), for firewall traversal (B β use federation or remote write), or for alerting (D β thatβs Alertmanager).
Question 16
What is the default value of evaluation_interval in Prometheus?
A) 10 seconds B) 15 seconds C) 30 seconds D) 1 minute
β Answer & Explanation
Correct Answer: D
The default evaluation_interval is 1 minute (60s), same as the default scrape_interval. This controls how often Prometheus evaluates alerting and recording rules. Itβs configured in the global section of prometheus.yml.
Question 17
Which of the following metric names follows Prometheus naming conventions correctly?
A) http_request_duration_milliseconds
B) httpRequestDurationSeconds
C) http_request_duration_seconds
D) http_request_duration_seconds_counter
β Answer & Explanation
Correct Answer: C
Prometheus naming conventions require:
- snake_case (eliminates B which is camelCase)
- Base units (eliminates A β should be seconds, not milliseconds)
- No type in the name (eliminates D β donβt put βcounterβ in the name)
http_request_duration_seconds follows all conventions: snake_case, base unit (seconds), and no redundant type suffix.
Question 18
What does the le label in a Histogram bucket represent?
A) βlast eventβ β the timestamp of the last observation B) βless than or equalβ β the upper bound of the bucket C) βlog entryβ β the log level of the observation D) βlatency estimateβ β the estimated latency
β Answer & Explanation
Correct Answer: B
The le label stands for βless than or equalβ. It represents the upper inclusive bound of a histogram bucket. For example, http_request_duration_seconds_bucket{le="0.5"} counts all requests that took β€ 0.5 seconds. Histogram buckets are cumulative, meaning the le="0.5" bucket includes all requests counted in the le="0.1" bucket.
Question 19
Which of the following is TRUE about Prometheus local storage (TSDB)?
A) It is a distributed database that replicates data across nodes B) It stores data in a single file for simplicity C) It stores data in time-ordered blocks, with the most recent data in an in-memory Head block D) It requires an external database like PostgreSQL
β Answer & Explanation
Correct Answer: C
Prometheus TSDB stores data in time-ordered blocks. The most recent ~2 hours of data is kept in an in-memory Head block (protected by the WAL). When the Head block is full, itβs compacted to disk as an immutable persisted block. Smaller blocks are periodically merged into larger ones through compaction. TSDB is NOT distributed (A), NOT a single file (B), and does NOT require an external database (D) β itβs fully self-contained.
Question 20
What is the purpose of external_labels in the Prometheus global configuration?
A) To label targets that are outside the network B) To add labels to all time series and alerts produced by this Prometheus instance, useful for federation and remote write C) To override labels from scraped targets D) To configure labels for the Alertmanager
β Answer & Explanation
Correct Answer: B
external_labels are added to all time series and alerts produced by this Prometheus server. They are primarily used in federation and remote write scenarios to identify which Prometheus instance produced the data. For example, you might set cluster: 'us-east-1' and replica: 'prometheus-1' so that when data is sent to a central Thanos or Cortex, you can distinguish the source.
Question 21
Which of the following correctly describes the relationship between a Prometheus βjobβ and an βinstanceβ?
A) A job is a single target; an instance is a group of targets B) A job is a logical group of targets; an instance is a specific target endpoint within that job C) They are the same thing D) A job runs on an instance
β Answer & Explanation
Correct Answer: B
A job is a logical grouping of targets that serve the same purpose (e.g., job="web-app" groups all web application servers). An instance is a specific target endpoint within that job (e.g., instance="app1:8080", instance="app2:8080"). One job can have many instances. The job label comes from job_name in the scrape config, and the instance label comes from the targetβs address.
Question 22
What is the exposition format content type header that Prometheus expects?
A) application/json
B) text/html
C) text/plain; version=0.0.4
D) application/xml
β Answer & Explanation
Correct Answer: C
The Prometheus exposition format uses the content type text/plain; version=0.0.4. This is the standard text-based format for the /metrics endpoint. The newer OpenMetrics format uses application/openmetrics-text. JSON (A), HTML (B), and XML (D) are not used for Prometheus metric exposition.
Question 23
Which of the following scenarios would cause a counter to reset to zero?
A) When the counter reaches its maximum value B) When the Prometheus server restarts C) When the application/process that exposes the counter restarts D) When the scrape interval changes
β Answer & Explanation
Correct Answer: C
A counter resets to zero when the application or process that maintains the counter restarts, because the counter is stored in the applicationβs memory. When the process restarts, all in-memory state is lost. Prometheus server restarts (B) donβt affect the counter β Prometheus stores historical data in TSDB. Counters donβt have a maximum value (A), and scrape interval changes (D) donβt affect counter values.
Question 24
What is the purpose of the /federate endpoint in Prometheus?
A) To send alerts to Alertmanager B) To allow one Prometheus server to scrape selected metrics from another Prometheus server C) To expose metrics to Grafana D) To push metrics to remote storage
β Answer & Explanation
Correct Answer: B
The /federate endpoint allows hierarchical Prometheus setups where a βglobalβ Prometheus server scrapes aggregated metrics from βregionalβ Prometheus servers. You specify which metrics to federate using match[] URL parameters. This is useful for multi-datacenter or multi-team setups where each team has their own Prometheus, and a central server aggregates key metrics.
Question 25
Which of the following is NOT a valid metric type in Prometheus?
A) Counter B) Gauge C) Histogram D) Timer
β Answer & Explanation
Correct Answer: D
Prometheus has exactly four metric types: Counter, Gauge, Histogram, and Summary. βTimerβ is NOT a Prometheus metric type. Some other monitoring systems (like StatsD or Dropwizard) have a Timer type, but in Prometheus, you would use a Histogram or Summary to track durations/timers.
Question 26
What does the scrape_timeout configuration control?
A) How long Prometheus waits before marking a target as permanently down B) The maximum time Prometheus waits for a single scrape request to complete C) How long metrics are retained in the TSDB D) The interval between scrape retries after a failure
β Answer & Explanation
Correct Answer: B
scrape_timeout defines the maximum duration Prometheus will wait for a single scrape HTTP request to complete. If the target doesnβt respond within this time, the scrape is considered failed. The default is 10 seconds. It must be less than or equal to scrape_interval. It does NOT control permanent down detection (A), retention (C), or retry intervals (D).
Question 27
In Prometheus TSDB, what is βcompactionβ?
A) Deleting old data to free disk space B) Compressing individual samples to save memory C) Merging smaller time-ordered blocks into larger blocks for efficiency D) Encrypting stored data for security
β Answer & Explanation
Correct Answer: C
Compaction is the process of merging smaller blocks into larger blocks. For example, three 2-hour blocks are merged into one 6-hour block, and three 6-hour blocks into one 18-hour block. This reduces the number of files on disk, improves query performance (fewer blocks to scan), and achieves better compression ratios. Compaction is NOT deletion (A β thatβs retention), NOT individual sample compression (B), and NOT encryption (D).
Question 28
Which of the following PromQL queries correctly identifies targets that are currently down?
A) up == 1
B) up == 0
C) down == 1
D) target_status == "down"
β Answer & Explanation
Correct Answer: B
The up metric is automatically generated by Prometheus. up == 1 means the target is healthy (scrape succeeded). up == 0 means the target is DOWN (scrape failed). There is no down metric (C) or target_status metric (D) in Prometheus. A common alerting rule is: expr: up == 0 with for: 5m to alert when a target has been down for 5 minutes.
Question 29
What is the recommended approach for long-term storage of Prometheus data?
A) Increase --storage.tsdb.retention.time to 10 years
B) Use remote write to send data to a long-term storage solution like Thanos, Cortex, or VictoriaMetrics
C) Export all data to CSV files daily
D) Use the Pushgateway for long-term storage
β Answer & Explanation
Correct Answer: B
The recommended approach for long-term storage is to use remote write to send data to a purpose-built long-term storage solution like Thanos, Cortex (Grafana Mimir), or VictoriaMetrics. While you CAN increase local retention (A), Prometheusβs local TSDB is not designed for years of data β itβs single-node and limited by local disk. CSV export (C) is impractical. Pushgateway (D) is for batch jobs, not storage.
Question 30
Which of the following statements about Prometheus metric labels is TRUE?
A) Label names starting with __ (double underscore) are reserved for internal use
B) Labels can contain any character including spaces and special symbols
C) High-cardinality labels like user_id are recommended for detailed monitoring
D) Label values must be numeric
β Answer & Explanation
Correct Answer: A
Label names starting with __ (double underscore) are reserved for internal use by Prometheus. Examples include __address__, __scheme__, __metrics_path__, and __meta_* labels from service discovery. These are used during relabeling and are typically dropped before storage. Label names must match [a-zA-Z_][a-zA-Z0-9_]* (B is wrong). High-cardinality labels cause memory explosion (C is wrong). Label values are strings, not numeric (D is wrong).
β DOMAIN 2 REVISION SUMMARY CHEAT SHEET
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DOMAIN 2 CHEAT SHEET β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β ARCHITECTURE: β
β Prometheus Server = Retrieval + TSDB + HTTP Server β
β Ecosystem = Server + Exporters + Client Libs + Pushgateway β
β + Alertmanager + Service Discovery + Grafana β
β Alertmanager & Grafana are SEPARATE components β
β β
β DATA MODEL: β
β Time Series = Metric Name + Labels β
β Samples = Timestamp (ms) + Value (float64) β
β Auto labels: job, instance β
β Reserved labels: __ prefix (internal use) β
β β οΈ High cardinality labels = BAD (memory explosion) β
β β
β NAMING: snake_case, base units (seconds/bytes), _total suffix β
β β
β METRIC TYPES: β
β Counter β Only β (rate(), increase()) β
β Gauge β β and β (use directly) β
β Histogram β Buckets (le), _sum, _count, aggregatable β
β
β Summary β Quantiles, _sum, _count, NOT aggregatable β β
β β Prefer Histogram for distributed systems β
β β
β CONFIG (prometheus.yml): β
β global β scrape_interval (default 1m), evaluation_interval (1m) β
β scrape_configs β job_name, targets, SD, relabeling β
β rule_files β alerting + recording rules β
β alerting β Alertmanager connection β
β honor_labels: true β Preserve original labels (Pushgateway) β
β β
β STORAGE (TSDB): β
β Head block β ~2h in memory, WAL for crash safety β
β Persisted blocks β Compressed, immutable, on disk β
β Compaction β Merge small blocks into larger ones β
β Retention β --storage.tsdb.retention.time=15d (default) β
β Efficiency β ~1-2 bytes per sample β
β β
β EXPOSITION: text/plain, /metrics endpoint β
β Special metrics: up (0/1), scrape_duration_seconds β
β β
β STALENESS: 5 min lookback delta, stale = StaleNaN β
β FEDERATION: /federate endpoint, honor_labels: true β
β REMOTE: remote_write/remote_read β Thanos, Cortex, VictoriaM. β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
π― NEXT STEPS
Once youβve thoroughly studied Domain 2 and can confidently answer all 30 questions:
Reply with βDomain 2 Completeβ and Iβll provide the comprehensive guide for Domain 3: PromQL (28%) β the MOST heavily weighted domain on the exam!
This will cover:
- Selectors (Instant & Range Vectors)
- Label Matchers (=, !=, =
, !)- All Operators (Arithmetic, Comparison, Logical, Vector Matching)
- All Aggregation Operators (sum, avg, topk, etc.)
- All Functions (rate, irate, increase, histogram_quantile, etc.)
- Recording Rules
- Subqueries
- 30+ Exam-Style Questions
Since PromQL is 28% of the exam, this will be the most detailed domain! π