π PCA EXAM β DOMAIN 4: INSTRUMENTATION & EXPORTERS (16%)
Complete One-Stop Guide: Theory β Examples β Exam Questions
TABLE OF CONTENTS
4.1 What is Instrumentation?
4.2 Instrumentation Approaches
4.3 Prometheus Client Libraries (Overview)
4.4 Instrumenting with Go Client Library
4.5 Instrumenting with Python Client Library
4.6 Instrumenting with Java Client Library
4.7 Instrumentation Best Practices
4.8 Metric Naming Conventions (Deep Dive)
4.9 Base Units & Suffixes
4.10 Choosing the Right Metric Type (Decision Guide)
4.11 What are Exporters?
4.12 Node Exporter (Deep Dive)
4.13 Blackbox Exporter (Deep Dive)
4.14 cAdvisor (Container Metrics)
4.15 Database Exporters (MySQL, PostgreSQL, Redis)
4.16 Other Common Exporters
4.17 Custom Exporters
4.18 Pushgateway β Detailed Usage
4.19 Exposition Format Deep Dive
4.20 Real-World Instrumentation Scenarios
4.21 EXAM-STYLE QUESTIONS (30 Questions with Answers)
4.1 π WHAT IS INSTRUMENTATION?
Definition
Instrumentation is the process of adding monitoring code to your application so that it exposes metrics about its internal behavior, performance, and health.
Think of it as: Installing sensors inside a car engine to measure
temperature, oil pressure, RPM, and fuel flow β instead of just
looking at the speedometer from outside.
Why Instrumentation Matters
Without Instrumentation (Black-Box only):
β "The website is slow" β That's all you know
β You have to SSH into servers, read logs, guess the problem
With Instrumentation (White-Box):
β "The /api/checkout endpoint has p99 latency of 3s"
β "The database connection pool is at 95% capacity"
β "The payment-service error rate spiked to 12%"
β "The cache hit ratio dropped from 95% to 40%"
β You know EXACTLY where the problem is!
What Gets Instrumented?
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β INSTRUMENTATION LAYERS β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β Application Layer: β
β β Request counts, latencies, error rates β
β β Business metrics (signups, purchases) β
β β Queue depths, processing times β
β β Cache hit/miss ratios β
β β
β Runtime Layer: β
β β Garbage collection pauses β
β β Thread/goroutine counts β
β β Memory allocation rates β
β β Heap usage β
β β
β Infrastructure Layer (via Exporters): β
β β CPU, memory, disk, network β
β β Container metrics β
β β Database performance β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
4.2 π INSTRUMENTATION APPROACHES
Approach 1: Library Instrumentation (Most Common)
You add a Prometheus client library to your application code.
The library provides APIs to create and update metrics.
The library automatically exposes a /metrics HTTP endpoint.
Flow:
Your Code β Client Library β /metrics endpoint β Prometheus scrapes
Example:
import prometheus_client # Python
counter = Counter('requests_total', 'Total requests')
counter.inc() # Increment on each request
# Library auto-exposes at http://localhost:8000/metrics
Approach 2: Exporter (Third-Party Systems)
For systems you don't control (MySQL, Redis, Linux kernel),
you use a separate exporter process.
Flow:
Third-Party System β Exporter β /metrics endpoint β Prometheus scrapes
Example:
MySQL Database β MySQL Exporter (port 9104) β Prometheus scrapes
Linux Kernel β Node Exporter (port 9100) β Prometheus scrapes
Approach 3: Service Mesh / Sidecar (Advanced)
In Kubernetes with Istio/Linkerd, the sidecar proxy automatically
captures metrics without any code changes.
Flow:
App β Envoy Sidecar (auto-instruments) β /metrics β Prometheus
Limitation: Only captures network-level metrics (latency, errors, traffic)
Cannot capture business logic metrics
Approach 4: Pushgateway (Batch Jobs)
For short-lived jobs that can't be scraped.
Flow:
Cron Job β Push metrics β Pushgateway β Prometheus scrapes Pushgateway
4.3 π PROMETHEUS CLIENT LIBRARIES (Overview)
Official Client Libraries
| Language | Library | Maturity | Maintainer |
|---|---|---|---|
| Go | client_golang |
βββββ Most mature | Prometheus Team |
| Python | prometheus_client |
ββββ | Prometheus Team |
| Java | simpleclient / client_java |
ββββ | Prometheus Team |
| Ruby | prometheus-client |
βββ | Prometheus Team |
Community Client Libraries
| Language | Library | Notes |
|---|---|---|
| .NET/C# | prometheus-net |
Very popular, well-maintained |
| Rust | prometheus crate |
Good for Rust ecosystem |
| Node.js | prom-client |
Most popular for Node |
| PHP | promphp/prometheus_client_php |
Community maintained |
| C++ | prometheus-cpp |
Community maintained |
| Bash | prometheus-bash-exporter |
For shell scripts |
What All Client Libraries Provide
Every Prometheus client library provides:
1. Metric Types:
β Counter, Gauge, Histogram, Summary
2. Registry:
β Central place where all metrics are registered
β Default registry is created automatically
β Custom registries for isolation
3. HTTP Handler:
β Exposes /metrics endpoint in the correct exposition format
β Handles content negotiation
4. Default Metrics:
β Process metrics (CPU, memory, file descriptors)
β Runtime metrics (GC, threads, goroutines)
β These are collected automatically!
5. Label Support:
β Create metric families with label dimensions
β Validate label names and values
4.4 π INSTRUMENTING WITH GO CLIENT LIBRARY
Go is the language Prometheus is written in, so its client library is the most mature.
Installation
go get github.com/prometheus/client_golang/prometheus
go get github.com/prometheus/client_golang/prometheus/promhttp
Complete Example: Web Server Instrumentation
package main
import (
"net/http"
"time"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
// βββ STEP 1: Define Metrics βββ
// Counter: Total HTTP requests
var httpRequestsTotal = prometheus.NewCounterVec(
prometheus.CounterOpts{
Namespace: "myapp",
Name: "http_requests_total",
Help: "Total number of HTTP requests",
},
[]string{"method", "endpoint", "status"}, // Label names
)
// Gauge: Currently active requests
var activeRequests = prometheus.NewGauge(
prometheus.GaugeOpts{
Namespace: "myapp",
Name: "active_requests",
Help: "Number of currently active requests",
},
)
// Histogram: Request duration
var requestDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Namespace: "myapp",
Name: "http_request_duration_seconds",
Help: "HTTP request duration in seconds",
Buckets: []float64{0.01, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10},
},
[]string{"method", "endpoint"},
)
// βββ STEP 2: Register Metrics βββ
func init() {
prometheus.MustRegister(httpRequestsTotal)
prometheus.MustRegister(activeRequests)
prometheus.MustRegister(requestDuration)
}
// βββ STEP 3: Use Metrics in Handlers βββ
func helloHandler(w http.ResponseWriter, r *http.Request) {
start := time.Now()
activeRequests.Inc() // Gauge: increment
defer activeRequests.Dec() // Gauge: decrement when done
// ... do work ...
w.WriteHeader(http.StatusOK)
w.Write([]byte("Hello, World!"))
duration := time.Since(start).Seconds()
// Counter: increment with labels
httpRequestsTotal.WithLabelValues(r.Method, "/hello", "200").Inc()
// Histogram: observe duration
requestDuration.WithLabelValues(r.Method, "/hello").Observe(duration)
}
// βββ STEP 4: Expose /metrics Endpoint βββ
func main() {
http.HandleFunc("/hello", helloHandler)
http.Handle("/metrics", promhttp.Handler()) // Auto /metrics!
http.ListenAndServe(":8080", nil)
}
What the /metrics Endpoint Returns
$ curl http://localhost:8080/metrics
# HELP myapp_http_requests_total Total number of HTTP requests.
# TYPE myapp_http_requests_total counter
myapp_http_requests_total{endpoint="/hello",method="GET",status="200"} 42
# HELP myapp_active_requests Number of currently active requests.
# TYPE myapp_active_requests gauge
myapp_active_requests 3
# HELP myapp_http_request_duration_seconds HTTP request duration in seconds.
# TYPE myapp_http_request_duration_seconds histogram
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="0.01"} 10
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="0.05"} 25
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="0.1"} 35
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="0.25"} 40
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="0.5"} 41
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="1"} 42
myapp_http_request_duration_seconds_bucket{endpoint="/hello",method="GET",le="+Inf"} 42
myapp_http_request_duration_seconds_sum{endpoint="/hello",method="GET"} 3.456
myapp_http_request_duration_seconds_count{endpoint="/hello",method="GET"} 42
# Default Go runtime metrics (automatically included!)
# HELP go_goroutines Number of goroutines that currently exist.
# TYPE go_goroutines gauge
go_goroutines 15
# HELP process_resident_memory_bytes Resident memory size in bytes.
# TYPE process_resident_memory_bytes gauge
process_resident_memory_bytes 25600000
4.5 π INSTRUMENTING WITH PYTHON CLIENT LIBRARY
Installation
pip install prometheus_client
Complete Example
from prometheus_client import (
Counter, Gauge, Histogram, Summary,
start_http_server, generate_latest
)
import time
import random
# βββ STEP 1: Define Metrics βββ
# Counter
REQUESTS = Counter(
'http_requests_total',
'Total HTTP requests',
['method', 'endpoint', 'status']
)
# Gauge
ACTIVE_CONNECTIONS = Gauge(
'active_connections',
'Number of active database connections'
)
# Histogram
REQUEST_LATENCY = Histogram(
'http_request_duration_seconds',
'HTTP request latency in seconds',
['method', 'endpoint'],
buckets=[0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0]
)
# Summary
RESPONSE_SIZE = Summary(
'http_response_size_bytes',
'HTTP response size in bytes',
['endpoint']
)
# βββ STEP 2: Use Metrics βββ
def handle_request(method, endpoint):
start = time.time()
# Simulate work
time.sleep(random.uniform(0.01, 0.5))
duration = time.time() - start
# Update metrics
REQUESTS.labels(method=method, endpoint=endpoint, status='200').inc()
REQUEST_LATENCY.labels(method=method, endpoint=endpoint).observe(duration)
RESPONSE_SIZE.labels(endpoint=endpoint).observe(random.randint(100, 5000))
return "OK"
# βββ STEP 3: Start Metrics Server βββ
if __name__ == '__main__':
# Start HTTP server for /metrics on port 8000
start_http_server(8000)
print("Metrics available at http://localhost:8000/metrics")
# Simulate some requests
while True:
handle_request('GET', '/api/users')
ACTIVE_CONNECTIONS.set(random.randint(5, 50))
time.sleep(1)
Flask Integration Example
from flask import Flask, request
from prometheus_client import Counter, Histogram, generate_latest
import time
app = Flask(__name__)
REQUESTS = Counter('flask_requests_total', 'Total Flask requests',
['method', 'endpoint', 'status'])
LATENCY = Histogram('flask_request_duration_seconds', 'Request latency',
['method', 'endpoint'])
@app.before_request
def before_request():
request.start_time = time.time()
@app.after_request
def after_request(response):
latency = time.time() - request.start_time
REQUESTS.labels(
method=request.method,
endpoint=request.path,
status=response.status_code
).inc()
LATENCY.labels(
method=request.method,
endpoint=request.path
).observe(latency)
return response
@app.route('/metrics')
def metrics():
return generate_latest(), 200, {'Content-Type': 'text/plain'}
@app.route('/hello')
def hello():
return "Hello, World!"
if __name__ == '__main__':
app.run(port=5000)
4.6 π INSTRUMENTING WITH JAVA CLIENT LIBRARY
Maven Dependency
<dependency>
<groupId>io.prometheus</groupId>
<artifactId>simpleclient</artifactId>
<version>0.16.0</version>
</dependency>
<dependency>
<groupId>io.prometheus</groupId>
<artifactId>simpleclient_httpserver</artifactId>
<version>0.16.0</version>
</dependency>
Complete Example
import io.prometheus.client.Counter;
import io.prometheus.client.Gauge;
import io.prometheus.client.Histogram;
import io.prometheus.client.exporter.HTTPServer;
public class MyApp {
// βββ Define Metrics βββ
static final Counter requests = Counter.build()
.name("http_requests_total")
.help("Total HTTP requests")
.labelNames("method", "status")
.register();
static final Gauge activeRequests = Gauge.build()
.name("active_requests")
.help("Currently active requests")
.register();
static final Histogram requestLatency = Histogram.build()
.name("http_request_duration_seconds")
.help("Request latency in seconds")
.buckets(0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0)
.labelNames("method")
.register();
public static void main(String[] args) throws Exception {
// Start metrics HTTP server on port 9090
HTTPServer server = new HTTPServer(9090);
// Simulate handling a request
handleRequest("GET");
}
static void handleRequest(String method) {
activeRequests.inc();
Histogram.Timer timer = requestLatency.labels(method).startTimer();
try {
// ... process request ...
requests.labels(method, "200").inc();
} finally {
timer.observeDuration(); // Record latency
activeRequests.dec();
}
}
}
Spring Boot Integration (Micrometer)
// Spring Boot 2+ uses Micrometer as the metrics facade
// Add dependency: spring-boot-starter-actuator + micrometer-registry-prometheus
// application.yml:
// management:
// endpoints:
// web:
// exposure:
// include: prometheus
// metrics:
// export:
// prometheus:
// enabled: true
// Metrics auto-exposed at: /actuator/prometheus
import io.micrometer.core.instrument.Counter;
import io.micrometer.core.instrument.MeterRegistry;
import org.springframework.stereotype.Service;
@Service
public class OrderService {
private final Counter ordersCounter;
public OrderService(MeterRegistry registry) {
ordersCounter = Counter.builder("orders_total")
.description("Total orders placed")
.tag("type", "online")
.register(registry);
}
public void placeOrder() {
ordersCounter.increment();
// ... process order
}
}
4.7 π INSTRUMENTATION BEST PRACTICES
β DO: Instrument at the Right Level
Good: Instrument at the application/endpoint level
β http_requests_total{method="GET", endpoint="/api/users", status="200"}
β http_request_duration_seconds{method="GET", endpoint="/api/users"}
Bad: Instrument at too low a level (every function call)
β myapp_function_doSomething_duration_seconds β Too granular!
β Creates too many time series, adds overhead
β DO: Use Labels Wisely (Avoid High Cardinality!)
Good: Low cardinality labels (bounded set of values)
β method: GET, POST, PUT, DELETE (4 values)
β status: 200, 404, 500 (limited values)
β endpoint: /api/users, /api/orders (bounded)
BAD: High cardinality labels (unbounded values!)
β user_id: "user_12345" β Millions of users!
β session_id: "abc123" β Unique per session!
β request_id: "xyz789" β Unique per request!
β email: "user@email.com" β Unbounded!
Why it's bad:
Each unique label combination = a new time series
1M users Γ 10 endpoints Γ 5 methods = 50M time series!
β Prometheus will run out of memory!
β DO: Use Base Units
Good:
β http_request_duration_seconds (seconds, not milliseconds)
β process_memory_bytes (bytes, not megabytes)
β network_transmit_bytes_total (bytes)
Bad:
β http_request_duration_milliseconds β Wrong unit!
β process_memory_megabytes β Wrong unit!
β DO: Use the Right Metric Type
Counting events? β Counter
Measuring current state? β Gauge
Measuring distributions? β Histogram (prefer over Summary)
Pre-computed quantiles? β Summary (single instance only)
β DO: Provide HELP Strings
Good:
Counter('http_requests_total', 'Total number of HTTP requests processed')
Bad:
Counter('http_requests_total', '') β No help text!
β DO: Use Consistent Naming
Good:
myapp_http_requests_total
myapp_http_request_duration_seconds
myapp_active_connections
Bad:
myapp_requests β Missing _total suffix for counter
myappHttpLatency β camelCase, missing unit
my-app-errors β Hyphens not allowed
β DONβT: Put Rates in Metric Names
Bad:
http_requests_per_second β Don't pre-calculate rates!
Good:
http_requests_total β Let Prometheus calculate rates with rate()
Why: Prometheus is designed to store raw counters and compute rates at query time.
Pre-computing rates loses information and makes aggregation impossible.
β DONβT: Create Metrics Dynamically Based on Unbounded Input
Bad:
for url in all_urls_visited:
Counter(f'requests_{url}_total', ...).inc()
β Creates a new metric for every URL ever visited!
Good:
Counter('requests_total', ..., ['url']).labels(url=sanitize(url)).inc()
β Use labels with bounded values instead
π‘ Key Takeaway for Exam
Low cardinality labels only (no user_id, session_id, request_id) Base units (seconds, bytes) _total suffix for counters HELP strings always Donβt pre-compute rates β store counters, use
rate()in PromQL Histogram over Summary for distributed systems
4.8 π METRIC NAMING CONVENTIONS (Deep Dive)
The Full Naming Pattern
[<namespace>_]<subsystem>_<name>_<unit>[_<suffix>]
namespace β Application or organization name (optional)
subsystem β Component within the application (optional)
name β What is being measured (required)
unit β Base unit (seconds, bytes, total, ratio) (recommended)
suffix β _total, _bucket, _sum, _count, _info, _created (auto)
Examples of Well-Named Metrics
β
http_requests_total (subsystem_name_unit)
β
node_cpu_seconds_total (namespace_subsystem_unit_suffix)
β
process_resident_memory_bytes (subsystem_name_unit)
β
go_goroutines (namespace_name)
β
jvm_memory_used_bytes (namespace_name_unit)
β
mysql_global_status_queries_total (namespace_subsystem_name_suffix)
β
http_request_duration_seconds (subsystem_name_unit)
β
app_build_info (namespace_name_suffix)
Naming Rules (Exam Critical!)
1. MUST match regex: [a-zA-Z_:][a-zA-Z0-9_:]*
β Start with letter, underscore, or colon
β Only letters, digits, underscores, colons
2. MUST use snake_case
β httpRequestsTotal (camelCase)
β http-requests-total (kebab-case)
β HTTP_REQUESTS (UPPER_CASE)
3. MUST use base units
β seconds (not ms, not minutes)
β bytes (not KB, not MB)
β celsius (not fahrenheit)
β ratio (0 to 1, not percentage 0 to 100)
4. SHOULD include unit in name
β
http_request_duration_seconds
β
node_memory_MemTotal_bytes
5. MUST use _total suffix for counters
β
http_requests_total
β http_requests (ambiguous β is it a counter or gauge?)
6. MUST NOT include metric type in name
β http_requests_counter_total (redundant!)
β cpu_usage_gauge (redundant!)
7. Colons (:) are RESERVED for recording rules
β http:requests:total (in raw instrumentation)
β
job:http_requests_total:rate5m (in recording rules)
8. Labels starting with __ are RESERVED
β __custom_label (reserved for internal use)
β
custom_label (fine)
4.9 π BASE UNITS & SUFFIXES
Standard Base Units
| Unit | Suffix | Example |
|---|---|---|
| Seconds | _seconds |
http_request_duration_seconds |
| Bytes | _bytes |
process_resident_memory_bytes |
| Ratio (0-1) | _ratio |
go_memstats_alloc_ratio |
| Celsius | _celsius |
node_hwmon_temp_celsius |
| Meters | _meters |
altitude_meters |
| Joules | _joules |
energy_consumed_joules |
| Volts | _volts |
psu_voltage_volts |
| Amperes | _amperes |
psu_current_amperes |
| Grams | _grams |
weight_grams |
Standard Suffixes
| Suffix | Used For | Example |
|---|---|---|
_total |
Counters | http_requests_total |
_bucket |
Histogram buckets (auto) | http_request_duration_seconds_bucket |
_sum |
Histogram/Summary sum (auto) | http_request_duration_seconds_sum |
_count |
Histogram/Summary count (auto) | http_request_duration_seconds_count |
_info |
Info metrics (gauge = 1) | node_os_info |
_created |
Creation timestamp | process_start_time_seconds_created |
Info Metrics Pattern
# Info metrics are gauges that always have value 1
# They carry metadata as labels
node_os_info{os="linux", version="22.04", name="Ubuntu"} 1
app_build_info{version="1.2.3", commit="abc123", branch="main"} 1
jvm_info{version="17.0.1", vendor="OpenJDK"} 1
# Useful for joining with other metrics in PromQL:
sum(rate(http_requests_total[5m])) by (instance)
* on(instance) group_left(version)
app_build_info
4.10 π CHOOSING THE RIGHT METRIC TYPE (Decision Guide)
START HERE: What are you measuring?
β
ββ "How many times did something happen?" (cumulative count)
β ββ COUNTER β
β Examples: requests, errors, bytes sent, tasks completed
β PromQL: rate(), increase()
β
ββ "What is the current value right now?" (snapshot)
β ββ GAUGE β
β Examples: temperature, memory usage, active connections, queue size
β PromQL: use directly, avg(), max(), min()
β
ββ "What is the distribution of values?" (latency, size)
β β
β ββ Multiple instances? Need to aggregate?
β β ββ HISTOGRAM β
β β Examples: request latency across 10 servers
β β PromQL: histogram_quantile()
β β
β ββ Single instance? Need exact quantiles?
β ββ SUMMARY β
β Examples: request latency on a single server
β PromQL: read quantile label directly
β
ββ "Is it a yes/no or metadata?"
ββ GAUGE (value = 1) with info labels β
Examples: app_build_info, node_os_info
4.11 π WHAT ARE EXPORTERS?
Definition
An exporter is a standalone process that collects metrics from a third-party system and exposes them in Prometheus format on a /metrics endpoint.
Why exporters exist:
β Many systems (MySQL, Linux, Redis) don't natively speak Prometheus format
β Exporters act as translators/bridges
β They query the third-party system's native metrics API
β They convert the data to Prometheus exposition format
β They expose it on an HTTP endpoint for Prometheus to scrape
Architecture:
βββββββββββββ βββββββββββββββββ ββββββββββββββ
β Prometheusβββββββ Exporter βββββββ 3rd Party β
β Server βscrapeβ (Translator) βqueryβ System β
β βββββββ βββββββ β
β β/metricsβ βdata β (MySQL, β
β β β :9104 β β Linux, β
βββββββββββββ βββββββββββββββββ β Redis) β
ββββββββββββββ
Exporter vs Client Library
| Aspect | Client Library | Exporter |
|---|---|---|
| Where it runs | Inside your application | Separate process |
| What it monitors | Your application code | Third-party systems |
| Code changes | Yes (add instrumentation) | No (standalone) |
| Examples | Go/Python/Java client | Node Exporter, MySQL Exporter |
| Port | Same as your app (e.g., 8080) | Separate port (e.g., 9100) |
4.12 π NODE EXPORTER (Deep Dive)
What is Node Exporter?
The Node Exporter is the most widely used Prometheus exporter. It exposes hardware and OS-level metrics for Linux/Unix systems.
Key Facts
Port: 9100 (default)
Endpoint: /metrics
Binary: node_exporter
Maintainer: Prometheus Team (official)
OS Support: Linux (full), macOS, FreeBSD, OpenBSD (partial)
Installation & Running
# Download
wget https://github.com/prometheus/node_exporter/releases/download/v1.7.0/node_exporter-1.7.0.linux-amd64.tar.gz
tar xvfz node_exporter-*.tar.gz
cd node_exporter-*
# Run
./node_exporter
# Or as a systemd service:
# /etc/systemd/system/node_exporter.service
[Unit]
Description=Node Exporter
[Service]
ExecStart=/usr/local/bin/node_exporter
[Install]
WantedBy=multi-user.target
Key Metrics Exposed
$ curl http://localhost:9100/metrics | head -50
# βββ CPU βββ
node_cpu_seconds_total{cpu="0", mode="idle"} 78345.23
node_cpu_seconds_total{cpu="0", mode="user"} 12345.67
node_cpu_seconds_total{cpu="0", mode="system"} 5678.90
node_cpu_seconds_total{cpu="0", mode="iowait"} 234.56
# Modes: idle, user, system, iowait, nice, irq, softirq, steal
# βββ MEMORY βββ
node_memory_MemTotal_bytes 16777216000 # 16 GB total
node_memory_MemFree_bytes 2147483648 # 2 GB free
node_memory_MemAvailable_bytes 8589934592 # 8 GB available
node_memory_Buffers_bytes 536870912
node_memory_Cached_bytes 4294967296
node_memory_SwapTotal_bytes 2147483648
node_memory_SwapFree_bytes 1073741824
# βββ DISK βββ
node_filesystem_size_bytes{mountpoint="/", fstype="ext4"} 107374182400
node_filesystem_avail_bytes{mountpoint="/", fstype="ext4"} 53687091200
node_filesystem_free_bytes{mountpoint="/", fstype="ext4"} 59055800320
node_disk_read_bytes_total{device="sda"} 9876543210
node_disk_written_bytes_total{device="sda"} 5432109876
node_disk_io_time_seconds_total{device="sda"} 12345.67
# βββ NETWORK βββ
node_network_receive_bytes_total{device="eth0"} 9876543210
node_network_transmit_bytes_total{device="eth0"} 5432109876
node_network_receive_packets_total{device="eth0"} 12345678
node_network_receive_errs_total{device="eth0"} 0
node_network_transmit_errs_total{device="eth0"} 0
# βββ LOAD AVERAGE βββ
node_load1 1.23
node_load5 0.98
node_load15 0.87
# βββ SYSTEM INFO βββ
node_os_info{name="Ubuntu", version="22.04", id="ubuntu"} 1
node_uname_info{sysname="Linux", release="5.15.0", machine="x86_64"} 1
node_time_seconds 1699000000.123
node_boot_time_seconds 1698000000
# βββ FILE DESCRIPTORS βββ
node_filefd_allocated 1024
node_filefd_maximum 65536
Common PromQL Queries with Node Exporter
# CPU Usage % (per instance)
100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# Memory Usage %
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100
# Disk Usage %
(1 - (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"})) * 100
# Network Receive Rate (bytes/sec)
rate(node_network_receive_bytes_total{device="eth0"}[5m])
# Disk I/O Utilization
rate(node_disk_io_time_seconds_total{device="sda"}[5m]) * 100
# Load Average vs CPU Count
node_load1 / count(node_cpu_seconds_total{mode="idle"}) by (instance)
Collectors (Modules)
Node Exporter uses "collectors" to gather different types of metrics.
You can enable/disable collectors:
Enabled by default:
β cpu, meminfo, diskstats, filesystem, netdev, loadavg,
time, uname, filefd, stat, vmstat, etc.
Disabled by default (enable with --collector.<name>):
β systemd, processes, tcpstat, ntp, wifi, etc.
Example:
./node_exporter --collector.systemd --collector.processes
./node_exporter --no-collector.wifi # Disable a collector
4.13 π BLACKBOX EXPORTER (Deep Dive)
What is Blackbox Exporter?
The Blackbox Exporter performs black-box probing of endpoints over HTTP, HTTPS, DNS, TCP, and ICMP. It monitors from the outside, like a real user would.
Key Facts
Port: 9115 (default)
Endpoint: /probe (NOT /metrics for probing!)
Maintainer: Prometheus Team (official)
Protocols: HTTP, HTTPS, DNS, TCP, ICMP
How It Works (Different from other exporters!)
Normal Exporter:
Prometheus β scrapes /metrics β Exporter β queries system
Blackbox Exporter:
Prometheus β scrapes /probe?target=URL&module=http_2xx
β Blackbox Exporter β probes the target URL
β Returns probe results as metrics
The TARGET is passed as a URL parameter!
Configuration (blackbox.yml)
modules:
http_2xx:
prober: http
timeout: 5s
http:
valid_http_versions: ["HTTP/1.1", "HTTP/2.0"]
valid_status_codes: [200, 201, 204, 301, 302]
method: GET
follow_redirects: true
preferred_ip_protocol: "ip4"
http_post_2xx:
prober: http
http:
method: POST
valid_status_codes: [200, 201]
tcp_connect:
prober: tcp
timeout: 5s
icmp_check:
prober: icmp
timeout: 5s
icmp:
preferred_ip_protocol: "ip4"
dns_udp:
prober: dns
dns:
query_name: "example.com"
query_type: "A"
transport_protocol: "udp"
Prometheus Configuration for Blackbox
# prometheus.yml
scrape_configs:
- job_name: 'blackbox-http'
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets:
- https://example.com
- https://prometheus.io
- https://grafana.com
relabel_configs:
# Pass the target URL as a parameter
- source_labels: [__address__]
target_label: __param_target
# Set the instance label to the target URL
- source_labels: [__param_target]
target_label: instance
# Point to the actual Blackbox Exporter
- target_label: __address__
replacement: blackbox-exporter:9115
Key Metrics Exposed
# Was the probe successful? (1 = yes, 0 = no)
probe_success{instance="https://example.com"} 1
# HTTP status code received
probe_http_status_code{instance="https://example.com"} 200
# Total probe duration
probe_duration_seconds{instance="https://example.com"} 0.234
# DNS resolution time
probe_dns_lookup_time_seconds{instance="https://example.com"} 0.012
# TCP connection time
probe_http_duration_seconds{phase="connect", instance="https://example.com"} 0.045
# TLS handshake time
probe_http_duration_seconds{phase="tls", instance="https://example.com"} 0.089
# Response time
probe_http_duration_seconds{phase="resolve", instance="https://example.com"} 0.012
# SSL certificate expiry (Unix timestamp)
probe_ssl_earliest_cert_expiry{instance="https://example.com"} 1735689600
# SSL certificate is valid?
probe_ssl_last_chain_expiry_timestamp_seconds 1735689600
# HTTP version
probe_http_version{instance="https://example.com"} 2
# Content length
probe_http_content_length{instance="https://example.com"} 1256
Common PromQL Queries
# Which targets are down?
probe_success == 0
# SSL certificate expiring in less than 30 days?
(probe_ssl_earliest_cert_expiry - time()) / 86400 < 30
# Slow probes (> 2 seconds)
probe_duration_seconds > 2
# HTTP status code not 200
probe_http_status_code != 200
4.14 π cADVISOR (Container Metrics)
What is cAdvisor?
cAdvisor (Container Advisor) provides container-level resource usage and performance metrics. Itβs built into the Kubelet in Kubernetes.
Key Facts
Port: 8080 (standalone) or 10250 (Kubelet)
Endpoint: /metrics
Maintainer: Google
Use case: Docker/Kubernetes container metrics
Key Metrics
# Container CPU usage
container_cpu_usage_seconds_total{container="web-app", pod="web-abc123"}
# Container memory usage
container_memory_usage_bytes{container="web-app"}
container_memory_working_set_bytes{container="web-app"}
# Container network
container_network_receive_bytes_total{container="web-app"}
container_network_transmit_bytes_total{container="web-app"}
# Container filesystem
container_fs_usage_bytes{container="web-app"}
container_fs_limit_bytes{container="web-app"}
# Container restarts
container_last_seen{container="web-app"}
In Kubernetes
In Kubernetes, cAdvisor is embedded in the Kubelet.
You don't need to deploy it separately.
Metrics are available at:
https://<node-ip>:10250/metrics/cadvisor
Prometheus typically scrapes these via the kubernetes_sd_configs
with role: node and the appropriate relabeling.
4.15 π DATABASE EXPORTERS
MySQL Exporter
Port: 9104
Monitors: MySQL/MariaDB performance metrics
Key metrics:
mysql_global_status_queries_total # Total queries
mysql_global_status_threads_connected # Active connections
mysql_global_status_slow_queries_total # Slow queries
mysql_global_status_connections_total # Total connections
mysql_global_status_bytes_received_total # Bytes received
mysql_global_status_bytes_sent_total # Bytes sent
mysql_global_variables_max_connections # Max connections config
PostgreSQL Exporter
Port: 9187
Monitors: PostgreSQL performance metrics
Key metrics:
pg_stat_activity_count # Active connections
pg_stat_database_tup_fetched # Rows fetched
pg_stat_database_tup_inserted # Rows inserted
pg_stat_database_numbackends # Backend connections
pg_database_size_bytes # Database size
pg_stat_bgwriter_buffers_checkpoint # Checkpoint buffers
Redis Exporter
Port: 9121
Monitors: Redis performance metrics
Key metrics:
redis_connected_clients # Connected clients
redis_used_memory_bytes # Memory usage
redis_commands_processed_total # Total commands
redis_keyspace_hits_total # Cache hits
redis_keyspace_misses_total # Cache misses
redis_connected_slaves # Replication slaves
redis_blocked_clients # Blocked clients
4.16 π OTHER COMMON EXPORTERS
| Exporter | Port | Monitors |
|---|---|---|
| SNMP Exporter | 9116 | Network devices (routers, switches) via SNMP |
| JMX Exporter | varies | Java applications via JMX |
| Kafka Exporter | 9308 | Apache Kafka topics, consumer groups |
| Elasticsearch Exporter | 9114 | Elasticsearch cluster health, indices |
| MongoDB Exporter | 9216 | MongoDB performance, replication |
| Nginx Exporter | 9113 | Nginx request rates, connections |
| HAProxy Exporter | 9101 | HAProxy backend/frontend stats |
| Consul Exporter | 9107 | HashiCorp Consul health, services |
| CloudWatch Exporter | 9106 | AWS CloudWatch metrics |
| Stackdriver Exporter | 9255 | GCP Stackdriver metrics |
| Azure Exporter | 9276 | Azure Monitor metrics |
| StatsD Exporter | 9125 | StatsD to Prometheus bridge |
4.17 π CUSTOM EXPORTERS
When to Build a Custom Exporter
Build a custom exporter when:
β
No existing exporter for your system
β
You need metrics from a proprietary API
β
You need to aggregate data from multiple sources
β
You need custom business logic in metric collection
Don't build when:
β An existing exporter already covers your needs
β You can instrument the application directly with a client library
β The data is better suited for logs or traces
Custom Exporter Architecture
βββββββββββββββββββββββββββββββββββββββββββββββ
β CUSTOM EXPORTER β
β β
β 1. Connect to target system β
β (API, database, file, etc.) β
β β
β 2. Collect raw data β
β (query API, parse logs, etc.) β
β β
β 3. Convert to Prometheus metrics β
β (Counter, Gauge, Histogram) β
β β
β 4. Expose /metrics endpoint β
β (HTTP server on a port) β
β β
β 5. Prometheus scrapes the endpoint β
βββββββββββββββββββββββββββββββββββββββββββββββ
Simple Custom Exporter Example (Python)
from prometheus_client import start_http_server, Gauge
import requests
import time
# Define metrics
API_RESPONSE_TIME = Gauge(
'external_api_response_time_seconds',
'Response time of external API',
['endpoint']
)
API_STATUS = Gauge(
'external_api_up',
'Whether the external API is reachable',
['endpoint']
)
def collect_metrics():
"""Collect metrics from external API"""
endpoints = ['/users', '/orders', '/products']
for endpoint in endpoints:
try:
start = time.time()
response = requests.get(f'https://api.example.com{endpoint}')
duration = time.time() - start
API_RESPONSE_TIME.labels(endpoint=endpoint).set(duration)
API_STATUS.labels(endpoint=endpoint).set(
1 if response.status_code == 200 else 0
)
except Exception:
API_STATUS.labels(endpoint=endpoint).set(0)
if __name__ == '__main__':
start_http_server(9200)
print("Custom exporter running on port 9200")
while True:
collect_metrics()
time.sleep(30) # Collect every 30 seconds
4.18 π PUSHGATEWAY β DETAILED USAGE
When to Use Pushgateway
β
USE for:
β Cron jobs that run for seconds (shorter than scrape interval)
β CI/CD pipeline stages
β Serverless functions (AWS Lambda, etc.)
β Batch processing jobs
β Any job that may not be alive when Prometheus scrapes
β DON'T USE for:
β Long-running services (use direct scraping!)
β Replacing the pull model entirely
β Pushing metrics from behind firewalls (use remote_write)
β High-frequency metrics (use client library + scraping)
How It Works
Flow:
1. Batch job runs and completes
2. Job pushes metrics to Pushgateway via HTTP POST/PUT
3. Pushgateway stores the metrics
4. Prometheus scrapes Pushgateway at regular intervals
5. Pushgateway metrics appear in Prometheus with original labels
β οΈ Important: Pushgateway remembers the LAST pushed value forever!
If a job stops pushing, the old value persists.
You need to explicitly DELETE metrics when a job is done.
Pushing Metrics (Examples)
# Push a single metric (bash)
echo "backup_duration_seconds 45.2" | \
curl --data-binary @- http://pushgateway:9091/metrics/job/backup
# Push with instance label
echo "backup_duration_seconds 45.2" | \
curl --data-binary @- http://pushgateway:9091/metrics/job/backup/instance/server1
# Push multiple metrics
cat <<EOF | curl --data-binary @- http://pushgateway:9091/metrics/job/backup
# TYPE backup_duration_seconds gauge
backup_duration_seconds 45.2
# TYPE backup_files_total counter
backup_files_total 1234
# TYPE backup_size_bytes gauge
backup_size_bytes 5368709120
EOF
# DELETE metrics when job is done
curl -X DELETE http://pushgateway:9091/metrics/job/backup/instance/server1
Python Pushgateway Example
from prometheus_client import CollectorRegistry, Gauge, push_to_gateway
registry = CollectorRegistry()
duration = Gauge(
'batch_job_duration_seconds',
'Duration of batch job',
registry=registry
)
records = Gauge(
'batch_job_records_processed',
'Number of records processed',
registry=registry
)
# Do the batch work
duration.set(45.2)
records.set(10000)
# Push to Pushgateway
push_to_gateway(
'pushgateway:9091',
job='daily_etl',
registry=registry,
grouping_key={'instance': 'etl-server-1'}
)
Prometheus Configuration for Pushgateway
scrape_configs:
- job_name: 'pushgateway'
honor_labels: true # β οΈ CRITICAL! Preserves job/instance from push
static_configs:
- targets: ['pushgateway:9091']
π‘ Key Takeaway for Exam
Pushgateway is ONLY for short-lived batch jobs honor_labels: true is required in Prometheus config Pushgateway remembers last value until explicitly deleted Push via HTTP POST/PUT to
/metrics/job/<jobname>Delete via HTTP DELETE when job completes
4.19 π EXPOSITION FORMAT DEEP DIVE
Complete Format Specification
# Lines starting with # are comments
# Special comment lines: HELP and TYPE
# HELP <metric_name> <description>
# TYPE <metric_name> <type>
# <metric_name>{<label_name>="<label_value>",...} <value> [<timestamp_ms>]
# βββ Full Example βββ
# HELP http_requests_total The total number of HTTP requests.
# TYPE http_requests_total counter
http_requests_total{method="post",code="200"} 1027 1395066363000
http_requests_total{method="post",code="400"} 3 1395066363000
# HELP http_request_duration_seconds A histogram of request duration.
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{le="0.05"} 24054
http_request_duration_seconds_bucket{le="0.1"} 33444
http_request_duration_seconds_bucket{le="0.2"} 100392
http_request_duration_seconds_bucket{le="0.5"} 129389
http_request_duration_seconds_bucket{le="1"} 133988
http_request_duration_seconds_bucket{le="+Inf"} 144320
http_request_duration_seconds_sum 53423
http_request_duration_seconds_count 144320
# HELP temperature_celsius Current temperature.
# TYPE temperature_celsius gauge
temperature_celsius{location="outside"} 27.3
temperature_celsius{location="inside"} 22.1
Format Rules (Exam Critical!)
1. Encoding: UTF-8
2. Content-Type: text/plain; version=0.0.4; charset=utf-8
3. Lines ending: \n (Unix-style)
4. Metric names: [a-zA-Z_:][a-zA-Z0-9_:]*
5. Label names: [a-zA-Z_][a-zA-Z0-9_]*
6. Label values: any UTF-8 string, must be in double quotes
7. Special characters in label values must be escaped:
β \" for double quote
β \\ for backslash
β \n for newline
8. Value: float64 (can be NaN, +Inf, -Inf)
9. Timestamp: optional, Unix epoch in milliseconds
10. Empty lines are allowed and ignored
11. HELP and TYPE lines are optional but recommended
12. TYPE must be one of: counter, gauge, histogram, summary, untyped
13. The _bucket, _sum, _count suffixes are auto-generated for histograms
4.20 π REAL-WORLD INSTRUMENTATION SCENARIOS
Scenario 1: E-Commerce Checkout Service
# What to instrument:
# 1. Request rate (Counter)
checkout_requests_total = Counter(
'checkout_requests_total',
'Total checkout requests',
['payment_method', 'status'] # visa, paypal, crypto | success, failed
)
# 2. Checkout duration (Histogram)
checkout_duration = Histogram(
'checkout_duration_seconds',
'Time to complete checkout',
['payment_method'],
buckets=[0.5, 1, 2, 5, 10, 30]
)
# 3. Cart value (Histogram)
cart_value = Histogram(
'checkout_cart_value_dollars',
'Value of cart at checkout',
buckets=[10, 25, 50, 100, 250, 500, 1000]
)
# 4. Active checkouts (Gauge)
active_checkouts = Gauge(
'active_checkouts',
'Currently in-progress checkouts'
)
# 5. Payment gateway latency (Histogram)
payment_gateway_latency = Histogram(
'payment_gateway_duration_seconds',
'Payment gateway response time',
['gateway'] # stripe, paypal
)
Scenario 2: Batch Data Pipeline
#!/bin/bash
# daily_etl.sh β Runs as a cron job at 2 AM
START=$(date +%s)
# ... run ETL pipeline ...
RECORDS_PROCESSED=1500000
ERRORS=23
END=$(date +%s)
DURATION=$((END - START))
# Push metrics to Pushgateway (job is short-lived!)
cat <<EOF | curl --data-binary @- http://pushgateway:9091/metrics/job/daily_etl/instance/etl-server-1
# TYPE etl_duration_seconds gauge
etl_duration_seconds $DURATION
# TYPE etl_records_processed gauge
etl_records_processed $RECORDS_PROCESSED
# TYPE etl_errors_total gauge
etl_errors_total $ERRORS
EOF
4.21 π EXAM-STYLE QUESTIONS (30 Questions)
Question 1
What is the primary purpose of instrumentation in the context of Prometheus?
A) To monitor network traffic between servers B) To add monitoring code to an application so it exposes internal metrics for Prometheus to scrape C) To configure Prometheus alerting rules D) To set up Grafana dashboards
β Answer & Explanation
Correct Answer: B
Instrumentation is the process of adding monitoring code (using a Prometheus client library) to your application so that it exposes metrics about its internal behavior via a /metrics endpoint. Prometheus then scrapes this endpoint to collect the metrics. Itβs a white-box monitoring approach that provides deep visibility into application internals.
Question 2
Which of the following is NOT an official Prometheus client library?
A) Go (client_golang)
B) Python (prometheus_client)
C) Java (simpleclient)
D) PHP (prometheus-php-official)
β Answer & Explanation
Correct Answer: D
The official Prometheus client libraries maintained by the Prometheus team are for Go, Python, Java, and Ruby. PHP has community-maintained libraries but no official one. Other community libraries exist for .NET, Node.js, Rust, C++, etc.
Question 3
What is the key difference between a client library and an exporter?
A) Client libraries are for infrastructure; exporters are for applications B) Client libraries run inside the application code; exporters are separate processes that monitor third-party systems C) Exporters use the push model; client libraries use the pull model D) There is no difference; they are the same thing
β Answer & Explanation
Correct Answer: B
A client library is embedded in your application code and instruments it from within (e.g., adding counters for HTTP requests). An exporter is a standalone process that collects metrics from a third-party system (e.g., MySQL, Linux kernel) and exposes them in Prometheus format. Both expose a /metrics endpoint for Prometheus to scrape.
Question 4
Which metric type should you use to track the current number of active user sessions?
A) Counter B) Gauge C) Histogram D) Summary
β Answer & Explanation
Correct Answer: B
Active user sessions is a value that goes up (user logs in) and down (user logs out). This is the classic use case for a Gauge. Counters only go up, and Histograms/Summaries are for distributions of observations (like latency).
Question 5
Why should you avoid using high-cardinality labels like user_id in Prometheus metrics?
A) Because Prometheus doesnβt support string label values B) Because each unique label combination creates a new time series, which can cause memory exhaustion C) Because high-cardinality labels slow down the scrape process D) Because Grafana cannot display high-cardinality labels
β Answer & Explanation
Correct Answer: B
In Prometheus, every unique combination of metric name + label values creates a separate time series. If you use user_id as a label with 1 million users, you create 1 million time series per metric. This consumes massive amounts of memory and can crash Prometheus. Labels should have a bounded, low-cardinality set of values (e.g., method, status, endpoint).
Question 6
What is the default port for the Node Exporter?
A) 9090 B) 9100 C) 9115 D) 8080
β Answer & Explanation
Correct Answer: B
The Node Exporter runs on port 9100 by default. Port 9090 is the Prometheus server default. Port 9115 is the Blackbox Exporter default. Port 8080 is commonly used by cAdvisor (standalone) or application servers.
Question 7
Which exporter would you use to monitor the availability and response time of external websites?
A) Node Exporter B) cAdvisor C) Blackbox Exporter D) MySQL Exporter
β Answer & Explanation
Correct Answer: C
The Blackbox Exporter is designed for black-box probing of endpoints over HTTP, HTTPS, DNS, TCP, and ICMP. It checks if websites are reachable, measures response times, validates SSL certificates, and checks HTTP status codes. Node Exporter monitors Linux system metrics, cAdvisor monitors containers, and MySQL Exporter monitors databases.
Question 8
What is the correct naming convention for a counter metric tracking total HTTP requests?
A) http_requests_counter
B) httpRequestsTotal
C) http_requests_total
D) http-requests-total
β Answer & Explanation
Correct Answer: C
Prometheus naming conventions require: snake_case (no camelCase or hyphens), the _total suffix for counters, and no metric type in the name. http_requests_total follows all rules. Option A includes βcounterβ in the name (redundant). Option B is camelCase. Option D uses hyphens (invalid).
Question 9
Which base unit should you use for measuring request duration in Prometheus?
A) Milliseconds B) Microseconds C) Seconds D) Minutes
β Answer & Explanation
Correct Answer: C
Prometheus convention requires base units. For time/duration, the base unit is seconds. The metric name should include the unit: http_request_duration_seconds. Even if your application measures in milliseconds internally, convert to seconds before exposing the metric. This ensures consistency across all Prometheus metrics and tools.
Question 10
What does the Blackbox Exporterβs /probe endpoint do?
A) Returns the Blackbox Exporterβs own health metrics B) Probes a target specified via URL parameters and returns the probe results as Prometheus metrics C) Lists all configured probe modules D) Returns the configuration of the Blackbox Exporter
β Answer & Explanation
Correct Answer: B
The Blackbox Exporter works differently from other exporters. Instead of exposing static metrics on /metrics, it probes targets on demand via the /probe endpoint. You pass the target and module as URL parameters: /probe?target=https://example.com&module=http_2xx. The exporter then performs the probe and returns the results (success, duration, status code, SSL expiry, etc.) as Prometheus metrics.
Question 11
Which of the following is a valid use case for the Pushgateway?
A) Monitoring a long-running web application B) Collecting metrics from a cron job that runs for 10 seconds every hour C) Replacing the pull model for all monitoring targets D) Pushing metrics from behind a corporate firewall
β Answer & Explanation
Correct Answer: B
The Pushgateway is specifically designed for short-lived batch jobs that may complete before Prometheus can scrape them. A cron job running for 10 seconds is a perfect use case β the job pushes its metrics to the Pushgateway before exiting, and Prometheus scrapes the Pushgateway later. Long-running services should be scraped directly (A). Pushgateway should not replace the pull model (C) or be used for firewall traversal (D β use remote_write).
Question 12
What configuration is essential when scraping the Pushgateway from Prometheus?
A) scrape_interval: 5s
B) honor_labels: true
C) scheme: https
D) metrics_path: /probe
β Answer & Explanation
Correct Answer: B
honor_labels: true is essential when scraping the Pushgateway. Without it, Prometheus would overwrite the job and instance labels from the pushed metrics with the Pushgatewayβs own job/instance labels. With honor_labels: true, the original labels from the batch job are preserved, allowing you to distinguish metrics from different jobs and instances.
Question 13
What does the node_cpu_seconds_total metric from Node Exporter represent?
A) Current CPU usage percentage B) Total CPU time consumed since boot, broken down by CPU and mode C) Number of CPU cores D) CPU temperature
β Answer & Explanation
Correct Answer: B
node_cpu_seconds_total is a counter that tracks the total CPU time (in seconds) consumed since the system booted, broken down by cpu (core number) and mode (idle, user, system, iowait, nice, irq, softirq, steal). To get CPU usage percentage, you need to calculate: 100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100).
Question 14
Which metric type is most appropriate for measuring the distribution of HTTP request latencies across multiple application instances?
A) Counter B) Gauge C) Histogram D) Summary
β Answer & Explanation
Correct Answer: C
A Histogram is the best choice for measuring latency distributions across multiple instances because histogram buckets are aggregatable. You can sum the bucket counts from all instances and then calculate quantiles using histogram_quantile(). A Summary pre-calculates quantiles on each instance, and you cannot meaningfully aggregate percentiles across instances (averaging p99 values is mathematically invalid).
Question 15
What is the purpose of the # HELP line in the Prometheus exposition format?
A) It defines the metric type B) It provides a human-readable description of the metric C) It specifies the scrape interval D) It defines the label names
β Answer & Explanation
Correct Answer: B
The # HELP line provides a human-readable description of the metric. For example: # HELP http_requests_total The total number of HTTP requests. Itβs optional but strongly recommended. The # TYPE line defines the metric type (counter, gauge, etc.). HELP and TYPE lines are comment lines in the exposition format.
Question 16
Which of the following metric names is INVALID according to Prometheus naming conventions?
A) process_resident_memory_bytes
B) http_request_duration_seconds
C) my-app_requests_total
D) node_cpu_seconds_total
β Answer & Explanation
Correct Answer: C
my-app_requests_total is invalid because it contains a hyphen (-). Prometheus metric names must match the regex [a-zA-Z_:][a-zA-Z0-9_:]*, which only allows letters, digits, underscores, and colons. The correct name would be my_app_requests_total (using an underscore instead of a hyphen).
Question 17
What does cAdvisor primarily monitor?
A) Network switch performance via SNMP B) Container resource usage (CPU, memory, network, disk) C) Database query performance D) SSL certificate expiry
β Answer & Explanation
Correct Answer: B
cAdvisor (Container Advisor) provides container-level resource usage and performance metrics including CPU usage, memory usage, network I/O, and filesystem usage for Docker/Kubernetes containers. In Kubernetes, cAdvisor is embedded in the Kubelet and doesnβt need separate deployment.
Question 18
Which of the following is a best practice for Prometheus instrumentation?
A) Pre-calculate rates in your application and expose them as gauges
B) Use user_id as a label for detailed per-user monitoring
C) Store raw counters and let Prometheus calculate rates at query time using rate()
D) Create a separate metric name for each HTTP endpoint
β Answer & Explanation
Correct Answer: C
The Prometheus best practice is to store raw counters and let Prometheus calculate rates at query time using rate(), increase(), etc. Pre-calculating rates (A) loses information and makes aggregation impossible. Using user_id as a label (B) causes high cardinality and memory issues. Creating separate metric names per endpoint (D) is an anti-pattern β use a single metric with an endpoint label instead.
Question 19
What is the default port for the Blackbox Exporter?
A) 9100 B) 9090 C) 9115 D) 9121
β Answer & Explanation
Correct Answer: C
The Blackbox Exporter runs on port 9115 by default. 9100 is Node Exporter, 9090 is Prometheus server, and 9121 is Redis Exporter.
Question 20
Which exporter would you use to monitor a MySQL database?
A) Node Exporter B) MySQL Exporter (mysqld_exporter) C) Blackbox Exporter D) cAdvisor
β Answer & Explanation
Correct Answer: B
The MySQL Exporter (officially mysqld_exporter) is the dedicated exporter for MySQL/MariaDB databases. It connects to MySQL and exposes metrics like query counts, connection counts, slow queries, replication status, InnoDB metrics, etc., on port 9104. Node Exporter monitors the OS, Blackbox probes endpoints, and cAdvisor monitors containers.
Question 21
What happens to Pushgateway metrics when the batch job that pushed them finishes?
A) They are automatically deleted after the scrape interval B) They persist indefinitely until explicitly deleted via the API C) They expire after 5 minutes D) They are converted to counters
β Answer & Explanation
Correct Answer: B
Pushgateway metrics persist indefinitely until explicitly deleted. This is a common gotcha β if a batch job stops running, its last pushed metrics will continue to appear in Prometheus as if they were current. You should explicitly delete metrics via curl -X DELETE http://pushgateway:9091/metrics/job/<jobname> when a job completes, or use the push_to_gateway function with appropriate settings.
Question 22
Which of the following is the correct content type for the Prometheus exposition format?
A) application/json
B) text/plain; version=0.0.4
C) application/xml
D) text/csv
β Answer & Explanation
Correct Answer: B
The Prometheus exposition format uses the content type text/plain; version=0.0.4; charset=utf-8. The newer OpenMetrics format uses application/openmetrics-text. JSON, XML, and CSV are not used for Prometheus metric exposition.
Question 23
What is the _total suffix used for in Prometheus metric naming?
A) It indicates a gauge metric B) It indicates a histogram metric C) It indicates a counter metric D) It indicates the sum of all metrics
β Answer & Explanation
Correct Answer: C
The _total suffix is the naming convention for counter metrics in Prometheus. It indicates that the metric represents a cumulative count that only increases (or resets on restart). Examples: http_requests_total, node_cpu_seconds_total, process_network_transmit_bytes_total. This convention helps users immediately identify the metric type.
Question 24
Which of the following labels would cause high cardinality issues in Prometheus?
A) method
π PCA EXAM β DOMAIN 4: INSTRUMENTATION & EXPORTERS (Continued)
Questions 24β30 + Summary Cheat Sheet
Question 24 (Complete Answer)
Which of the following labels would cause high cardinality issues in Prometheus?
A) method (GET, POST, PUT, DELETE)
B) status (200, 404, 500)
C) user_id (unique per user, millions of values)
D) endpoint (/api/users, /api/orders)
β Answer & Explanation
Correct Answer: C
user_id would cause high cardinality issues because each unique user creates a new time series. With millions of users, this could create millions of time series per metric, exhausting Prometheus memory. Labels should have a bounded, low-cardinality set of values. method (4-5 values), status (limited HTTP codes), and endpoint (bounded set of routes) are all acceptable low-cardinality labels.
Question 25
What is the purpose of the # TYPE line in the Prometheus exposition format?
A) It specifies the data type of the label values B) It declares the metric type (counter, gauge, histogram, summary, or untyped) C) It defines the scrape interval for the metric D) It specifies the unit of measurement
β Answer & Explanation
Correct Answer: B
The # TYPE line declares the metric type so Prometheus knows how to handle the metric. Valid types are: counter, gauge, histogram, summary, and untyped. For example: # TYPE http_requests_total counter. While optional, itβs strongly recommended because it helps Prometheus validate the data and enables proper handling (e.g., knowing a counter can be used with rate()).
Question 26
Which of the following is NOT a valid Prometheus metric type in the exposition format?
A) counter B) gauge C) histogram D) timer
β Answer & Explanation
Correct Answer: D
The valid Prometheus metric types in the exposition format are: counter, gauge, histogram, summary, and untyped. βtimerβ is NOT a valid Prometheus type. Some other monitoring systems (StatsD, Dropwizard) have a timer type, but in Prometheus, you would use a Histogram or Summary to track durations/timers.
Question 27
What is the recommended approach for monitoring a third-party system like Redis that doesnβt natively expose Prometheus metrics?
A) Instrument the Redis source code with a Prometheus client library B) Use a dedicated Redis Exporter that queries Redis and exposes metrics in Prometheus format C) Use the Blackbox Exporter to probe Redis D) Use the Pushgateway to push Redis metrics
β Answer & Explanation
Correct Answer: B
For third-party systems you donβt control (like Redis, MySQL, Linux), the recommended approach is to use a dedicated exporter. The Redis Exporter connects to Redis, queries its INFO command and other stats, translates the data into Prometheus format, and exposes it on port 9121 for Prometheus to scrape. You canβt modify Redis source code (A), Blackbox only does external probing (C), and Pushgateway is for batch jobs (D).
Question 28
Which Node Exporter metric would you use to calculate disk usage percentage?
A) node_disk_read_bytes_total and node_disk_written_bytes_total
B) node_filesystem_size_bytes and node_filesystem_avail_bytes
C) node_disk_io_time_seconds_total
D) node_memory_MemTotal_bytes and node_memory_MemAvailable_bytes
β Answer & Explanation
Correct Answer: B
Disk usage percentage is calculated using node_filesystem_size_bytes (total disk size) and node_filesystem_avail_bytes (available space):
(1 - (node_filesystem_avail_bytes / node_filesystem_size_bytes)) * 100
Option A gives disk I/O throughput. Option C gives disk I/O utilization time. Option D is for memory, not disk.
Question 29
What is the key difference between white-box and black-box monitoring in the context of Prometheus instrumentation?
A) White-box uses counters; black-box uses gauges B) White-box monitors from inside the application using client libraries; black-box monitors from outside using probes like the Blackbox Exporter C) White-box is for Linux servers; black-box is for cloud services D) White-box uses the push model; black-box uses the pull model
β Answer & Explanation
Correct Answer: B
White-box monitoring involves instrumenting the application from within using client libraries, exposing internal metrics like request rates, error rates, queue depths, and connection pool usage. Black-box monitoring observes the system from the outside, like a real user, using tools like the Blackbox Exporter to probe endpoints for availability, response time, and SSL certificate validity. Both approaches are complementary and should be used together.
Question 30
Which of the following statements about Prometheus client libraries is TRUE?
A) Client libraries only support the Counter metric type
B) Client libraries automatically expose a /metrics endpoint and collect default process/runtime metrics
C) Client libraries use the push model to send metrics to Prometheus
D) Client libraries can only be used with Go applications
β Answer & Explanation
Correct Answer: B
Prometheus client libraries automatically:
- Expose a
/metricsHTTP endpoint in the correct exposition format - Collect default process metrics (CPU, memory, file descriptors)
- Collect runtime metrics (GC pauses, goroutines/threads, heap usage)
They support ALL four metric types (Counter, Gauge, Histogram, Summary) β not just counters (A). They work with the pull model (Prometheus scrapes the /metrics endpoint), not push (C). Official libraries exist for Go, Python, Java, and Ruby, with community libraries for many more languages (D is wrong).
β DOMAIN 4 REVISION SUMMARY CHEAT SHEET
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DOMAIN 4: INSTRUMENTATION & EXPORTERS CHEAT SHEET β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β INSTRUMENTATION = Adding monitoring code to expose internal metrics β
β Client Library = Inside app code (Go, Python, Java, Ruby) β
β Exporter = Separate process for 3rd-party systems β
β β
β CLIENT LIBRARIES: β
β Official: Go (most mature), Python, Java, Ruby β
β Community: .NET, Node.js, Rust, C++, PHP β
β Auto-provide: /metrics endpoint + process/runtime metrics β
β Support: Counter, Gauge, Histogram, Summary β
β β
β NAMING CONVENTIONS: β
β β
snake_case (no hyphens, no camelCase) β
β β
Base units: seconds, bytes, celsius, ratio β
β β
_total suffix for counters β
β β
Include unit in name: _seconds, _bytes β
β β No type in name (not _counter, not _gauge) β
β β No colons (:) in raw metrics (reserved for recording rules) β
β β No __ prefix labels (reserved for internal use) β
β β
β BEST PRACTICES: β
β β
Low cardinality labels (method, status, endpoint) β
β β High cardinality labels (user_id, session_id, request_id) β
β β
Store raw counters, use rate() in PromQL β
β β Don't pre-calculate rates in application β
β β
Always provide HELP strings β
β β
Histogram > Summary for distributed systems β
β β
β METRIC TYPE DECISION: β
β Counting events? β Counter (_total suffix) β
β Current state? β Gauge β
β Distribution (multi)? β Histogram (aggregatable β
) β
β Distribution (single)? β Summary (not aggregatable β) β
β β
β COMMON EXPORTERS: β
β ββββββββββββββββββββ¬βββββββ¬ββββββββββββββββββββββββββββββ β
β β Node Exporter β 9100 β Linux CPU, mem, disk, net β β
β β Blackbox Exporterβ 9115 β HTTP/TCP/ICMP/DNS probing β β
β β cAdvisor β 8080 β Container metrics β β
β β MySQL Exporter β 9104 β MySQL/MariaDB β β
β β PostgreSQL β 9187 β PostgreSQL β β
β β Redis Exporter β 9121 β Redis β β
β β SNMP Exporter β 9116 β Network devices β β
β β JMX Exporter β var β Java JMX β β
β β Kafka Exporter β 9308 β Apache Kafka β β
β ββββββββββββββββββββ΄βββββββ΄ββββββββββββββββββββββββββββββ β
β β
β NODE EXPORTER: β
β Port 9100, Linux system metrics β
β Key metrics: node_cpu_seconds_total, node_memory_*_bytes, β
β node_filesystem_*_bytes, node_network_*_bytes_total, node_load* β
β Collectors can be enabled/disabled via flags β
β β
β BLACKBOX EXPORTER: β
β Port 9115, /probe endpoint (NOT /metrics!) β
β Probes: HTTP, HTTPS, TCP, ICMP, DNS β
β Target passed as URL param: /probe?target=URL&module=http_2xx β
β Key metrics: probe_success, probe_duration_seconds, β
β probe_http_status_code, probe_ssl_earliest_cert_expiry β
β Requires relabeling in prometheus.yml β
β β
β PUSHGATEWAY: β
β Port 9091, ONLY for short-lived batch jobs β
β Push via HTTP POST/PUT to /metrics/job/<name> β
β Delete via HTTP DELETE when done β
β β οΈ Metrics persist until explicitly deleted! β
β β οΈ honor_labels: true required in Prometheus config β
β β NOT for long-running services β
β β NOT for firewall traversal (use remote_write) β
β β
β EXPOSITION FORMAT: β
β Content-Type: text/plain; version=0.0.4 β
β # HELP <name> <description> β
β # TYPE <name> <counter|gauge|histogram|summary|untyped> β
β <name>{labels} <value> [<timestamp>] β
β UTF-8, label values in double quotes, escape \" \\ \n β
β β
β INFO METRICS: gauge = 1 with metadata labels β
β Example: app_build_info{version="1.2", commit="abc"} 1 β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
π― NEXT STEPS
Youβve now completed Domain 4 (16%)! Combined with Domains 1-3, youβve covered 82% of the exam! π
Reply with βDomain 4 Completeβ and Iβll provide the guides for the remaining three domains:
Domain 5: Dashboarding & Visualization (8%)
- Grafana fundamentals, data sources, panel types, variables, templating
Domain 6: Service Discovery (6%)
- Static, File, DNS, Kubernetes, Consul, EC2 SD + Relabeling deep dive
Domain 7: Alerting & Alertmanager (4%)
- Alerting rules, Alertmanager config, routing, grouping, silencing, inhibition
These three domains together are only 18%, so I can cover them more concisely or combine them into one comprehensive guide β your choice! π