PCA Exam Guide

πŸ“˜ PCA EXAM β€” DOMAIN 1: OBSERVABILITY CONCEPTS (18%)

Complete One-Stop Guide: Theory β†’ Examples β†’ Exam Questions


TABLE OF CONTENTS

1.1  Monitoring vs Observability
1.2  The Three Pillars of Observability (Metrics, Logs, Traces)
1.3  White-Box vs Black-Box Monitoring
1.4  Push vs Pull Model
1.5  Why Prometheus Uses Pull Model
1.6  SLI, SLO, SLA (The Reliability Trio)
1.7  Golden Signals (Google SRE Approach)
1.8  RED Method (Tom Wilkie)
1.9  USE Method (Brendan Gregg)
1.10 Comparison: Golden Signals vs RED vs USE
1.11 Real-World Scenarios
1.12 EXAM-STYLE QUESTIONS (30 Questions with Answers)


1.1 πŸ“– MONITORING vs OBSERVABILITY

What is Monitoring?

Monitoring is the process of collecting, aggregating, and analyzing predefined metrics and logs to track the health and performance of a system.

Think of it as: A car dashboard showing speed, fuel level, engine temperature. You know WHAT is happening.

Key Characteristics:

  • Reactive in nature
  • Focuses on known unknowns (things you already know can go wrong)
  • Uses predefined dashboards and alerts
  • Answers: β€œIs the system healthy?”

What is Observability?

Observability is a measure of how well you can understand the internal state of a system from its external outputs. It allows you to explore and investigate why something is happening, even for problems you’ve never seen before.

Think of it as: Having a full diagnostic computer connected to your car that can tell you exactly which cylinder is misfiring and why. You know WHY it’s happening.

Key Characteristics:

  • Proactive and exploratory
  • Focuses on unknown unknowns (things you didn’t anticipate)
  • Uses high-cardinality data and ad-hoc queries
  • Answers: β€œWhy is the system behaving this way?”

The Key Difference (Exam Favorite!)

Aspect Monitoring Observability
Question it answers Is something wrong? WHY is something wrong?
Approach Reactive Proactive & Exploratory
Data Predefined metrics High-cardinality, contextual
Scope Known unknowns Unknown unknowns
Action Alert β†’ Fix Investigate β†’ Understand β†’ Fix
Analogy Smoke detector Fire investigator

πŸ” Real-World Example

Scenario: Your e-commerce website is slow.

Monitoring tells you:

β€œCPU usage on server-3 is at 95%. Response time is above 2 seconds.” βœ… You know SOMETHING is wrong.

Observability tells you:

β€œCPU spike on server-3 is caused by a specific database query from the checkout microservice, triggered only for users in the EU region using the new payment gateway, because of a missing database index on the transactions table.” βœ… You know EXACTLY WHY it’s wrong and how to fix it.

πŸ’‘ Key Takeaway for Exam

Monitoring is a SUBSET of Observability. You can have monitoring without observability, but you CANNOT have observability without monitoring. Observability is the evolution of monitoring.



1.2 πŸ“– THE THREE PILLARS OF OBSERVABILITY

The three pillars are the foundational data types that together provide full observability.

Pillar 1: METRICS πŸ“Š

Definition: Numerical measurements of system behavior over time. Metrics are aggregated, structured, and time-series data.

Characteristics:

  • Lightweight and cheap to store
  • Great for dashboards and alerting
  • Pre-aggregated (you lose individual event details)
  • Best for: β€œHow is the system performing overall?”

Examples:

# CPU Usage (Gauge)
node_cpu_seconds_total{cpu="0", mode="idle"} 78345.23

# HTTP Requests (Counter)
http_requests_total{method="GET", status="200"} 15432

# Request Duration (Histogram)
http_request_duration_seconds_bucket{le="0.1"} 500
http_request_duration_seconds_bucket{le="0.5"} 800
http_request_duration_seconds_bucket{le="1.0"} 950
http_request_duration_seconds_bucket{le="+Inf"} 1000

# Memory Usage (Gauge)
process_resident_memory_bytes 256000000

When to use Metrics:

  • Tracking trends over time
  • Setting up alerts (e.g., CPU > 90%)
  • Building dashboards
  • Capacity planning

Limitation: Metrics are aggregated. You know 1000 requests failed, but you don’t know WHICH specific requests or WHY.


Pillar 2: LOGS πŸ“

Definition: Discrete, timestamped records of events that happened in a system. Logs are detailed, unstructured or semi-structured text.

Characteristics:

  • Rich in detail and context
  • Can be structured (JSON) or unstructured (plain text)
  • Expensive to store and search at scale
  • Best for: β€œWhat exactly happened at this specific moment?”

Examples:

# Unstructured Log
[2024-11-15 10:23:45] ERROR Connection to database failed: 
timeout after 30s host=db-primary.internal:5432

# Structured Log (JSON)
{
  "timestamp": "2024-11-15T10:23:45.123Z",
  "level": "ERROR",
  "service": "payment-service",
  "trace_id": "abc123def456",
  "user_id": "user_789",
  "message": "Payment processing failed",
  "error": "Insufficient funds",
  "duration_ms": 234
}

When to use Logs:

  • Debugging specific errors
  • Auditing and compliance
  • Post-incident investigation
  • Understanding exact sequence of events

Limitation: Logs generate massive volume. Searching through billions of log lines is slow and expensive.


Pillar 3: TRACES πŸ”—

Definition: Traces represent the end-to-end journey of a single request as it flows through multiple services in a distributed system. A trace is made up of spans.

Characteristics:

  • Shows the full path of a request across microservices
  • Each span represents a single unit of work
  • Essential for distributed systems
  • Best for: β€œWhere is the bottleneck in this request’s journey?”

Key Terminology:

  • Trace: The entire journey of one request (has a unique Trace ID)
  • Span: A single operation within a trace (has a Span ID and Parent Span ID)
  • Context Propagation: Passing trace context between services

Example:

User clicks "Buy Now" β†’ Trace ID: xyz789

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Trace: xyz789 (Total: 850ms)                        β”‚
β”‚                                                     β”‚
β”‚ β”œβ”€ Span 1: API Gateway          (10ms)              β”‚
β”‚ β”‚   β”œβ”€ Span 2: Auth Service     (50ms)              β”‚
β”‚ β”‚   β”œβ”€ Span 3: Order Service    (200ms)             β”‚
β”‚ β”‚   β”‚   β”œβ”€ Span 4: Inventory    (150ms)  ← SLOW!    β”‚
β”‚ β”‚   β”‚   └─ Span 5: Payment      (400ms)  ← SLOWER!  β”‚
β”‚ β”‚   └─ Span 6: Notification     (40ms)              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

β†’ Trace reveals: Payment service is the bottleneck (400ms)

When to use Traces:

  • Debugging latency in microservices
  • Finding bottlenecks in distributed systems
  • Understanding service dependencies
  • Root cause analysis across services

Limitation: Traces are sampled (you can’t trace 100% of requests at scale due to cost). You might miss the specific failing request.


How the Three Pillars Work Together

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              OBSERVABILITY WORKFLOW                       β”‚
β”‚                                                          β”‚
β”‚  Step 1: METRICS alert you                               β”‚
β”‚          "Error rate spiked to 5% on payment-service"     β”‚
β”‚                          ↓                               β”‚
β”‚  Step 2: TRACES help you find                            β”‚
β”‚          "Requests to /api/pay are taking 3s,             β”‚
β”‚           bottleneck is the database call"                β”‚
β”‚                          ↓                               β”‚
β”‚  Step 3: LOGS give you the detail                        β”‚
β”‚          "Connection pool exhausted at 10:23:45,          β”‚
β”‚           max_connections=100, active=100"                β”‚
β”‚                          ↓                               β”‚
β”‚  Result: You fix the connection pool size!                β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ’‘ Key Takeaway for Exam

Metrics = WHAT is happening (numbers, trends, alerts) Logs = WHAT happened in detail (events, errors, context) Traces = WHERE it happened (request flow across services)

Prometheus is primarily a METRICS system. It does NOT handle logs or traces natively (though it can integrate with tools that do).



1.3 πŸ“– WHITE-BOX vs BLACK-BOX MONITORING

White-Box Monitoring (Inside-Out)

Definition: Monitoring the internal state of a system by instrumenting the application code and infrastructure from within.

How it works:

  • You add instrumentation code INSIDE your application
  • The application exposes metrics about its internal workings
  • You monitor things like: memory usage, queue depth, active connections, internal error rates

Example:

# Your application exposes these metrics internally:
http_requests_total{handler="/api/users", method="GET"} 5000
process_open_fds 256
go_goroutines 150
jvm_memory_used_bytes{area="heap"} 536870912
app_active_database_connections 45
app_queue_messages_pending 1200

Tools: Prometheus (with client libraries), Datadog APM, New Relic

Analogy: A doctor performing blood tests and X-rays to see what’s happening INSIDE your body.


Black-Box Monitoring (Outside-In)

Definition: Monitoring the external behavior of a system by probing it from the outside, just like a real user would.

How it works:

  • You send synthetic requests to your system from outside
  • You measure response time, availability, and correctness
  • You DON’T need access to internal code or metrics

Example:

# Blackbox Exporter probing your website:
probe_success{instance="https://myapp.com"} 1
probe_http_status_code{instance="https://myapp.com"} 200
probe_duration_seconds{instance="https://myapp.com"} 0.234
probe_http_ssl 1
probe_ssl_earliest_cert_expiry 1735689600

Tools: Prometheus Blackbox Exporter, Pingdom, UptimeRobot, Synthetic monitoring

Analogy: A patient checking their own temperature with a thermometer β€” they can tell if they have a fever, but not WHY.


Comparison Table

Aspect White-Box Black-Box
Perspective Internal (developer) External (user)
What it monitors Internal metrics, code paths Endpoints, availability
Instrumentation Required in code No code changes needed
Depth Deep, granular Surface-level
Use case Debugging, optimization Uptime, SLA verification
Example β€œDB connection pool at 90%” β€œWebsite returned 200 OK in 200ms”
Prometheus tool Client libraries, exporters Blackbox Exporter

πŸ’‘ Key Takeaway for Exam

Best practice: Use BOTH.

  • Black-box tells you IF there’s a problem (from user’s perspective)
  • White-box tells you WHY there’s a problem (from system’s perspective)

In Prometheus: White-box = application metrics via client libraries Black-box = Blackbox Exporter for probing endpoints



1.4 πŸ“– PUSH vs PULL MODEL

This is one of the most frequently tested topics in the PCA exam!

PUSH Model

Definition: The monitored targets actively send (push) their metrics to a central monitoring server at regular intervals.

How it works:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  App Server │──push──→│             │←──push──│  App Server β”‚
β”‚  (Agent)    β”‚  every  β”‚  Central    β”‚  every  β”‚  (Agent)    β”‚
β”‚             β”‚  10s    β”‚  Monitoring β”‚  10s    β”‚             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β”‚  Server     β”‚         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚             β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”‚  (e.g.,     β”‚
β”‚  Database   │──push──→│  Graphite,  β”‚
β”‚  (Agent)    β”‚  every  β”‚  InfluxDB,  β”‚
β”‚             β”‚  10s    β”‚  Datadog)   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Tools using Push: Graphite, InfluxDB (Telegraf), Datadog Agent, StatsD, Nagios (NRPE)

Advantages of Push:

  • βœ… Works well behind firewalls (agent initiates outbound connection)
  • βœ… Server doesn’t need to know about targets in advance
  • βœ… Good for short-lived jobs (they push before dying)
  • βœ… Easier to scale horizontally (just add more agents)

Disadvantages of Push:

  • ❌ Hard to detect if a target dies (it just stops pushing β€” is it dead or just slow?)
  • ❌ Configuration is distributed (each agent needs to know where to push)
  • ❌ Can overwhelm the server if too many agents push simultaneously
  • ❌ Harder to debug (is the agent running? is the network blocking?)
  • ❌ No single source of truth for what should be monitored

PULL Model

Definition: The central monitoring server actively fetches (pulls/scrapes) metrics from targets at regular intervals.

How it works:

                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                        β”‚  Prometheus β”‚
                        β”‚  Server     β”‚
                        β”‚             β”‚
                        β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚          β”‚          β”‚
                scrape      scrape     scrape
                every       every      every
                15s         15s        15s
                    β”‚          β”‚          β”‚
                    β–Ό          β–Ό          β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚App Serverβ”‚ β”‚ Database β”‚ β”‚  Cache   β”‚
              β”‚ :9090    β”‚ β”‚ Exporter β”‚ β”‚ Exporter β”‚
              β”‚ /metrics β”‚ β”‚ :9104    β”‚ β”‚ :9121    β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ /metrics β”‚ β”‚ /metrics β”‚
                           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Tools using Pull: Prometheus, Nagios (partially)

Advantages of Pull:

  • βœ… Easy to detect if a target is down (scrape fails = target is dead)
  • βœ… Centralized configuration (prometheus.yml defines all targets)
  • βœ… Easier debugging (just curl the /metrics endpoint yourself!)
  • βœ… Targets don’t need to know about the monitoring server
  • βœ… Natural rate limiting (server controls scrape frequency)
  • βœ… Health check is built-in (if Prometheus can’t scrape, something is wrong)

Disadvantages of Pull:

  • ❌ Targets must be reachable from the server (firewall issues)
  • ❌ Server needs to know all targets (solved by Service Discovery)
  • ❌ Not ideal for short-lived batch jobs (solved by Pushgateway)
  • ❌ Can be harder to scale to millions of targets

Detailed Comparison Table

Aspect Push Model Pull Model
Who initiates? Target (agent) Monitoring server
Direction Target β†’ Server Server β†’ Target
Dead target detection ❌ Difficult (timeout-based) βœ… Easy (scrape failure)
Configuration Distributed (on each agent) Centralized (on server)
Debugging Hard (check agent logs) Easy (curl /metrics)
Firewall βœ… Easier (outbound) ❌ Harder (inbound needed)
Short-lived jobs βœ… Natural fit ❌ Needs Pushgateway
Rate control Agent decides Server decides
Examples Graphite, Datadog, StatsD Prometheus

πŸ” Real-World Example

Push Model Scenario (Datadog):

Your app server has a Datadog agent installed.
Every 10 seconds, the agent collects CPU, memory, and custom metrics
and PUSHES them to Datadog's cloud servers via HTTPS.

If the app server crashes, the agent stops pushing.
Datadog notices "no data received for 2 minutes" β†’ alerts.
But was it a crash? Network issue? Agent bug? Hard to tell immediately.

Pull Model Scenario (Prometheus):

Prometheus is configured to scrape your app server at http://app:8080/metrics
every 15 seconds.

If the app server crashes, the next scrape FAILS immediately.
Prometheus marks the target as DOWN and fires an alert within 15 seconds.
You can also try `curl http://app:8080/metrics` yourself to debug.


1.5 πŸ“– WHY PROMETHEUS USES THE PULL MODEL

This is a philosophical design decision by the Prometheus creators (at SoundCloud). Here are the key reasons:

Reason 1: Reliability & Health Detection

In a pull model, if Prometheus can't scrape a target, it IMMEDIATELY knows:
  β†’ The target might be down
  β†’ The network might be broken
  β†’ The target is overloaded and can't respond

In a push model, if data stops arriving:
  β†’ Is the target dead?
  β†’ Is the agent crashed?
  β†’ Is the network partitioned?
  β†’ Is the agent misconfigured?
  β†’ You can't easily tell!

Reason 2: Easier Development & Debugging

# During development, you can simply run your app and check metrics:
$ curl http://localhost:8080/metrics

# Output:
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",status="200"} 42
http_requests_total{method="POST",status="500"} 3

# No need to set up a monitoring server to see your metrics!
# This makes local development and debugging trivially easy.

Reason 3: Centralized Configuration

# prometheus.yml - Everything in ONE place
scrape_configs:
  - job_name: 'web-app'
    scrape_interval: 15s
    static_configs:
      - targets: ['app1:8080', 'app2:8080', 'app3:8080']
      
  - job_name: 'database'
    scrape_interval: 30s
    static_configs:
      - targets: ['db-exporter:9104']

You know exactly what is being monitored by looking at ONE file.

Reason 4: Service Discovery Integration

Prometheus can automatically discover targets using:
  β†’ Kubernetes API
  β†’ AWS EC2 API
  β†’ Consul
  β†’ DNS records
  β†’ File-based configs

When a new pod is created in Kubernetes, Prometheus automatically
discovers it and starts scraping. No agent configuration needed!

Reason 5: Natural Rate Limiting

In push model: 10,000 agents all push at the same time β†’ server overload!
In pull model: Prometheus controls the pace. It scrapes when IT's ready.
              If the server is overloaded, scrapes slow down naturally.

When Pull Model Doesn’t Work (The Exception)

Short-lived batch jobs (e.g., a cron job that runs for 5 seconds):
  β†’ The job finishes before Prometheus can scrape it
  β†’ Solution: Use the Pushgateway!
  
  Batch Job β†’ pushes metrics β†’ Pushgateway β†’ Prometheus scrapes Pushgateway
  
  This is the ONLY officially recommended use case for pushing in Prometheus.

πŸ’‘ Key Takeaway for Exam

Prometheus uses pull because:

  1. Better failure detection
  2. Easier debugging (curl /metrics)
  3. Centralized configuration
  4. Works with service discovery
  5. Natural rate limiting

Exception: Pushgateway for short-lived batch jobs



1.6 πŸ“– SLI, SLO, SLA (The Reliability Trio)

These three concepts are fundamental to Site Reliability Engineering (SRE) and frequently appear on the PCA exam.

SLI β€” Service Level INDICATOR

Definition: A quantitative measure of some aspect of the service level. It’s the actual metric you measure.

Format: Usually a ratio or percentage.

Examples:

SLI 1: Availability
  β†’ Proportion of successful HTTP requests
  β†’ Measurement: (successful requests / total requests) Γ— 100
  β†’ Current value: 99.95%

SLI 2: Latency
  β†’ Proportion of requests served faster than 200ms
  β†’ Measurement: (requests < 200ms / total requests) Γ— 100
  β†’ Current value: 98.5%

SLI 3: Throughput
  β†’ Number of requests processed per second
  β†’ Current value: 5,000 req/s

SLI 4: Error Rate
  β†’ Proportion of requests returning 5xx errors
  β†’ Current value: 0.02%

PromQL Examples for SLIs:

# Availability SLI (last 5 minutes)
sum(rate(http_requests_total{status=~"2.."}[5m])) 
/ 
sum(rate(http_requests_total[5m])) 
* 100

# Latency SLI: % of requests under 200ms
sum(rate(http_request_duration_seconds_bucket{le="0.2"}[5m])) 
/ 
sum(rate(http_request_duration_seconds_bucket{le="+Inf"}[5m])) 
* 100

# Error Rate SLI
sum(rate(http_requests_total{status=~"5.."}[5m])) 
/ 
sum(rate(http_requests_total[5m])) 
* 100

SLO β€” Service Level OBJECTIVE

Definition: A target value or range for an SLI. It’s the GOAL you want to achieve.

Format: β€œSLI should be β‰₯ X% over Y time period”

Examples:

SLO 1: "99.9% of HTTP requests should succeed over a 30-day window"
  β†’ SLI: Availability
  β†’ Target: 99.9%
  β†’ Window: 30 days

SLO 2: "95% of requests should complete within 500ms over 7 days"
  β†’ SLI: Latency
  β†’ Target: 95%
  β†’ Threshold: 500ms
  β†’ Window: 7 days

SLO 3: "Error rate should be less than 0.1% over 30 days"
  β†’ SLI: Error Rate
  β†’ Target: < 0.1%
  β†’ Window: 30 days

Key Concept β€” Error Budget:

If your SLO is 99.9% availability over 30 days:
  β†’ Total minutes in 30 days = 43,200 minutes
  β†’ Allowed downtime = 0.1% Γ— 43,200 = 43.2 minutes
  β†’ This 43.2 minutes is your ERROR BUDGET
  
  β†’ If you've used 30 minutes of downtime this month,
     you have 13.2 minutes of error budget remaining.
  β†’ If error budget is exhausted β†’ stop releasing new features,
     focus on reliability!

SLA β€” Service Level AGREEMENT

Definition: A formal contract between a service provider and customer that includes consequences (usually financial penalties) if the SLO is not met.

Key Point: SLA is a business/legal concept, not a technical one.

Examples:

SLA: "We guarantee 99.95% uptime per month. If we fail to meet this,
      you will receive a 10% credit on your monthly bill."

  β†’ SLO (internal): 99.95% uptime
  β†’ Consequence: 10% billing credit
  β†’ This is a LEGAL document signed by both parties

The Relationship (Exam Critical!)

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                                                         β”‚
β”‚   SLI (Indicator)    β†’  What you MEASURE                β”‚
β”‚        ↓                                                β”‚
β”‚   SLO (Objective)    β†’  What you AIM FOR (target)       β”‚
β”‚        ↓                                                β”‚
β”‚   SLA (Agreement)    β†’  What you PROMISE (contract)     β”‚
β”‚                                                         β”‚
β”‚   Hierarchy:  SLI βŠ‚ SLO βŠ‚ SLA                          β”‚
β”‚                                                         β”‚
β”‚   β€’ Every SLO is based on an SLI                        β”‚
β”‚   β€’ Every SLA contains one or more SLOs                 β”‚
β”‚   β€’ Not every SLO has an SLA (internal services)        β”‚
β”‚   β€’ SLO should be STRICTER than SLA (internal buffer)   β”‚
β”‚                                                         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ” Real-World Example

Netflix Streaming Service:

SLI:  Proportion of video play requests that start within 2 seconds
      Current measurement: 99.7%

SLO:  99.5% of video play requests should start within 2 seconds
      (Internal target β€” stricter than SLA)

SLA:  "We guarantee 99.0% of play requests start within 2 seconds.
       If we fall below this, enterprise customers get a refund."
      (External promise β€” less strict than SLO to give buffer)

Why SLO > SLA?
  β†’ If your SLO is 99.5% and SLA is 99.0%, you have a 0.5% buffer.
  β†’ You can breach your SLO internally and still not breach the SLA.
  β†’ This gives you time to react before customers are affected!

πŸ’‘ Key Takeaway for Exam

SLI = The metric (what you measure) β€” e.g., β€œ99.95% success rate” SLO = The target (what you aim for) β€” e.g., β€œβ‰₯ 99.9% success rate” SLA = The contract (what you promise + penalty) β€” e.g., β€œβ‰₯ 99.5% or refund”

SLO is always stricter than SLA (internal target > external promise) Not all SLOs have SLAs (internal services don’t need contracts)



1.7 πŸ“– GOLDEN SIGNALS (Google SRE Approach)

The Four Golden Signals were introduced by Google in their SRE Book. They represent the most important aspects of monitoring a user-facing system.

Signal 1: LATENCY ⏱️

Definition: The time it takes to service a request.

Important: Distinguish between success latency and error latency!

βœ… Good: "99th percentile latency for successful requests is 120ms"
❌ Bad:  "Average latency is 200ms" 
         (This hides the fact that 1% of requests take 10 seconds!)

Why separate success and error latency?
  β†’ A fast error (50ms 500 error) shouldn't drag down your latency metric
  β†’ A slow success (2s 200 OK) is still a problem even if errors are fast

PromQL Example:

# 99th percentile latency for successful requests
histogram_quantile(0.99, 
  sum(rate(http_request_duration_seconds_bucket{status=~"2.."}[5m])) 
  by (le)
)

# Median (50th percentile) latency
histogram_quantile(0.5, 
  sum(rate(http_request_duration_seconds_bucket[5m])) 
  by (le)
)

Signal 2: TRAFFIC 🚦

Definition: A measure of how much demand is being placed on your system.

Examples:

  • HTTP requests per second
  • Concurrent active users
  • Database queries per second
  • Messages per second in a queue

PromQL Example:

# Requests per second
sum(rate(http_requests_total[5m]))

# Traffic by endpoint
sum(rate(http_requests_total[5m])) by (handler)

# Concurrent connections
sum(node_netstat_Tcp_CurrEstab)

Why it matters:

  • Helps with capacity planning
  • High traffic + high latency = scaling problem
  • Sudden traffic spike = possible DDoS or viral event

Signal 3: ERRORS ❌

Definition: The rate of requests that fail, either explicitly (5xx) or implicitly (wrong content, slow response counted as error).

Types of Errors:

Explicit Errors:
  β†’ HTTP 500 Internal Server Error
  β†’ HTTP 503 Service Unavailable
  β†’ Connection refused
  β†’ Timeout

Implicit Errors:
  β†’ HTTP 200 but with wrong/empty content
  β†’ Response took longer than SLO threshold (e.g., > 2s)
  β†’ Partial data returned

PromQL Example:

# Error rate (5xx responses)
sum(rate(http_requests_total{status=~"5.."}[5m])) 
/ 
sum(rate(http_requests_total[5m])) 
* 100

# Error rate including timeouts
sum(rate(http_requests_total{status=~"5..|timeout"}[5m])) 
/ 
sum(rate(http_requests_total[5m])) 
* 100

Signal 4: SATURATION πŸ“ˆ

Definition: How β€œfull” your system is. A measure of resource utilization against capacity.

Examples:

  • CPU utilization: 85% (capacity: 100%)
  • Memory usage: 12GB / 16GB = 75%
  • Disk usage: 450GB / 500GB = 90%
  • Thread pool: 180 / 200 active threads = 90%
  • Connection pool: 95 / 100 connections = 95%

PromQL Example:

# CPU Saturation
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)

# Memory Saturation
(1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100

# Disk Saturation
(1 - (node_filesystem_avail_bytes / node_filesystem_size_bytes)) * 100

# Thread Pool Saturation
app_active_threads / app_max_threads * 100

Why it matters:

  • Saturation > 80% = time to scale up
  • Saturation at 100% = system is degraded or will crash
  • Saturation often LEADS to increased latency and errors

The Golden Signals Relationship

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                                                      β”‚
β”‚   TRAFFIC increases                                  β”‚
β”‚       ↓                                              β”‚
β”‚   SATURATION increases (resources getting full)      β”‚
β”‚       ↓                                              β”‚
β”‚   LATENCY increases (system struggling)              β”‚
β”‚       ↓                                              β”‚
β”‚   ERRORS increase (system failing)                   β”‚
β”‚                                                      β”‚
β”‚   This is the typical degradation cascade!           β”‚
β”‚   Monitor all four to catch problems early.          β”‚
β”‚                                                      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ’‘ Key Takeaway for Exam

The Four Golden Signals are: Latency, Traffic, Errors, Saturation

  • Latency: Time to serve (separate success vs error!)
  • Traffic: Demand on the system
  • Errors: Rate of failures (explicit + implicit)
  • Saturation: How full your resources are

These are for user-facing systems (from Google SRE book)



1.8 πŸ“– RED METHOD (Tom Wilkie)

The RED Method was introduced by Tom Wilkie (Grafana Labs) and is specifically designed for microservices and request-driven systems.

R β€” Rate (Request Rate)

Definition: The number of requests per second your service is handling.

# Request rate per service
sum(rate(http_requests_total{job="payment-service"}[5m]))

# Request rate by endpoint
sum(rate(http_requests_total[5m])) by (handler, method)

# Example output:
# {handler="/api/pay", method="POST"} β†’ 250 req/s
# {handler="/api/refund", method="POST"} β†’ 15 req/s

E β€” Errors (Error Rate)

Definition: The number of failed requests per second (or as a percentage of total).

# Error rate (absolute)
sum(rate(http_requests_total{status=~"5.."}[5m]))

# Error rate (percentage)
sum(rate(http_requests_total{status=~"5.."}[5m])) 
/ 
sum(rate(http_requests_total[5m])) 
* 100

# Example: 2.5% of requests are failing

D β€” Duration (Request Duration/Latency)

Definition: The distribution of request durations (how long requests take).

# 99th percentile duration
histogram_quantile(0.99, 
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)

# Average duration
sum(rate(http_request_duration_seconds_sum[5m])) 
/ 
sum(rate(http_request_duration_seconds_count[5m]))

# Example: p99 = 450ms, avg = 120ms

When to Use RED

βœ… USE RED for:
  β†’ Microservices
  β†’ API endpoints
  β†’ Request-driven architectures
  β†’ HTTP/gRPC services
  β†’ When you care about the REQUEST perspective

❌ DON'T use RED for:
  β†’ Infrastructure (servers, disks, networks)
  β†’ Batch processing systems
  β†’ Storage systems
  β†’ When you care about the RESOURCE perspective

πŸ’‘ Key Takeaway for Exam

RED = Rate, Errors, Duration Designed for microservices and request-driven systems Think of it as the β€œper-service” version of Golden Signals



1.9 πŸ“– USE METHOD (Brendan Gregg)

The USE Method was created by Brendan Gregg (Netflix) and is designed for infrastructure and resource monitoring.

U β€” Utilization

Definition: The average time that a resource was busy servicing work (as a percentage).

# CPU Utilization
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# Example: 75% β†’ CPU is busy 75% of the time

# Disk I/O Utilization
rate(node_disk_io_time_seconds_total[5m]) * 100
# Example: 60% β†’ Disk is busy 60% of the time

# Network Utilization
rate(node_network_transmit_bytes_total[5m]) / node_network_speed_bytes * 100

S β€” Saturation

Definition: The amount of work a resource has to do, often measured as queue length or wait time. It indicates how much work is waiting because the resource is full.

# CPU Saturation (load average vs CPU count)
node_load1 / count(node_cpu_seconds_total{mode="idle"})
# Example: 2.5 β†’ On average, 2.5 processes are waiting for CPU per core

# Memory Saturation (swap usage indicates memory saturation)
rate(node_vmstat_pswpout[5m])
# Example: High swap-out rate = memory is saturated

# Disk Saturation (I/O wait queue)
rate(node_disk_io_time_weighted_seconds_total[5m])

Key Difference: Utilization vs Saturation

Utilization: "The CPU is busy 90% of the time" (it's working)
Saturation:  "There are 50 processes waiting in the run queue" (it's overwhelmed)

A resource can be 100% utilized but NOT saturated (if it's keeping up).
A resource is saturated when work is QUEUED and waiting.

E β€” Errors

Definition: The count of error events for a resource.

# Network errors
rate(node_network_receive_errs_total[5m])
rate(node_network_transmit_errs_total[5m])

# Disk errors
rate(node_disk_io_errors_total[5m])

# Memory errors (ECC errors, OOM kills)
rate(node_vmstat_oom_kill[5m])

# File descriptor errors
node_filefd_allocated / node_filefd_maximum * 100

When to Use USE

βœ… USE USE for:
  β†’ Infrastructure monitoring (servers, VMs, containers)
  β†’ Hardware resources (CPU, memory, disk, network)
  β†’ Finding bottlenecks at the resource level
  β†’ Capacity planning

❌ DON'T use USE for:
  β†’ Application-level metrics
  β†’ Microservice request flows
  β†’ Business metrics

πŸ’‘ Key Takeaway for Exam

USE = Utilization, Saturation, Errors Designed for infrastructure and resource monitoring Apply to EVERY resource: CPU, memory, disk, network Utilization β‰  Saturation (busy vs overwhelmed)



1.10 πŸ“– COMPARISON: GOLDEN SIGNALS vs RED vs USE

Aspect Golden Signals RED USE
Creator Google SRE Tom Wilkie Brendan Gregg
Components Latency, Traffic, Errors, Saturation Rate, Errors, Duration Utilization, Saturation, Errors
Best for User-facing systems Microservices Infrastructure
Perspective User experience Request/Service Resource/Hardware
Scope Broad Narrow (per service) Narrow (per resource)
Example β€œUsers see 2s latency” β€œPayment API: 500 req/s, 2% errors, p99=300ms” β€œCPU: 85% util, load=4, 0 errors”
When to use Overall system health Debugging a specific service Debugging a specific server

How They Work Together in Practice

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    FULL OBSERVABILITY                    β”‚
β”‚                                                          β”‚
β”‚  Layer 1: USE Method (Infrastructure)                    β”‚
β”‚    β†’ "Server-3 CPU is 95% utilized, load average = 8"    β”‚
β”‚                                                          β”‚
β”‚  Layer 2: RED Method (Microservices)                     β”‚
β”‚    β†’ "Payment-service: error rate spiked to 5%"          β”‚
β”‚    β†’ "Order-service: p99 latency is 3 seconds"           β”‚
β”‚                                                          β”‚
β”‚  Layer 3: Golden Signals (User Experience)               β”‚
β”‚    β†’ "Checkout page latency is 5s for 10% of users"      β”‚
β”‚    β†’ "Overall error rate is 3%"                          β”‚
β”‚                                                          β”‚
β”‚  All three layers together give you complete picture!    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜


1.11 πŸ“– REAL-WORLD SCENARIOS

Scenario 1: E-Commerce Black Friday

Situation: It's Black Friday. Traffic spikes 10x.

Golden Signals:
  β†’ Traffic: 50,000 req/s (normally 5,000) ⚠️
  β†’ Latency: p99 went from 200ms to 2s ⚠️
  β†’ Errors: 5xx rate at 8% (normally 0.1%) πŸ”΄
  β†’ Saturation: CPU at 98%, memory at 95% πŸ”΄

RED (per service):
  β†’ Checkout-service: Rate=5000/s, Errors=12%, Duration p99=5s πŸ”΄
  β†’ Catalog-service: Rate=30000/s, Errors=0.5%, Duration p99=100ms βœ…
  β†’ Payment-service: Rate=4000/s, Errors=15%, Duration p99=8s πŸ”΄

USE (per server):
  β†’ Web-server-1: CPU=99%, Saturation=load 20, Errors=0 πŸ”΄
  β†’ DB-server-1: CPU=95%, Saturation=IOPS at max, Errors=timeout πŸ”΄

Conclusion: Database is the bottleneck. Scale DB and add caching.

Scenario 2: Microservice Deployment Gone Wrong

Situation: New version of user-service deployed. Users report slowness.

SLI: Success rate dropped from 99.95% to 97%
SLO: Target is 99.9% β†’ BREACHED! πŸ”΄

Investigation using RED:
  β†’ user-service: Rate=1000/s (normal), Errors=3% (was 0.05%), Duration p99=1s (was 100ms)
  
Root cause found via logs: New code introduced N+1 query problem.
Each user request now makes 50 DB queries instead of 1.

Fix: Rollback deployment, fix the query, redeploy.


1.12 πŸ“ EXAM-STYLE QUESTIONS (30 Questions)

Question 1

What is the primary difference between monitoring and observability?

A) Monitoring uses metrics while observability uses logs B) Monitoring is proactive while observability is reactive C) Monitoring answers β€œis something wrong?” while observability answers β€œwhy is something wrong?” D) Monitoring is for infrastructure while observability is for applications

βœ… Answer & Explanation

Correct Answer: C

Monitoring tells you IF there is a problem (reactive, predefined alerts). Observability helps you understand WHY there is a problem (proactive, exploratory, ad-hoc queries). Monitoring is a subset of observability. A is wrong because both use metrics, logs, and traces. B is reversed. D is incorrect β€” both apply to all layers.


Question 2

Which of the following is NOT one of the three pillars of observability?

A) Metrics B) Logs C) Traces D) Alerts

βœ… Answer & Explanation

Correct Answer: D

The three pillars of observability are Metrics, Logs, and Traces. Alerts are a feature built ON TOP of metrics (or logs/traces), not a pillar themselves. Alerts are derived from the data collected through the three pillars.


Question 3

In the context of observability, what is a β€œtrace”?

A) A log entry that tracks a specific error B) A metric that shows the trend of a value over time C) The end-to-end journey of a single request through a distributed system D) A recording rule in Prometheus

βœ… Answer & Explanation

Correct Answer: C

A trace represents the complete path of a single request as it flows through multiple services in a distributed system. It is composed of multiple spans, each representing a unit of work. Traces are essential for understanding latency and bottlenecks in microservice architectures.


Question 4

Which monitoring model does Prometheus use?

A) Push model B) Pull model C) Hybrid model (push and pull equally) D) Agent-based push model

βœ… Answer & Explanation

Correct Answer: B

Prometheus primarily uses a pull model where the Prometheus server scrapes metrics from targets at configured intervals. The only exception is the Pushgateway, which is used for short-lived batch jobs that cannot be scraped.


Question 5

What is a key advantage of the pull model over the push model?

A) It works better behind firewalls B) It is easier to detect when a target goes down C) It requires less network bandwidth D) It supports short-lived batch jobs natively

βœ… Answer & Explanation

Correct Answer: B

In a pull model, if a target stops responding to scrapes, Prometheus immediately knows the target is down (scrape failure). In a push model, if data stops arriving, it’s ambiguous β€” the target could be down, the agent could have crashed, or there could be a network issue. A and D are advantages of the push model. C is not necessarily true.


Question 6

When should you use the Prometheus Pushgateway?

A) For all monitoring targets behind a firewall B) For long-running services that are hard to reach C) For short-lived batch jobs that may not exist long enough to be scraped D) For high-traffic services that generate too many metrics

βœ… Answer & Explanation

Correct Answer: C

The Pushgateway is specifically designed for short-lived batch jobs (e.g., cron jobs, CI/CD pipelines) that may complete before Prometheus has a chance to scrape them. The batch job pushes its metrics to the Pushgateway, and Prometheus scrapes the Pushgateway. It should NOT be used as a general replacement for the pull model.


Question 7

What does SLI stand for in the context of reliability?

A) Service Level Integration B) Service Level Indicator C) System Latency Index D) Service Log Interface

βœ… Answer & Explanation

Correct Answer: B

SLI stands for Service Level Indicator. It is a quantitative measure of some aspect of the service level, such as availability percentage, latency percentile, or error rate. It’s the actual metric you measure to determine if you’re meeting your objectives.


Question 8

Which of the following correctly describes the relationship between SLI, SLO, and SLA?

A) SLA is stricter than SLO, which is stricter than SLI B) SLI is the metric, SLO is the target for that metric, SLA is the contract with consequences C) SLO and SLA are the same thing D) SLI is the contract, SLO is the metric, SLA is the target

βœ… Answer & Explanation

Correct Answer: B

SLI (Service Level Indicator) is the quantitative measure (e.g., β€œ99.95% success rate”). SLO (Service Level Objective) is the target you aim for (e.g., β€œβ‰₯ 99.9% success rate”). SLA (Service Level Agreement) is the formal contract that includes consequences for not meeting the SLO (e.g., β€œβ‰₯ 99.5% or you get a refund”). The hierarchy is SLI β†’ SLO β†’ SLA.


Question 9

If your SLO is 99.9% availability over a 30-day period, what is your error budget?

A) 43.2 minutes of allowed downtime B) 4.32 minutes of allowed downtime C) 432 minutes of allowed downtime D) 0.1 minutes of allowed downtime

βœ… Answer & Explanation

Correct Answer: A

30 days = 30 Γ— 24 Γ— 60 = 43,200 minutes. Error budget = (100% - 99.9%) Γ— 43,200 = 0.1% Γ— 43,200 = 43.2 minutes. This means you can afford 43.2 minutes of downtime per month while still meeting your SLO.


Question 10

Which of the following is NOT one of the Four Golden Signals?

A) Latency B) Traffic C) Utilization D) Saturation

βœ… Answer & Explanation

Correct Answer: C

The Four Golden Signals (from Google SRE) are: Latency, Traffic, Errors, and Saturation. Utilization is part of the USE Method (Utilization, Saturation, Errors), not the Golden Signals. This is a common trick question on the exam!


Question 11

In the context of the Golden Signals, why is it important to distinguish between success latency and error latency?

A) Error latency is always higher than success latency B) Fast errors should not skew the latency measurement of successful requests C) Prometheus cannot measure error latency D) Success latency is not important for SLOs

βœ… Answer & Explanation

Correct Answer: B

A fast error response (e.g., a 500 error returned in 10ms) should not artificially lower your overall latency metric. Similarly, a slow successful response is still a problem. By separating success and error latency, you get a more accurate picture of user experience. Users who get errors have a different experience than users who get slow successes.


Question 12

What does the RED method stand for?

A) Reliability, Errors, Duration B) Rate, Errors, Duration C) Requests, Exceptions, Delays D) Response, Errors, Distribution

βœ… Answer & Explanation

Correct Answer: B

RED stands for Rate (requests per second), Errors (failed requests), and Duration (distribution of request latencies). It was introduced by Tom Wilkie and is specifically designed for monitoring microservices and request-driven systems.


Question 13

The USE method is best suited for monitoring which of the following?

A) Microservice request flows B) User experience metrics C) Infrastructure resources (CPU, memory, disk, network) D) Business KPIs

βœ… Answer & Explanation

Correct Answer: C

The USE Method (Utilization, Saturation, Errors) by Brendan Gregg is designed for infrastructure and resource-level monitoring. You apply it to every hardware resource: CPU, memory, disk, network, etc. For microservices, use the RED method. For user-facing systems, use Golden Signals.


Question 14

What is the difference between utilization and saturation in the USE method?

A) They are the same thing B) Utilization measures how busy a resource is; saturation measures how much work is queued waiting for the resource C) Utilization is for CPU; saturation is for memory D) Saturation is always lower than utilization

βœ… Answer & Explanation

Correct Answer: B

Utilization = the percentage of time a resource is busy (e.g., CPU at 80%). Saturation = the amount of work waiting in queue because the resource is full (e.g., 50 processes waiting for CPU). A resource can be 100% utilized but not saturated (if it’s keeping up with demand). Saturation indicates the resource is overwhelmed.


Question 15

Which of the following is an example of white-box monitoring?

A) Using the Blackbox Exporter to probe https://myapp.com B) Checking if a website returns HTTP 200 from an external location C) Instrumenting application code to expose internal metrics like active database connections D) Pinging a server to check if it’s reachable

βœ… Answer & Explanation

Correct Answer: C

White-box monitoring involves instrumenting the internals of a system to expose metrics about its inner workings (e.g., active connections, queue depth, internal error rates). Options A, B, and D are all examples of black-box monitoring β€” observing the system from the outside without knowledge of its internals.


Question 16

Which Prometheus component is used for black-box monitoring?

A) Node Exporter B) Pushgateway C) Blackbox Exporter D) Alertmanager

βœ… Answer & Explanation

Correct Answer: C

The Blackbox Exporter is used for black-box monitoring. It probes endpoints over HTTP, HTTPS, DNS, TCP, and ICMP to check availability, response time, SSL certificate expiry, etc. Node Exporter is for white-box infrastructure metrics. Pushgateway is for batch jobs. Alertmanager handles alerts.


Question 17

In a pull-based monitoring system, what happens when a target becomes unreachable?

A) The monitoring server continues to use the last known metrics B) The monitoring server immediately detects the failure because the scrape fails C) The target sends an alert to the monitoring server D) The monitoring server switches to push mode automatically

βœ… Answer & Explanation

Correct Answer: B

In a pull model, the monitoring server (Prometheus) actively scrapes targets. If a target becomes unreachable, the scrape fails immediately, and Prometheus marks the target as DOWN. This is one of the key advantages of the pull model β€” immediate and unambiguous failure detection. Note: Prometheus does keep the last known value for a short time (staleness period of 5 minutes), but the target is marked as down.


Question 18

Which of the following best describes an β€œunknown unknown” in the context of observability?

A) A known issue that you have an alert for B) A problem you didn’t anticipate and don’t have a predefined dashboard for C) A metric that has no data D) A target that is not configured in Prometheus

βœ… Answer & Explanation

Correct Answer: B

β€œUnknown unknowns” are problems you didn’t anticipate and can’t predict. Observability enables you to investigate these issues using ad-hoc queries and high-cardinality data, even without predefined dashboards. β€œKnown unknowns” are problems you expect and have alerts for (monitoring handles these).


Question 19

Which of the following PromQL expressions correctly calculates an availability SLI?

A) sum(http_requests_total{status="200"}) / sum(http_requests_total) B) sum(rate(http_requests_total{status=~"2.."}[5m])) / sum(rate(http_requests_total[5m])) C) count(http_requests_total{status="200"}) / count(http_requests_total) D) avg(http_requests_total{status=~"2.."}) * 100

βœ… Answer & Explanation

Correct Answer: B

To calculate an availability SLI, you need to use rate() because http_requests_total is a counter. You calculate the rate of successful requests (2xx status codes) divided by the rate of all requests over a time window. Option A uses raw counter values (incorrect for counters). Option C uses count() which counts time series, not values. Option D uses avg() which is meaningless for counters.


Question 20

What is an β€œerror budget” in the context of SLOs?

A) The maximum number of errors your application can generate per day B) The amount of acceptable unreliability, calculated as 100% minus the SLO target C) The budget allocated for fixing errors in the application D) The number of alerts you can receive before escalating

βœ… Answer & Explanation

Correct Answer: B

An error budget is the amount of acceptable unreliability. If your SLO is 99.9%, your error budget is 0.1%. This represents the maximum amount of downtime or errors you can have while still meeting your SLO. When the error budget is exhausted, teams should focus on reliability instead of new features.


Question 21

Which monitoring approach would be MOST appropriate for detecting that a new deployment caused a regression in a specific microservice?

A) USE Method on the underlying servers B) Black-box monitoring of the website homepage C) RED Method on the specific microservice D) Monitoring network bandwidth utilization

βœ… Answer & Explanation

Correct Answer: C

The RED Method (Rate, Errors, Duration) is designed for microservices. By monitoring the specific microservice’s request rate, error rate, and request duration, you can quickly detect if a new deployment caused increased errors or latency. USE method is for infrastructure, black-box is too high-level, and network bandwidth is too low-level.


Question 22

In the push model, why is it difficult to detect a dead target?

A) The monitoring server cannot process push data fast enough B) The target might just be slow to push, or the network might be partitioned, making it ambiguous C) Push model does not support alerting D) The monitoring server has no configuration for the target

βœ… Answer & Explanation

Correct Answer: B

In a push model, when data stops arriving, it’s ambiguous: the target could be dead, the agent could have crashed, the network could be partitioned, or the target could just be slow. In contrast, in a pull model, a failed scrape is a clear signal that something is wrong with the target or the network path to it.


Question 23

Which of the following is a characteristic of metrics as an observability pillar?

A) They provide detailed per-request information B) They are expensive to store at scale C) They are lightweight, pre-aggregated, and ideal for dashboards and alerting D) They show the full request path across microservices

βœ… Answer & Explanation

Correct Answer: C

Metrics are lightweight numerical measurements that are pre-aggregated over time. They are cheap to store, ideal for dashboards and alerting, and great for trend analysis. However, they lose individual event details (A is logs/traces). B describes logs. D describes traces.


Question 24

What is the relationship between monitoring and observability?

A) They are exactly the same thing B) Observability is a subset of monitoring C) Monitoring is a subset of observability D) They are completely unrelated concepts

βœ… Answer & Explanation

Correct Answer: C

Monitoring is a subset of observability. Monitoring focuses on predefined metrics and alerts for known issues. Observability encompasses monitoring but goes further, enabling you to explore and understand unknown issues through metrics, logs, and traces. You can have monitoring without observability, but not observability without monitoring.


Question 25

Which of the following is an example of an implicit error?

A) HTTP 500 Internal Server Error B) HTTP 503 Service Unavailable C) HTTP 200 OK with incorrect or empty response body D) Connection timeout

βœ… Answer & Explanation

Correct Answer: C

An implicit error is when the response appears successful (HTTP 200) but the content is wrong, empty, or stale. The system didn’t explicitly signal an error, but from the user’s perspective, the request failed. Options A, B, and D are explicit errors β€” the system clearly indicates failure.


Question 26

You are monitoring a Kubernetes cluster. Which method would you use to check if the nodes have enough CPU and memory capacity?

A) RED Method B) Golden Signals only C) USE Method D) Black-box probing

βœ… Answer & Explanation

Correct Answer: C

The USE Method (Utilization, Saturation, Errors) is designed for infrastructure resources. You would check CPU utilization (%), CPU saturation (load average/run queue), memory utilization (%), memory saturation (swap usage/OOM kills), and any hardware errors. RED is for microservices. Golden Signals are for user-facing systems. Black-box probing checks external availability.


Question 27

Why is centralized configuration an advantage of the pull model?

A) It allows targets to configure themselves automatically B) All monitoring targets and scrape settings are defined in one place on the monitoring server C) It eliminates the need for service discovery D) It reduces the number of metrics collected

βœ… Answer & Explanation

Correct Answer: B

In the pull model, all target configurations (what to scrape, how often, which labels to add) are defined in a single configuration file (prometheus.yml) on the Prometheus server. This makes it easy to audit, manage, and understand what is being monitored. In the push model, each agent needs to be individually configured with the server address.


Question 28

Which of the following statements about SLOs and SLAs is TRUE?

A) Every SLO must have a corresponding SLA B) SLOs should typically be stricter than SLAs to provide an internal buffer C) SLAs are internal targets and SLOs are external contracts D) SLOs and SLAs always have the same target value

βœ… Answer & Explanation

Correct Answer: B

SLOs (internal targets) should be stricter than SLAs (external promises) to provide a buffer. For example, your SLO might be 99.95% while your SLA promises 99.9%. This gives you time to detect and fix issues before you breach the SLA. Not every SLO needs an SLA (internal services). C is reversed. D is incorrect β€” they should differ.


Question 29

What is a β€œspan” in the context of distributed tracing?

A) A time range in a PromQL query B) A single unit of work within a trace, representing one operation C) The total duration of a trace D) A type of Prometheus metric

βœ… Answer & Explanation

Correct Answer: B

A span is a single unit of work within a trace. Each span has a span ID, a parent span ID (except the root span), a start time, a duration, and metadata (tags/logs). A trace is composed of multiple spans that together represent the full journey of a request through a distributed system.


Question 30

You notice that your application’s p99 latency has increased from 200ms to 2s, but the average latency is still 150ms. What does this indicate?

A) The system is healthy because the average is low B) A small percentage of requests are experiencing very high latency, which the average hides C) The p99 metric is broken and should be ignored D) All requests are now taking 2 seconds

βœ… Answer & Explanation

Correct Answer: B

This is a classic example of why averages are misleading. The average latency of 150ms hides the fact that the worst 1% of requests (p99) are taking 2 seconds. This could affect a significant number of users in a high-traffic system. This is why the Golden Signals recommend looking at percentile latency (p50, p95, p99) rather than averages. The p99 is not broken β€” it’s revealing a real problem.



βœ… DOMAIN 1 REVISION SUMMARY CHEAT SHEET

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    DOMAIN 1 CHEAT SHEET                      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                              β”‚
β”‚  Monitoring = "Is something wrong?" (Reactive)               β”‚
β”‚  Observability = "WHY is something wrong?" (Proactive)       β”‚
β”‚  Monitoring βŠ‚ Observability                                  β”‚
β”‚                                                              β”‚
β”‚  3 Pillars: Metrics (numbers) + Logs (events) + Traces (path)β”‚
β”‚  Prometheus = Metrics system primarily                       β”‚
β”‚                                                              β”‚
β”‚  White-Box = Internal instrumentation (client libraries)     β”‚
β”‚  Black-Box = External probing (Blackbox Exporter)            β”‚
β”‚                                                              β”‚
β”‚  Pull Model = Prometheus scrapes targets                     β”‚
β”‚    βœ… Easy failure detection, centralized config, debugging   β”‚
β”‚    ❌ Firewall issues, needs service discovery                β”‚
β”‚  Push Model = Targets send to server (Graphite, Datadog)     β”‚
β”‚    βœ… Works behind firewalls, good for batch jobs             β”‚
β”‚    ❌ Hard to detect dead targets, distributed config         β”‚
β”‚  Pushgateway = ONLY for short-lived batch jobs               β”‚
β”‚                                                              β”‚
β”‚  SLI = Measure (e.g., 99.95% success rate)                   β”‚
β”‚  SLO = Target (e.g., β‰₯ 99.9%)  [Internal, stricter]         β”‚
β”‚  SLA = Contract (e.g., β‰₯ 99.5% or refund) [External]        β”‚
β”‚  Error Budget = 100% - SLO%                                  β”‚
β”‚                                                              β”‚
β”‚  Golden Signals: Latency, Traffic, Errors, Saturation        β”‚
β”‚  RED Method: Rate, Errors, Duration (microservices)          β”‚
β”‚  USE Method: Utilization, Saturation, Errors (infra)         β”‚
β”‚                                                              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

🎯 NEXT STEPS

Once you’ve thoroughly studied this material and can answer all 30 questions correctly:

Reply with β€œDomain 1 Complete” and I’ll provide the same comprehensive guide for Domain 2: Prometheus Fundamentals (20%) covering Architecture, Data Model, Configuration, Metric Types, Storage, and more!

Good luck with your revision! πŸ’ͺ