LearnThatStack Ace your next interview
Part of System Design Concepts

Observability & Monitoring.

Start free Change topic Change
Practice · Questions

All questions

Showing of 133
Beginner 31
01

What is observability and how does it differ from monitoring?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Observability is the ability to understand the internal state of a system by examining its external outputs. It's about being able to ask arbitrary questions about your system's behavior and get answers from the data you collect.
Monitoring is the practice of collecting, aggregating, and acting on metrics and logs from your systems. It typically involves predefined dashboards and alerts for known issues.
Key differences:

  • Scope: Monitoring focuses on known problems; observability helps discover unknown issues
  • Approach: Monitoring is reactive; observability is more exploratory
  • Data: Monitoring uses predetermined metrics; observability requires rich, high-cardinality data
  • Questions: Monitoring answers "Is the system working?"; observability answers "Why isn't it working?"
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

02

What are the three pillars of observability?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

The three pillars of observability are:

  1. Metrics: Quantitative data about system performance (CPU usage, response time, error rates)
  2. Logs: Discrete events that happened in the system with context and details
  3. Traces: Records of requests as they flow through multiple services in distributed systems
    These pillars work together to provide comprehensive visibility:
  • Metrics show you WHAT is happening
  • Logs tell you WHY it's happening
  • Traces show you WHERE it's happening across services
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

03

What is telemetry data and why is it important?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Telemetry data is automatically collected information about the behavior, performance, and health of systems, applications, and infrastructure. It includes metrics, logs, traces, and events generated by software and hardware components.
Importance:

  • Proactive issue detection: Identify problems before users experience them
  • Performance optimization: Understand bottlenecks and optimization opportunities
  • Capacity planning: Make informed decisions about scaling
  • Debugging: Troubleshoot issues in production environments
  • Business insights: Understand user behavior and system usage patterns
  • Compliance: Meet regulatory requirements for system monitoring
    Types of telemetry:
  • Performance metrics (latency, throughput)
  • Error rates and exceptions
  • Resource utilization (CPU, memory, disk)
  • Business metrics (user actions, revenue)
  • Security events and audit logs
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

04

What are the different log levels and when should you use each?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Standard log levels in order of severity:
TRACE: Finest level of detail, typically for debugging specific code paths

TRACE: Entering method calculateTax() with amount=100.50

DEBUG: Detailed information for debugging, not for production

DEBUG: Database query: SELECT * FROM users WHERE id = 123

INFO: General information about application flow

INFO: User 123 successfully logged in

WARN: Indicates a potential issue that doesn't stop the application

WARN: API rate limit approaching: 950/1000 requests used

ERROR: Error conditions that don't stop the application

ERROR: Failed to send email notification: SMTP timeout

FATAL: Severe errors that may cause the application to terminate

FATAL: Cannot connect to database, shutting down

Best practices:

  • Use INFO for business events and milestones
  • Use WARN for recoverable errors and degraded states
  • Use ERROR for actual failures that need attention
  • Use DEBUG/TRACE only in development (performance impact)
  • Be consistent across your organization
  • Include context (user ID, request ID, session ID)
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

05

What is Grafana and how does it integrate with different data sources?

Part of Pro
06

What is Prometheus and how does it differ from traditional monitoring systems?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Prometheus is an open-source monitoring and alerting toolkit designed for reliability and scalability. Key differences from traditional monitoring systems:

  • Pull-based model: Prometheus scrapes metrics from targets rather than receiving pushed data
  • Time-series database: Built-in TSDB optimized for metric storage
  • Multi-dimensional data model: Metrics identified by metric name and key-value pairs (labels)
  • Powerful query language: PromQL for flexible data analysis
  • Service discovery: Automatic target discovery from various sources
  • No external dependencies: Self-contained system

Traditional systems often use push-based models, require external databases, and have less flexible data models.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

07

Explain the main components of Prometheus architecture.

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Prometheus architecture consists of several key components:

  • Prometheus Server: Core component that scrapes and stores metrics, serves queries
  • Client Libraries: For instrumenting application code to expose metrics
  • Pushgateway: For short-lived jobs that can't be scraped directly
  • Exporters: Proxy metrics from third-party systems (node_exporter, blackbox_exporter)
  • Alertmanager: Handles alerts sent by Prometheus, manages routing, grouping, and notifications
  • Service Discovery: Mechanisms to automatically discover targets
  • Grafana: Common visualization layer (though separate project)
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

08

What are the four metric types in Prometheus?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Prometheus supports four core metric types:

  1. Counter: Cumulative metric that only increases (or resets to zero). Used for requests served, errors occurred, etc.
  2. Gauge: Metric that can go up and down. Used for temperature, memory usage, concurrent requests
  3. Histogram: Samples observations and counts them in configurable buckets. Also provides sum and count
  4. Summary: Similar to histogram but calculates configurable quantiles over a sliding time window

Example:

# Counter
http_requests_total{method="GET",status="200"} 1234

# Gauge  
memory_usage_bytes 8589934592

# Histogram
http_request_duration_seconds_bucket{le="0.1"} 100
http_request_duration_seconds_bucket{le="0.5"} 150
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

09

How would you install Prometheus on a Linux system?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Multiple installation methods:

Binary Installation:

# Download and extract
wget https://github.com/prometheus/prometheus/releases/download/v2.45.0/prometheus-2.45.0.linux-amd64.tar.gz
tar xvfz prometheus-*.tar.gz
cd prometheus-*

# Create user and directories
sudo useradd --no-create-home --shell /bin/false prometheus
sudo mkdir /etc/prometheus /var/lib/prometheus
sudo chown prometheus:prometheus /var/lib/prometheus

# Copy binaries and config
sudo cp prometheus promtool /usr/local/bin/
sudo cp -r consoles/ console_libraries/ /etc/prometheus/
sudo cp prometheus.yml /etc/prometheus/
sudo chown -R prometheus:prometheus /etc/prometheus/

Package Manager:

# Ubuntu/Debian
sudo apt update && sudo apt install prometheus

# RHEL/CentOS
sudo yum install prometheus

Docker:

docker run -p 9090:9090 prom/prometheus
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

10

What is the default port for Prometheus and how do you change it?

Part of Pro
11

How do you reload Prometheus configuration without restarting?

Part of Pro
12

Explain file-based service discovery with an example.

Part of Pro
13

How do you silence alerts in Alertmanager?

Part of Pro
14

How do you configure data retention in Prometheus?

Part of Pro
15

How do you integrate Prometheus with Grafana?

Part of Pro
16

What is Kibana and how does it fit into the Elastic Stack?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Kibana is a data visualization and exploration tool that serves as the frontend interface for the Elastic Stack (formerly ELK Stack). It's designed to work with Elasticsearch as its primary data source and provides a web-based interface for searching, viewing, and interacting with data stored in Elasticsearch indices.
In the Elastic Stack architecture:

  • Elasticsearch stores and indexes the data
  • Logstash processes and transforms data before sending to Elasticsearch
  • Beats are lightweight data shippers that collect data
  • Kibana visualizes and explores the data
    Kibana allows users to create dashboards, visualizations, and perform real-time data analysis without needing to write complex queries directly against Elasticsearch.
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

17

Explain the main components and features of Kibana.

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Kibana consists of several key components:
Core Applications:

  • Discover: Interactive data exploration and search interface
  • Visualize: Create charts, graphs, and other visual representations
  • Dashboard: Combine multiple visualizations into unified views
  • Canvas: Create custom, pixel-perfect presentations
  • Maps: Geospatial data visualization and analysis
    Management Tools:
  • Dev Tools: Console for direct Elasticsearch API interaction
  • Stack Management: Configure index patterns, saved objects, and system settings
  • Stack Monitoring: Monitor Elastic Stack health and performance
  • Machine Learning: Anomaly detection and forecasting
  • Security: User authentication and role-based access control
  • Alerting: Create and manage alerts based on data conditions
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

18

What are the system requirements for running Kibana?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Kibana system requirements include:

  • RAM: Minimum 1GB, recommended 2GB or more
  • CPU: Multi-core processor recommended
  • Disk Space: Varies based on usage, typically 200MB for installation

Software Requirements:

  • Node.js: Built-in (comes with Kibana installation)
  • Operating System: Linux, macOS, or Windows
  • Browser: Modern browsers (Chrome, Firefox, Safari, Edge)

Network Requirements:

  • Network connectivity to Elasticsearch cluster
  • Default port 5601 for Kibana web interface
  • HTTPS configuration recommended for production

Elasticsearch Compatibility:

  • Kibana version must match Elasticsearch major version
  • Minor version differences are typically supported
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

19

How do you install and configure Kibana to connect to an Elasticsearch cluster?

Part of Pro
20

What are index patterns in Kibana and how do you create them?

Part of Pro
21

Describe the different types of visualizations available in Kibana and their use cases.

Part of Pro
22

How do you use Kibana's Discover interface for data exploration?

Part of Pro
23

What is Jaeger and what problem does it solve?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Jaeger is an open-source, end-to-end distributed tracing system originally developed by Uber. It helps monitor and troubleshoot complex microservices architectures by tracking requests as they flow through multiple services.

Problems it solves:

  • Performance bottlenecks: Identifies slow services in request chains
  • Error tracking: Pinpoints where failures occur in distributed systems
  • Dependency analysis: Maps service interactions and dependencies
  • Root cause analysis: Helps debug issues across multiple services
  • Service optimization: Provides insights for performance improvements

Jaeger follows the OpenTracing standard and is now part of the Cloud Native Computing Foundation (CNCF).

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

24

Explain the concept of distributed tracing and its key components.

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

Distributed tracing tracks requests as they travel through multiple services in a distributed system. It creates a complete picture of how a request is processed across different components.

Key components:

  • Trace: Complete journey of a request through the system
  • Span: Individual unit of work (e.g., HTTP request, database call)
  • SpanContext: Carries trace information between services
  • Tags: Key-value pairs that add metadata to spans
  • Logs: Timestamped events within spans
  • Baggage: Cross-service propagated key-value data

Example flow:

User Request → Service A → Service B → Database
     |            |          |          |
   Trace ID    Span 1     Span 2     Span 3
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

25

What are the differences between logging, metrics, and tracing?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

These are the three pillars of observability, each serving different purposes:

Logging:

  • Records discrete events with timestamps
  • Good for debugging specific issues
  • High volume, text-based
  • Example: "User 123 failed login at 2023-10-15 14:30:00"

Metrics:

  • Numerical measurements over time
  • Good for alerting and dashboards
  • Low volume, aggregated data
  • Example: CPU usage: 75%, Request rate: 100/sec

Tracing:

  • Shows request flow across services
  • Good for understanding system behavior
  • Medium volume, structured data
  • Example: Request path: Frontend → API → Database (150ms total)

Integration: Modern observability uses all three together - traces provide context for logs and metrics.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

26

What are the different deployment strategies in Jaeger?

Part of Pro
27

How do you instrument a microservice to work with Jaeger?

Part of Pro
28

What is OpenTelemetry and what are the three pillars of observability it addresses?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

OpenTelemetry (OTel) is an open-source observability framework that provides a unified approach to collecting, processing, and exporting telemetry data. It addresses the three pillars of observability:

  • Traces: Track requests as they flow through distributed systems, showing the path and timing of operations
  • Metrics: Numerical measurements collected over time intervals (counters, gauges, histograms)
  • Logs: Discrete event records with timestamps and contextual information

OpenTelemetry provides APIs, SDKs, and tools to instrument applications and infrastructure, making it easier to understand system behavior and performance.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

29

Explain the difference between OpenTelemetry and OpenTracing/OpenCensus.

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

OpenTelemetry is the successor and merger of two previous projects:

  • OpenTracing: Focused primarily on distributed tracing with vendor-neutral APIs
  • OpenCensus: Provided both tracing and metrics collection capabilities

Key differences:

  • OpenTelemetry combines both projects' capabilities into a single, comprehensive framework
  • Provides unified APIs for traces, metrics, and logs (the three pillars)
  • Better standardization and broader vendor support
  • More mature ecosystem with extensive instrumentation libraries
  • Backward compatibility with both OpenTracing and OpenCensus

OpenTracing and OpenCensus are now in maintenance mode, with new development focused on OpenTelemetry.

Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

30

What are the main components of the OpenTelemetry ecosystem?

Beginner ·

Answer it yourself first - out loud, or typed below.

How should your speech become text?

Listening… your words appear above as you speak - tap Stop when you're done.

Recording · cr - tap Stop & transcribe when you're done.

Transcribing with AI…

Voice:

Keep going - a few more words and AI can grade it.

Last attempt -

Your answer

Re-explain

The OpenTelemetry ecosystem consists of several key components:

  • APIs: Language-specific interfaces for creating telemetry data
  • SDKs: Implementations of the APIs with configuration capabilities
  • Instrumentation Libraries: Pre-built integrations for popular frameworks and libraries
  • Collector: Vendor-agnostic service for receiving, processing, and exporting telemetry data
  • Exporters: Components that send data to various backends (Jaeger, Prometheus, etc.)
  • Resources: Metadata describing the entity producing telemetry (service name, version, etc.)
  • Semantic Conventions: Standard attribute names and meanings for consistent data representation
Rewriting in plainer words…

This answer doesn't lend itself to a diagram - it reads best . No credits were charged.

Why there's no diagram: “”

The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓

Related concept

Tailored explanation · switch back to · ·
What should the new diagram focus on?
How well did you know this?
AI:

31

How do you configure a basic OpenTelemetry Collector pipeline?

Part of Pro
Intermediate 59
32

Explain the difference between white-box and black-box monitoring.

Part of Pro
33

Explain the different types of metrics: counter, gauge, histogram, and summary.

Part of Pro
34

What are the RED and USE methodologies?

Part of Pro
35

What is metric cardinality and why is it important to manage?

Part of Pro
36

Explain structured logging and its benefits over unstructured logging.

Part of Pro
37

What is the ELK stack and how does each component work?

Part of Pro
38

What is distributed tracing and why is it essential for microservices?

Part of Pro
39

Explain the concepts of trace, span, and span context in distributed tracing.

Part of Pro
40

How do you implement correlation IDs for request tracing?

Part of Pro
41

Explain the Prometheus architecture and data model.

Part of Pro
42

Explain the difference between SLI, SLO, and SLA with examples.

Part of Pro
43

What makes a good alert? Explain alert fatigue and how to prevent it.

Part of Pro
44

Explain the structure of prometheus.yml configuration file.

Part of Pro
45

What are recording rules and why are they useful?

Part of Pro
46

What is service discovery in Prometheus and why is it important?

Part of Pro
47

How do you configure Kubernetes service discovery?

Part of Pro
48

How do you configure alerting rules in Prometheus?

Part of Pro
49

What is Alertmanager and how does it work with Prometheus?

Part of Pro
50

How does Prometheus store data and what are the storage considerations?

Part of Pro
51

What is cardinality and why is it important for Prometheus performance?

Part of Pro
52

How do you monitor Prometheus itself?

Part of Pro
53

How do you secure a Prometheus installation?

Part of Pro
54

How do you implement authentication for Prometheus?

Part of Pro
55

How do you troubleshoot Prometheus scraping issues?

Part of Pro
56

How do you backup and restore Prometheus data?

Part of Pro
57

What is the Pushgateway and when should you use it?

Part of Pro
58

What are some best practices for Prometheus administration?

Part of Pro
59

Explain the Kibana configuration file structure and important settings.

Part of Pro
60

How do you handle multiple data sources in Kibana?

Part of Pro
61

How do you create and optimize dashboards in Kibana?

Part of Pro
62

Explain dashboard filters and how they work across multiple visualizations.

Part of Pro
63

What is KQL (Kibana Query Language) and how does it differ from Lucene query syntax?

Part of Pro
64

How do you use Kibana's Dev Tools Console and what are its main features?

Part of Pro
65

Explain Elasticsearch Query DSL and provide examples of common queries used in Kibana.

Part of Pro
66

How do you implement user authentication and authorization in Kibana?

Part of Pro
67

Explain Kibana Spaces and how they provide multi-tenancy.

Part of Pro
68

How do you set up monitoring for Kibana and the Elastic Stack?

Part of Pro
69

Explain how to create and manage alerts in Kibana.

Part of Pro
70

What is Kibana Canvas and how is it used for custom presentations?

Part of Pro
71

What are common Kibana startup issues and how do you resolve them?

Part of Pro
72

Describe Jaeger's architecture and its main components.

Part of Pro
73

What is the role of Jaeger Agent and why is it important?

Part of Pro
74

Explain the difference between Jaeger's push and pull models.

Part of Pro
75

How would you deploy Jaeger in a Kubernetes environment?

Part of Pro
76

How do you configure sampling in Jaeger?

Part of Pro
77

Explain how trace context propagation works across services.

Part of Pro
78

How do you handle tracing in asynchronous operations?

Part of Pro
79

What are the performance considerations when implementing Jaeger?

Part of Pro
80

What are the storage backend options and their trade-offs?

Part of Pro
81

How do you troubleshoot missing or incomplete traces in Jaeger?

Part of Pro
82

Explain how to perform trace analysis and create custom dashboards.

Part of Pro
83

Describe the OpenTelemetry Collector architecture and its main components.

Part of Pro
84

What is the OTLP protocol and why is it important?

Part of Pro
85

Explain the concept of Resources in OpenTelemetry.

Part of Pro
86

What are the different deployment patterns for OpenTelemetry Collector?

Part of Pro
87

How do you secure OpenTelemetry Collector communications?

Part of Pro
88

How do you monitor the health and performance of OpenTelemetry Collector itself?

Part of Pro
89

How do you troubleshoot common OpenTelemetry Collector issues?

Part of Pro
90

What are the security best practices for OpenTelemetry deployments?

Part of Pro
Expert 43
91

How do you choose appropriate metrics for monitoring a microservices architecture?

Part of Pro
92

How do you implement centralized logging in a microservices architecture?

Part of Pro
93

Compare Jaeger, Zipkin, and AWS X-Ray for distributed tracing.

Part of Pro
94

Compare APM tools: New Relic, AppDynamics, and Datadog.

Part of Pro
95

How do you implement error budgets and what are their benefits?

Part of Pro
96

Explain different alerting strategies: threshold-based, anomaly-based, and SLO-based.

Part of Pro
97

How do you design an effective on-call rotation and incident response process?

Part of Pro
98

How do you monitor and optimize application performance using observability data?

Part of Pro
99

Explain observability in cloud-native environments and containerized applications.

Part of Pro
100

What are some common observability anti-patterns and how do you avoid them?

Part of Pro
101

What is Prometheus federation and when would you use it?

Part of Pro
102

How do you scale Prometheus for large environments?

Part of Pro
103

What are common Prometheus performance issues and how do you resolve them?

Part of Pro
104

What are metric relabeling and target relabeling?

Part of Pro
105

How do you implement high availability for Prometheus?

Part of Pro
106

What are the security considerations when deploying Prometheus?

Part of Pro
107

What are the performance optimization techniques for Prometheus?

Part of Pro
108

What are the common performance issues in Kibana and how do you troubleshoot them?

Part of Pro
109

How do you optimize Elasticsearch queries for better Kibana performance?

Part of Pro
110

How do you implement Machine Learning capabilities in Kibana?

Part of Pro
111

Explain Kibana's integration with Elasticsearch Cross-Cluster Search.

Part of Pro
112

How do you debug performance issues in Kibana dashboards?

Part of Pro
113

How do you handle index pattern conflicts and mapping issues in Kibana?

Part of Pro
114

How do you scale Jaeger collectors?

Part of Pro
115

What are the security considerations for Jaeger deployment?

Part of Pro
116

What are the best practices for Jaeger in production?

Part of Pro
117

How do you integrate Jaeger with service mesh (Istio)?

Part of Pro
118

Explain how to implement custom span processors and samplers.

Part of Pro
119

How do you monitor Jaeger itself and set up alerting?

Part of Pro
120

How do you implement distributed tracing for batch jobs and scheduled tasks?

Part of Pro
121

How do you handle trace data retention and archival?

Part of Pro
122

How do you handle backpressure and reliability in OpenTelemetry Collector?

Part of Pro
123

Explain sampling strategies in OpenTelemetry and when to use each.

Part of Pro
124

How do you handle data loss prevention in OpenTelemetry deployments?

Part of Pro
125

How do you optimize OpenTelemetry Collector performance for high-throughput environments?

Part of Pro
126

What are the considerations for deploying OpenTelemetry Collector in Kubernetes at scale?

Part of Pro
127

How do you implement data governance and compliance in OpenTelemetry?

Part of Pro
128

How do you implement custom processors in OpenTelemetry Collector?

Part of Pro
129

Explain auto-instrumentation strategies and their trade-offs.

Part of Pro
130

How do you handle multi-tenancy in OpenTelemetry deployments?

Part of Pro
131

How do you implement distributed tracing correlation across service boundaries?

Part of Pro
132

How would you implement blue-green deployments with OpenTelemetry observability?

Part of Pro
133

How do you handle OpenTelemetry in serverless environments like AWS Lambda?

Part of Pro

No matches

Try a different filter or search term.

Know someone prepping for Observability & Monitoring? Send them this set.
Learn · Videos

Observability & Monitoring, in short videos.

Pro · $10/mo

116 of 133 Observability & Monitoring answers are in Pro.

Full answers, code samples, and AI explanations that go simpler or deeper. Cancel anytime.

  • Full answers + code
  • AI explanations, simpler or deeper
  • 1,000 AI credits / month
  • Cancel anytime

Change topic

Pick a different technology or stack. Your current topic stays put until you choose a new one.

Technologies
No technologies match “”.
Cross-cutting topics
No topics match “”.
By role
Stacks & frameworks

MEAN

MongoDB, Express, Angular, Node.js

MERN

MongoDB, Express, React, Node.js

LAMP

Linux, Apache, MySQL, PHP

Django

Python Full-Stack Development

Ruby on Rails

Convention over Configuration

Serverless on AWS

Serverless Architecture on AWS

Flutter Mobile

Flutter Cross-Platform Mobile Development

Spring Boot

Enterprise Java Development

.NET

Microsoft Ecosystem

Vue

Vue.js, Vite, TypeScript, Tailwind, Node.js

Go Backend

Golang, gRPC, PostgreSQL, Redis, RabbitMQ

FastAPI

Python, FastAPI, SQLAlchemy, PostgreSQL

React Native

React, TypeScript, Redux, Firebase

iOS Native

Swift, SwiftUI, UIKit, Firebase

Android Native

Java, Jetpack Compose, Firebase

DevOps / Platform

Docker, Kubernetes, Terraform, CI/CD

AI Engineer

LLMs, RAG, Agents, Evals

AI-Powered Developer

Claude Code, Copilot, Agentic Workflows

Core SWE Interview Prep

Data structures, algorithms, OS, concurrency, networking, git