LearnThatStack Ace your next interview

Observability System Design.
Cheat sheet.

Quick reference for Observability System Design - sectioned for fast scanning. Skim the part you're shaky on, walk in confident.

System Design Concepts 16-section reference ~5 min read

Summary

Comprehensive guide to observability for system design interviews. Covers the three pillars (metrics, logs, traces), distributed tracing, monitoring strategies, and tools. Critical for maintaining and debugging complex distributed systems.

What is Observability?

Definition: The ability to understand the internal state of a system by examining its outputs. Unlike monitoring (which tracks predefined metrics), observability helps you debug unknown unknowns.

Key Principle: "You can't fix what you can't see"

Three Pillars of Observability

1. Metrics (What happened?)

  • Definition: Numeric measurements collected at regular intervals
  • Characteristics: Aggregated, time-series data, low cardinality
  • Use Cases: Alerting, trending, capacity planning
  • Examples: CPU usage, request latency, error rates
# Example: Recording a metric
metric.increment('api.requests', tags=['endpoint:users', 'status:200'])
metric.gauge('memory.usage', 85.5)
metric.histogram('request.duration', 145, tags=['service:api'])

2. Logs (Why did it happen?)

  • Definition: Timestamped records of discrete events
  • Characteristics: High cardinality, structured or unstructured
  • Use Cases: Debugging, auditing, forensic analysis
  • Best Practice: Structured logging (JSON)
# Example: Structured logging
logger.info({
    "event": "user_login",
    "user_id": "12345",
    "ip": "192.168.1.1",
    "timestamp": "2024-01-20T10:30:00Z",
    "duration_ms": 230
})

3. Traces (How did it happen?)

  • Definition: End-to-end journey of a request through distributed systems
  • Components: Spans (units of work), trace ID, parent-child relationships
  • Use Cases: Performance optimization, dependency mapping, root cause analysis
# Example: Creating a trace span
with tracer.start_span('process_order') as span:
    span.set_tag('order.id', order_id)
    span.set_tag('customer.tier', 'premium')
    # Business logic here
    process_payment(order_id)

Key Concepts

Cardinality

  • Low Cardinality: Limited unique values (e.g., status codes: 200, 404, 500)
  • High Cardinality: Many unique values (e.g., user IDs, request IDs)
  • Impact: High cardinality increases storage costs and query complexity

Sampling

  • Head Sampling: Decide at the start whether to trace
  • Tail Sampling: Decide after completion based on characteristics
  • Adaptive Sampling: Adjust rate based on traffic patterns

Context Propagation

  • Passing correlation IDs across service boundaries
  • Standards: W3C Trace Context, B3 Headers
# Example: Context propagation
headers = {
    'X-Trace-Id': '1234567890abcdef',
    'X-Parent-Span-Id': 'abcdef123456',
    'X-Span-Id': 'fedcba098765'
}

Common Tools & Technologies

Metrics

  • Time Series DBs: Prometheus, InfluxDB, TimescaleDB
  • Commercial: DataDog, New Relic, CloudWatch
  • Visualization: Grafana, Kibana

Logging

  • Collection: Fluentd, Logstash, Vector
  • Storage: Elasticsearch, Splunk, CloudWatch Logs
  • Processing: Apache Kafka, AWS Kinesis

Tracing

  • Standards: OpenTelemetry, OpenTracing
  • Systems: Jaeger, Zipkin, AWS X-Ray
  • APM: AppDynamics, Dynatrace

📐 Architecture Patterns

1. Agent-Based

Application → Local Agent → Collector → Backend
  • Pros: Reduced application complexity
  • Cons: Resource overhead, another component to manage

2. Sidecar Pattern

Pod: [Application Container | Sidecar Container] → Backend
  • Pros: Language agnostic, separation of concerns
  • Cons: Increased resource usage

3. Direct Instrumentation

Application → Backend API
  • Pros: Simple, no intermediaries
  • Cons: Application must handle failures, buffering

Best Practices

1. Instrumentation

  • Start with SLIs (Service Level Indicators)
  • Use semantic conventions (OpenTelemetry)
  • Instrument at service boundaries
  • Add business context to technical metrics

2. Data Management

  • Implement retention policies
  • Use sampling for high-volume data
  • Compress and archive old data
  • Consider hot/warm/cold storage tiers

3. Alerting

  • Alert on symptoms, not causes
  • Use SLO-based alerts
  • Implement alert fatigue reduction
  • Include runbooks in alerts

4. Cost Optimization

  • Sample traces intelligently
  • Aggregate metrics at the edge
  • Use appropriate retention periods
  • Filter unnecessary data early

Interview Topics

Common Questions

  1. "How would you implement distributed tracing?"

    • Trace ID generation and propagation
    • Span collection and storage
    • Sampling strategies
    • Visualization needs
  2. "Design a logging system for a large-scale application"

    • Log aggregation architecture
    • Storage and indexing strategy
    • Search and analytics capabilities
    • Retention and compliance
  3. "How do you handle high cardinality metrics?"

    • Sampling techniques
    • Aggregation strategies
    • Storage optimization
    • Query performance

System Design Considerations

Scalability

  • Metrics: Pre-aggregation, downsampling
  • Logs: Partitioning, indexing strategies
  • Traces: Sampling, stream processing

Reliability

  • Buffering: Handle backend failures
  • Circuit Breakers: Prevent cascading failures
  • Graceful Degradation: Core functionality without observability

Performance

  • Async Collection: Don't block main thread
  • Batching: Reduce network overhead
  • Compression: Minimize data transfer

🚨 Red Flags & Anti-Patterns

  1. Over-instrumentation: Impacting application performance
  2. Under-sampling: Missing critical events
  3. No Correlation: Inability to connect metrics, logs, and traces
  4. Alert Fatigue: Too many non-actionable alerts
  5. Data Silos: Separate tools without integration

📝 Quick Reference

When to Use What?

Use Case Metrics Logs Traces
"Is there a problem?" ✅ ❌ ❌
"Where is the problem?" ⚠️ ✅ ✅
"What is the impact?" ✅ ⚠️ ⚠️
"Why did it happen?" ❌ ✅ ✅
"How often does it happen?" ✅ ⚠️ ❌

Cost Comparison (Relative)

Data Type Storage Cost Query Cost Retention
Metrics Low Low Long (1-2 years)
Logs High Medium Medium (30-90 days)
Traces Very High High Short (7-30 days)

SRE Golden Signals

  1. Latency: Time to service requests
  2. Traffic: Request rate
  3. Errors: Rate of failed requests
  4. Saturation: Resource utilization

Debugging Workflow

  1. Alert fires → Check metrics dashboards
  2. Identify anomaly → Query relevant logs
  3. Find affected requests → Examine traces
  4. Correlate data → Root cause analysis
  5. Fix & verify → Monitor metrics

💬 Interview Tips

  1. Start with requirements: Data volume, retention, query patterns
  2. Consider trade-offs: Cost vs granularity, real-time vs batch
  3. Think about failures: What happens when observability fails?
  4. Mention standards: OpenTelemetry, Prometheus format
  5. Discuss evolution: How to migrate from existing systems

📚 Advanced Topics

Distributed Context

# OpenTelemetry example
from opentelemetry import trace, baggage

tracer = trace.get_tracer(__name__)

with tracer.start_as_current_span("process_request") as span:
    # Add baggage for cross-cutting concerns
    baggage.set_baggage("user.id", "12345")
    baggage.set_baggage("tenant.id", "acme-corp")

Custom Metrics

# Prometheus client example
from prometheus_client import Counter, Histogram, Gauge

request_count = Counter('http_requests_total', 
                       'Total HTTP requests', 
                       ['method', 'endpoint', 'status'])

request_duration = Histogram('http_request_duration_seconds',
                           'HTTP request latency',
                           ['method', 'endpoint'])

active_connections = Gauge('active_connections',
                         'Number of active connections')

Log Correlation

# Structured logging with trace context
import logging
import json

class JSONFormatter(logging.Formatter):
    def format(self, record):
        log_obj = {
            'timestamp': record.created,
            'level': record.levelname,
            'message': record.getMessage(),
            'trace_id': getattr(record, 'trace_id', None),
            'span_id': getattr(record, 'span_id', None),
            'service': 'user-service'
        }
        return json.dumps(log_obj)

🎪 Real-World Scenarios

Scenario 1: E-commerce Platform

  • Metrics: Orders/min, cart abandonment rate, payment success rate
  • Logs: Order events, payment gateway responses, inventory updates
  • Traces: Complete order flow from cart to delivery

Scenario 2: Video Streaming Service

  • Metrics: Concurrent viewers, buffer ratio, bitrate distribution
  • Logs: Player events, CDN errors, quality switches
  • Traces: Video request from player to CDN to origin

Scenario 3: Financial Trading System

  • Metrics: Trade latency percentiles, order book depth, fill rate
  • Logs: All trades (regulatory compliance), system events
  • Traces: Order flow from submission to execution

Remember for Interviews

  1. No solution fits all: Tailor to specific requirements
  2. Start simple: MVP then iterate
  3. Think holistically: Observability is more than just tools
  4. Consider the human factor: Who will use this? How?
  5. Budget matters: Discuss cost implications
  6. Compliance: GDPR, data residency, audit requirements

Golden Rule: Observability should answer: "What's happening?", "Why is it happening?", and "What should I do about it?"

Found this useful? Pass it on.
Pro · $10/mo

The sheet is free. Pro goes deeper.

Pro opens the full question library behind every sheet, every refresher and a monthly AI allowance. One subscription, all formats.

Full question library All refreshers Cancel anytime