LearnThatStack Ace your next interview

Microservices System Design.
Interview cheat sheet.

Quick reference for Microservices System Design - sectioned for fast scanning. Skim the part you're shaky on, walk in confident.

System Design Concepts 20-section reference ~5 min read

Summary

Master microservices architecture for system design interviews. Covers service decomposition, communication patterns, data management, resilience, and deployment strategies. Essential for designing scalable, maintainable distributed systems used by companies like Netflix, Amazon, and Uber.

1. Core Concepts

Definition

  • Microservices: Architectural style where applications are built as a collection of small, autonomous services
  • Each service is independently deployable, scalable, and owns its data
  • Services communicate through well-defined APIs

Microservices vs Monolithic

Aspect Monolithic Microservices
Deployment Single unit Independent services
Scaling Scale entire app Scale specific services
Technology Single stack Polyglot
Data Shared database Service-owned databases
Failure Impact Can affect entire app Isolated failures

2. Key Principles

Single Responsibility

  • One service = one business capability
  • Example: User Service, Order Service, Payment Service

Autonomous Teams

  • Each team owns their service end-to-end
  • Decentralized governance and decision-making

Design for Failure

  • Services can fail independently
  • Build resilient systems with fallbacks

Decentralized Data Management

  • Each service manages its own database
  • No shared database between services

3. Communication Patterns

Synchronous Communication

REST APIs

# Service A calling Service B
import requests

response = requests.get('http://service-b/api/users/123')
user = response.json()

gRPC

// user.proto
service UserService {
  rpc GetUser(UserRequest) returns (User) {}
}

Asynchronous Communication

Message Queues (RabbitMQ, SQS)

# Publisher
import pika
channel.basic_publish(exchange='orders', 
                     routing_key='order.created',
                     body=json.dumps(order_data))

# Consumer
def callback(ch, method, properties, body):
    order = json.loads(body)
    process_order(order)

Event Streaming (Kafka)

# Producer
producer.send('order-events', value=order_event)

# Consumer
for message in consumer:
    process_event(message.value)

4. Service Discovery

Client-Side Discovery

  • Client queries service registry
  • Examples: Netflix Eureka

Server-Side Discovery

  • Load balancer queries registry
  • Examples: AWS ELB, Kubernetes Services
# Kubernetes Service
apiVersion: v1
kind: Service
metadata:
  name: user-service
spec:
  selector:
    app: user
  ports:
    - port: 80

5. API Gateway Pattern

Benefits

  • Single entry point for clients
  • Authentication/authorization
  • Rate limiting
  • Request routing
  • Protocol translation

Implementation

// Simple API Gateway
app.get('/api/orders/:id', async (req, res) => {
  const order = await orderService.getOrder(req.params.id);
  const user = await userService.getUser(order.userId);
  res.json({ order, user });
});

6. Data Management Patterns

Database per Service

  • Each service has its own database
  • Can use different database types (polyglot persistence)

Saga Pattern

Manages distributed transactions across services

Choreography-based Saga

Order Service -> (Order Created Event) -> Payment Service
                                      -> (Payment Processed) -> Shipping Service

Orchestration-based Saga

class OrderSaga:
    def execute(self, order):
        try:
            payment = payment_service.process(order)
            shipping = shipping_service.create(order)
            order_service.complete(order)
        except:
            # Compensating transactions
            payment_service.refund(payment)
            shipping_service.cancel(shipping)

CQRS (Command Query Responsibility Segregation)

  • Separate read and write models
  • Optimized for different use cases

7. Resilience Patterns

Circuit Breaker

from pybreaker import CircuitBreaker

db_breaker = CircuitBreaker(fail_max=5, reset_timeout=60)

@db_breaker
def call_service():
    return external_service.call()

Retry with Exponential Backoff

import time

def retry_with_backoff(func, max_retries=3):
    for i in range(max_retries):
        try:
            return func()
        except Exception as e:
            if i == max_retries - 1:
                raise
            time.sleep(2 ** i)

Bulkhead Pattern

  • Isolate resources to prevent cascade failures
  • Example: Separate thread pools for different services

Timeout Pattern

import requests

response = requests.get('http://service-b/api/data', timeout=5)

8. Security

Service-to-Service Authentication

  • mTLS: Mutual TLS for service authentication
  • JWT Tokens: For API authentication
# JWT validation
from jose import jwt

def validate_token(token):
    payload = jwt.decode(token, SECRET_KEY, algorithms=['HS256'])
    return payload

API Key Management

# API Gateway validation
def validate_api_key(request):
    api_key = request.headers.get('X-API-Key')
    if not is_valid_key(api_key):
        raise Unauthorized()

9. Monitoring & Observability

The Three Pillars

1. Metrics

  • Response time, throughput, error rate
  • Tools: Prometheus, Grafana

2. Logging

  • Centralized logging with correlation IDs
  • Tools: ELK Stack, Splunk
# Correlation ID
import uuid

def add_correlation_id(request):
    request.correlation_id = request.headers.get('X-Correlation-ID', 
                                                str(uuid.uuid4()))

3. Distributed Tracing

  • Track requests across services
  • Tools: Jaeger, Zipkin

10. Deployment Strategies

Blue-Green Deployment

  • Two identical environments
  • Switch traffic between them

Canary Deployment

# Kubernetes canary deployment
spec:
  replicas: 10
  strategy:
    canary:
      steps:
      - setWeight: 10  # 10% traffic
      - pause: {duration: 10m}
      - setWeight: 50  # 50% traffic
      - pause: {duration: 10m}
      - setWeight: 100 # 100% traffic

Feature Flags

if feature_flag.is_enabled('new_payment_flow'):
    return new_payment_process()
else:
    return legacy_payment_process()

11. Interview Focus Areas

System Design Scenarios

  • URL Shortener: Service separation, caching strategy, analytics service
  • E-commerce Platform: Order, Inventory, Payment, Notification services
  • Ride-sharing System: User, Driver, Trip, Payment, Location services
  • Social Media Platform: Post, User, Timeline, Notification services

Critical Architecture Concepts

  • Distributed Transactions: Saga pattern, two-phase commit, eventual consistency
  • Data Consistency: CQRS, event sourcing, distributed locks
  • Service Failures: Circuit breakers, retries, bulkheads, timeouts
  • Cross-Service Authentication: JWT, OAuth, service mesh, API gateway

12. Best Practices

Do's

  • Start with a modular monolith
  • Define clear service boundaries
  • Implement comprehensive monitoring
  • Use API versioning
  • Implement health checks
  • Document APIs thoroughly
  • Use containers (Docker/Kubernetes)

Don'ts

  • Create too many fine-grained services
  • Share databases between services
  • Ignore distributed system complexity
  • Forget about data consistency
  • Overlook security between services

13. Anti-Patterns

Distributed Monolith

  • Services too tightly coupled
  • Must deploy services together

Chatty Services

  • Too many synchronous calls between services
  • Solution: Aggregate data, use async communication

Shared Database

  • Multiple services using same database
  • Solution: Database per service pattern

14. Technology Stack

Container Orchestration

  • Kubernetes: De facto standard
  • Docker Swarm: Simpler alternative
  • Amazon ECS: AWS managed solution

Service Mesh

  • Istio: Traffic management, security
  • Linkerd: Lightweight option
  • Consul Connect: HashiCorp solution

Message Brokers

  • Kafka: Event streaming
  • RabbitMQ: Traditional message queue
  • AWS SQS/SNS: Managed solutions

15. Migration Strategy

Strangler Fig Pattern

1. Identify bounded contexts
2. Create new microservice
3. Redirect traffic gradually
4. Decommission old code

Branch by Abstraction

# Abstract interface
class PaymentProcessor:
    def process(self, amount):
        if use_new_service():
            return new_payment_service.process(amount)
        else:
            return legacy_payment.process(amount)

16. Performance Considerations

Caching Strategies

  • Service-level caching: Redis, Memcached
  • CDN: For static content
  • API Gateway caching: Response caching

Optimization Techniques

# Batch API calls
users = user_service.get_users_batch(user_ids)

# Parallel calls
import asyncio

async def get_data():
    user, order = await asyncio.gather(
        get_user(user_id),
        get_order(order_id)
    )
    return user, order

17. Testing Strategies

Testing Pyramid

  1. Unit Tests: Individual service logic
  2. Integration Tests: Service interactions
  3. Contract Tests: API contracts
  4. End-to-End Tests: Complete workflows

Consumer-Driven Contract Testing

# Pact example
@provider_state("user 123 exists")
def test_get_user():
    response = user_service.get_user(123)
    assert response.status_code == 200
    assert response.json()["id"] == 123

Quick Reference - Key Takeaways

  1. Service Boundaries: Define based on business capabilities
  2. Communication: Prefer async over sync when possible
  3. Data: Each service owns its data
  4. Resilience: Design for failure with circuit breakers, retries
  5. Monitoring: Implement distributed tracing early
  6. Security: Never trust service-to-service communication
  7. Deployment: Use containers and orchestration
  8. Testing: Focus on contract testing between services
  9. Performance: Cache aggressively, minimize network calls
  10. Migration: Start small, use strangler fig pattern

Interview Success Tips

Must-Mention Topics

  • Service discovery mechanisms
  • Distributed transaction handling
  • Monitoring and observability strategy
  • Inter-service security
  • Data consistency approaches
  • API versioning strategy
  • Failure handling patterns

Key Trade-offs to Discuss

  • Complexity vs Modularity: More services = more operational overhead
  • Consistency vs Performance: Strong consistency impacts latency
  • Sync vs Async: Trade-offs between simplicity and resilience
  • Centralized vs Distributed: Gateway vs service mesh
Found this useful? Pass it on.
Pro · $10/mo

The sheet is free. Pro goes deeper.

Pro opens the full question library behind every sheet, every refresher and a monthly AI allowance. One subscription, all formats.

Full question library All refreshers Cancel anytime