Summary
Master microservices architecture for system design interviews. Covers service decomposition, communication patterns, data management, resilience, and deployment strategies. Essential for designing scalable, maintainable distributed systems used by companies like Netflix, Amazon, and Uber.
1. Core Concepts
Definition
- Microservices: Architectural style where applications are built as a collection of small, autonomous services
- Each service is independently deployable, scalable, and owns its data
- Services communicate through well-defined APIs
Microservices vs Monolithic
| Aspect | Monolithic | Microservices |
|---|---|---|
| Deployment | Single unit | Independent services |
| Scaling | Scale entire app | Scale specific services |
| Technology | Single stack | Polyglot |
| Data | Shared database | Service-owned databases |
| Failure Impact | Can affect entire app | Isolated failures |
2. Key Principles
Single Responsibility
- One service = one business capability
- Example: User Service, Order Service, Payment Service
Autonomous Teams
- Each team owns their service end-to-end
- Decentralized governance and decision-making
Design for Failure
- Services can fail independently
- Build resilient systems with fallbacks
Decentralized Data Management
- Each service manages its own database
- No shared database between services
3. Communication Patterns
Synchronous Communication
REST APIs
# Service A calling Service B
import requests
response = requests.get('http://service-b/api/users/123')
user = response.json()
gRPC
// user.proto
service UserService {
rpc GetUser(UserRequest) returns (User) {}
}
Asynchronous Communication
Message Queues (RabbitMQ, SQS)
# Publisher
import pika
channel.basic_publish(exchange='orders',
routing_key='order.created',
body=json.dumps(order_data))
# Consumer
def callback(ch, method, properties, body):
order = json.loads(body)
process_order(order)
Event Streaming (Kafka)
# Producer
producer.send('order-events', value=order_event)
# Consumer
for message in consumer:
process_event(message.value)
4. Service Discovery
Client-Side Discovery
- Client queries service registry
- Examples: Netflix Eureka
Server-Side Discovery
- Load balancer queries registry
- Examples: AWS ELB, Kubernetes Services
# Kubernetes Service
apiVersion: v1
kind: Service
metadata:
name: user-service
spec:
selector:
app: user
ports:
- port: 80
5. API Gateway Pattern
Benefits
- Single entry point for clients
- Authentication/authorization
- Rate limiting
- Request routing
- Protocol translation
Implementation
// Simple API Gateway
app.get('/api/orders/:id', async (req, res) => {
const order = await orderService.getOrder(req.params.id);
const user = await userService.getUser(order.userId);
res.json({ order, user });
});
6. Data Management Patterns
Database per Service
- Each service has its own database
- Can use different database types (polyglot persistence)
Saga Pattern
Manages distributed transactions across services
Choreography-based Saga
Order Service -> (Order Created Event) -> Payment Service
-> (Payment Processed) -> Shipping Service
Orchestration-based Saga
class OrderSaga:
def execute(self, order):
try:
payment = payment_service.process(order)
shipping = shipping_service.create(order)
order_service.complete(order)
except:
# Compensating transactions
payment_service.refund(payment)
shipping_service.cancel(shipping)
CQRS (Command Query Responsibility Segregation)
- Separate read and write models
- Optimized for different use cases
7. Resilience Patterns
Circuit Breaker
from pybreaker import CircuitBreaker
db_breaker = CircuitBreaker(fail_max=5, reset_timeout=60)
@db_breaker
def call_service():
return external_service.call()
Retry with Exponential Backoff
import time
def retry_with_backoff(func, max_retries=3):
for i in range(max_retries):
try:
return func()
except Exception as e:
if i == max_retries - 1:
raise
time.sleep(2 ** i)
Bulkhead Pattern
- Isolate resources to prevent cascade failures
- Example: Separate thread pools for different services
Timeout Pattern
import requests
response = requests.get('http://service-b/api/data', timeout=5)
8. Security
Service-to-Service Authentication
- mTLS: Mutual TLS for service authentication
- JWT Tokens: For API authentication
# JWT validation
from jose import jwt
def validate_token(token):
payload = jwt.decode(token, SECRET_KEY, algorithms=['HS256'])
return payload
API Key Management
# API Gateway validation
def validate_api_key(request):
api_key = request.headers.get('X-API-Key')
if not is_valid_key(api_key):
raise Unauthorized()
9. Monitoring & Observability
The Three Pillars
1. Metrics
- Response time, throughput, error rate
- Tools: Prometheus, Grafana
2. Logging
- Centralized logging with correlation IDs
- Tools: ELK Stack, Splunk
# Correlation ID
import uuid
def add_correlation_id(request):
request.correlation_id = request.headers.get('X-Correlation-ID',
str(uuid.uuid4()))
3. Distributed Tracing
- Track requests across services
- Tools: Jaeger, Zipkin
10. Deployment Strategies
Blue-Green Deployment
- Two identical environments
- Switch traffic between them
Canary Deployment
# Kubernetes canary deployment
spec:
replicas: 10
strategy:
canary:
steps:
- setWeight: 10 # 10% traffic
- pause: {duration: 10m}
- setWeight: 50 # 50% traffic
- pause: {duration: 10m}
- setWeight: 100 # 100% traffic
Feature Flags
if feature_flag.is_enabled('new_payment_flow'):
return new_payment_process()
else:
return legacy_payment_process()
11. Interview Focus Areas
System Design Scenarios
- URL Shortener: Service separation, caching strategy, analytics service
- E-commerce Platform: Order, Inventory, Payment, Notification services
- Ride-sharing System: User, Driver, Trip, Payment, Location services
- Social Media Platform: Post, User, Timeline, Notification services
Critical Architecture Concepts
- Distributed Transactions: Saga pattern, two-phase commit, eventual consistency
- Data Consistency: CQRS, event sourcing, distributed locks
- Service Failures: Circuit breakers, retries, bulkheads, timeouts
- Cross-Service Authentication: JWT, OAuth, service mesh, API gateway
12. Best Practices
Do's
- Start with a modular monolith
- Define clear service boundaries
- Implement comprehensive monitoring
- Use API versioning
- Implement health checks
- Document APIs thoroughly
- Use containers (Docker/Kubernetes)
Don'ts
- Create too many fine-grained services
- Share databases between services
- Ignore distributed system complexity
- Forget about data consistency
- Overlook security between services
13. Anti-Patterns
Distributed Monolith
- Services too tightly coupled
- Must deploy services together
Chatty Services
- Too many synchronous calls between services
- Solution: Aggregate data, use async communication
Shared Database
- Multiple services using same database
- Solution: Database per service pattern
14. Technology Stack
Container Orchestration
- Kubernetes: De facto standard
- Docker Swarm: Simpler alternative
- Amazon ECS: AWS managed solution
Service Mesh
- Istio: Traffic management, security
- Linkerd: Lightweight option
- Consul Connect: HashiCorp solution
Message Brokers
- Kafka: Event streaming
- RabbitMQ: Traditional message queue
- AWS SQS/SNS: Managed solutions
15. Migration Strategy
Strangler Fig Pattern
1. Identify bounded contexts
2. Create new microservice
3. Redirect traffic gradually
4. Decommission old code
Branch by Abstraction
# Abstract interface
class PaymentProcessor:
def process(self, amount):
if use_new_service():
return new_payment_service.process(amount)
else:
return legacy_payment.process(amount)
16. Performance Considerations
Caching Strategies
- Service-level caching: Redis, Memcached
- CDN: For static content
- API Gateway caching: Response caching
Optimization Techniques
# Batch API calls
users = user_service.get_users_batch(user_ids)
# Parallel calls
import asyncio
async def get_data():
user, order = await asyncio.gather(
get_user(user_id),
get_order(order_id)
)
return user, order
17. Testing Strategies
Testing Pyramid
- Unit Tests: Individual service logic
- Integration Tests: Service interactions
- Contract Tests: API contracts
- End-to-End Tests: Complete workflows
Consumer-Driven Contract Testing
# Pact example
@provider_state("user 123 exists")
def test_get_user():
response = user_service.get_user(123)
assert response.status_code == 200
assert response.json()["id"] == 123
Quick Reference - Key Takeaways
- Service Boundaries: Define based on business capabilities
- Communication: Prefer async over sync when possible
- Data: Each service owns its data
- Resilience: Design for failure with circuit breakers, retries
- Monitoring: Implement distributed tracing early
- Security: Never trust service-to-service communication
- Deployment: Use containers and orchestration
- Testing: Focus on contract testing between services
- Performance: Cache aggressively, minimize network calls
- Migration: Start small, use strangler fig pattern
Interview Success Tips
Must-Mention Topics
- Service discovery mechanisms
- Distributed transaction handling
- Monitoring and observability strategy
- Inter-service security
- Data consistency approaches
- API versioning strategy
- Failure handling patterns
Key Trade-offs to Discuss
- Complexity vs Modularity: More services = more operational overhead
- Consistency vs Performance: Strong consistency impacts latency
- Sync vs Async: Trade-offs between simplicity and resilience
- Centralized vs Distributed: Gateway vs service mesh