Summary
Master load balancing concepts for system design interviews. Covers algorithms, Layer 4 vs Layer 7 balancing, health checks, session persistence, and real-world implementations. Critical for designing scalable, highly available systems that can handle millions of requests.
Overview
Load balancing distributes incoming network traffic across multiple servers to ensure no single server bears too much demand.
Why Load Balancing?
- High Availability: Eliminates single points of failure
- Scalability: Handle more traffic by adding servers
- Performance: Reduced response time and increased throughput
- Flexibility: Maintenance without downtime
Types of Load Balancers
1. Hardware Load Balancers
- Physical devices (F5, Citrix NetScaler)
- High performance but expensive
- Limited flexibility
2. Software Load Balancers
- HAProxy, NGINX, Apache
- Cost-effective and flexible
- Can run on commodity hardware
3. Cloud Load Balancers
- AWS ELB, Google Cloud Load Balancer, Azure Load Balancer
- Managed service, auto-scaling
- Pay-per-use model
Load Balancing Algorithms
1. Round Robin
class RoundRobinLB:
def __init__(self, servers):
self.servers = servers
self.current = 0
def get_server(self):
server = self.servers[self.current]
self.current = (self.current + 1) % len(self.servers)
return server
Use Case: When servers have similar capabilities
2. Weighted Round Robin
class WeightedRoundRobinLB:
def __init__(self, servers):
self.servers = servers # [(server, weight), ...]
self.current_weights = [0] * len(servers)
self.total_weight = sum(w for _, w in servers)
def get_server(self):
max_weight_index = 0
for i, (server, weight) in enumerate(self.servers):
self.current_weights[i] += weight
if self.current_weights[i] > self.current_weights[max_weight_index]:
max_weight_index = i
self.current_weights[max_weight_index] -= self.total_weight
return self.servers[max_weight_index][0]
Use Case: When servers have different capacities
3. Least Connections
class LeastConnectionsLB:
def __init__(self, servers):
self.servers = {server: 0 for server in servers}
def get_server(self):
return min(self.servers, key=self.servers.get)
def connect(self, server):
self.servers[server] += 1
def disconnect(self, server):
self.servers[server] = max(0, self.servers[server] - 1)
Use Case: When requests have varying processing times
4. Least Response Time
- Routes to server with fastest response time
- Considers both active connections and response time
5. IP Hash
import hashlib
class IPHashLB:
def __init__(self, servers):
self.servers = servers
def get_server(self, client_ip):
hash_val = int(hashlib.md5(client_ip.encode()).hexdigest(), 16)
return self.servers[hash_val % len(self.servers)]
Use Case: When session persistence is needed
6. Consistent Hashing
import hashlib
class ConsistentHashLB:
def __init__(self, servers, virtual_nodes=150):
self.servers = servers
self.virtual_nodes = virtual_nodes
self.ring = {}
self._build_ring()
def _hash(self, key):
return int(hashlib.md5(key.encode()).hexdigest(), 16)
def _build_ring(self):
for server in self.servers:
for i in range(self.virtual_nodes):
virtual_key = f"{server}:{i}"
hash_val = self._hash(virtual_key)
self.ring[hash_val] = server
def get_server(self, key):
if not self.ring:
return None
hash_val = self._hash(key)
sorted_keys = sorted(self.ring.keys())
for ring_key in sorted_keys:
if hash_val <= ring_key:
return self.ring[ring_key]
return self.ring[sorted_keys[0]]
Use Case: Distributed caching systems
Load Balancer Placement
1. Layer 4 (Transport Layer)
- Works with IP addresses and ports
- Cannot inspect application data
- Faster, less CPU intensive
- Example: TCP/UDP load balancing
2. Layer 7 (Application Layer)
- Can inspect HTTP headers, URLs, cookies
- Content-based routing possible
- More CPU intensive
- Example: HTTP/HTTPS load balancing
# NGINX Layer 7 Load Balancing Example
upstream backend {
server backend1.example.com weight=5;
server backend2.example.com;
server backend3.example.com;
}
server {
location /api/ {
proxy_pass http://backend;
}
location /static/ {
proxy_pass http://static-servers;
}
}
Health Checks
Active Health Checks
import requests
import time
class HealthChecker:
def __init__(self, servers, check_interval=30):
self.servers = servers
self.check_interval = check_interval
self.healthy_servers = set(servers)
def check_health(self, server):
try:
response = requests.get(f"http://{server}/health", timeout=5)
return response.status_code == 200
except:
return False
def update_healthy_servers(self):
for server in self.servers:
if self.check_health(server):
self.healthy_servers.add(server)
else:
self.healthy_servers.discard(server)
Passive Health Checks
- Monitor actual traffic responses
- Mark server unhealthy after N failures
- No additional health check traffic
Session Persistence (Sticky Sessions)
Methods:
- Cookie-based: Insert server ID in cookie
- IP-based: Route same IP to same server
- Session ID: Use application session ID
Drawbacks:
- Uneven load distribution
- Complex failover handling
- Limits horizontal scaling
Advanced Concepts
1. Global Server Load Balancing (GSLB)
- DNS-based load balancing across data centers
- Geo-routing capabilities
- Disaster recovery
2. Auto-scaling Integration
# AWS Auto Scaling with ELB
AutoScalingGroup:
MinSize: 2
MaxSize: 10
TargetGroupARNs:
- !Ref TargetGroup
HealthCheckType: ELB
HealthCheckGracePeriod: 300
3. Circuit Breaker Pattern
class CircuitBreaker:
def __init__(self, failure_threshold=5, timeout=60):
self.failure_threshold = failure_threshold
self.timeout = timeout
self.failures = 0
self.last_failure_time = None
self.state = 'CLOSED' # CLOSED, OPEN, HALF_OPEN
def call(self, func, *args, **kwargs):
if self.state == 'OPEN':
if time.time() - self.last_failure_time > self.timeout:
self.state = 'HALF_OPEN'
else:
raise Exception("Circuit breaker is OPEN")
try:
result = func(*args, **kwargs)
if self.state == 'HALF_OPEN':
self.state = 'CLOSED'
self.failures = 0
return result
except Exception as e:
self.failures += 1
self.last_failure_time = time.time()
if self.failures >= self.failure_threshold:
self.state = 'OPEN'
raise e
4. Rate Limiting
from collections import defaultdict
import time
class RateLimiter:
def __init__(self, max_requests=100, window_seconds=60):
self.max_requests = max_requests
self.window_seconds = window_seconds
self.requests = defaultdict(list)
def is_allowed(self, client_id):
now = time.time()
# Remove old requests
self.requests[client_id] = [
req_time for req_time in self.requests[client_id]
if now - req_time < self.window_seconds
]
if len(self.requests[client_id]) < self.max_requests:
self.requests[client_id].append(now)
return True
return False
Performance Metrics
Key Metrics to Monitor:
- Response Time: Average, P95, P99
- Throughput: Requests per second
- Error Rate: 4xx, 5xx responses
- Connection Count: Active connections per server
- CPU/Memory Usage: Resource utilization
- Queue Length: Pending requests
Critical Interview Topics
Scaling to Millions of Users
- Multi-layer Load Balancing: DNS → L4 → L7 hierarchy
- Geographic Distribution: GSLB with regional load balancers
- CDN Integration: Offload static content
- Caching Layers: Application, database, CDN caches
- Database Strategies: Read replicas, sharding, partitioning
Zero Downtime Deployments
- Blue-Green Deployment: Instant switchover between environments
- Rolling Updates: Gradual replacement of instances
- Canary Releases: Test with small traffic percentage
- Health Check Gates: Verify before routing traffic
High Availability Design
- Active-Active Configuration: Multiple active load balancers
- Floating IP/VRRP: Automatic failover
- DNS Failover: Multiple A records or health-based routing
- Cross-Region Redundancy: Geographic distribution
SSL/TLS Strategies
- SSL Termination: Decrypt at load balancer (performance)
- SSL Passthrough: End-to-end encryption (security)
- SSL Bridging: Re-encrypt after inspection
Design Considerations
1. Capacity Planning
- Peak traffic estimation
- Growth projections
- Buffer capacity (N+1 redundancy)
2. Cost Optimization
- Right-sizing instances
- Auto-scaling policies
- Reserved capacity for baseline
3. Security
- DDoS protection
- Web Application Firewall (WAF)
- IP whitelisting/blacklisting
- SSL/TLS configuration
Best Practices
- Start Simple: Begin with round-robin, optimize later
- Monitor Everything: You can't improve what you don't measure
- Plan for Failure: Design for N+1 redundancy
- Test Load Balancing: Regular failover drills
- Document Configuration: Keep configs in version control
- Consider Caching: Reduce backend load
- Use Health Checks: Both active and passive
- Implement Graceful Shutdown: Drain connections before removal
Quick Implementation Examples
HAProxy Configuration
global
maxconn 4096
defaults
mode http
timeout connect 5000ms
timeout client 50000ms
timeout server 50000ms
backend servers
balance roundrobin
option httpchk GET /health
server server1 192.168.1.10:80 check
server server2 192.168.1.11:80 check
NGINX Configuration
http {
upstream myapp {
least_conn;
server srv1.example.com;
server srv2.example.com;
server srv3.example.com down;
server srv4.example.com backup;
}
server {
listen 80;
location / {
proxy_pass http://myapp;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
}
Interview Strategy
- Clarify Requirements: Ask about scale, budget, existing infrastructure
- Start High-Level: Draw the architecture before diving into details
- Consider Trade-offs: Discuss pros/cons of each approach
- Think About Edge Cases: What happens during failures?
- Mention Monitoring: Show you think about operations
- Be Realistic: Don't over-engineer for the given scale
Key Takeaways
- Algorithm Selection: Match to traffic patterns (RR for uniform, LC for varied)
- Layer Choice: L4 for speed, L7 for intelligence
- Health Checks: Critical for reliability, use both active and passive
- Session Handling: Avoid sticky sessions when possible
- Monitoring: Track latency percentiles, not just averages
- Failure Planning: N+1 redundancy minimum
- Security: Consider DDoS, SSL termination, WAF integration
- Cloud vs On-Premise: Managed services reduce operational overhead