LearnThatStack Ace your next interview

Load Balancing System Design.
Cheat sheet.

Quick reference for Load Balancing System Design - sectioned for fast scanning. Skim the part you're shaky on, walk in confident.

System Design Concepts 15-section reference ~5 min read

Summary

Master load balancing concepts for system design interviews. Covers algorithms, Layer 4 vs Layer 7 balancing, health checks, session persistence, and real-world implementations. Critical for designing scalable, highly available systems that can handle millions of requests.

Overview

Load balancing distributes incoming network traffic across multiple servers to ensure no single server bears too much demand.

Why Load Balancing?

  • High Availability: Eliminates single points of failure
  • Scalability: Handle more traffic by adding servers
  • Performance: Reduced response time and increased throughput
  • Flexibility: Maintenance without downtime

Types of Load Balancers

1. Hardware Load Balancers

  • Physical devices (F5, Citrix NetScaler)
  • High performance but expensive
  • Limited flexibility

2. Software Load Balancers

  • HAProxy, NGINX, Apache
  • Cost-effective and flexible
  • Can run on commodity hardware

3. Cloud Load Balancers

  • AWS ELB, Google Cloud Load Balancer, Azure Load Balancer
  • Managed service, auto-scaling
  • Pay-per-use model

Load Balancing Algorithms

1. Round Robin

class RoundRobinLB:
    def __init__(self, servers):
        self.servers = servers
        self.current = 0
    
    def get_server(self):
        server = self.servers[self.current]
        self.current = (self.current + 1) % len(self.servers)
        return server

Use Case: When servers have similar capabilities

2. Weighted Round Robin

class WeightedRoundRobinLB:
    def __init__(self, servers):
        self.servers = servers  # [(server, weight), ...]
        self.current_weights = [0] * len(servers)
        self.total_weight = sum(w for _, w in servers)
    
    def get_server(self):
        max_weight_index = 0
        for i, (server, weight) in enumerate(self.servers):
            self.current_weights[i] += weight
            if self.current_weights[i] > self.current_weights[max_weight_index]:
                max_weight_index = i
        
        self.current_weights[max_weight_index] -= self.total_weight
        return self.servers[max_weight_index][0]

Use Case: When servers have different capacities

3. Least Connections

class LeastConnectionsLB:
    def __init__(self, servers):
        self.servers = {server: 0 for server in servers}
    
    def get_server(self):
        return min(self.servers, key=self.servers.get)
    
    def connect(self, server):
        self.servers[server] += 1
    
    def disconnect(self, server):
        self.servers[server] = max(0, self.servers[server] - 1)

Use Case: When requests have varying processing times

4. Least Response Time

  • Routes to server with fastest response time
  • Considers both active connections and response time

5. IP Hash

import hashlib

class IPHashLB:
    def __init__(self, servers):
        self.servers = servers
    
    def get_server(self, client_ip):
        hash_val = int(hashlib.md5(client_ip.encode()).hexdigest(), 16)
        return self.servers[hash_val % len(self.servers)]

Use Case: When session persistence is needed

6. Consistent Hashing

import hashlib

class ConsistentHashLB:
    def __init__(self, servers, virtual_nodes=150):
        self.servers = servers
        self.virtual_nodes = virtual_nodes
        self.ring = {}
        self._build_ring()
    
    def _hash(self, key):
        return int(hashlib.md5(key.encode()).hexdigest(), 16)
    
    def _build_ring(self):
        for server in self.servers:
            for i in range(self.virtual_nodes):
                virtual_key = f"{server}:{i}"
                hash_val = self._hash(virtual_key)
                self.ring[hash_val] = server
    
    def get_server(self, key):
        if not self.ring:
            return None
        
        hash_val = self._hash(key)
        sorted_keys = sorted(self.ring.keys())
        
        for ring_key in sorted_keys:
            if hash_val <= ring_key:
                return self.ring[ring_key]
        
        return self.ring[sorted_keys[0]]

Use Case: Distributed caching systems

Load Balancer Placement

1. Layer 4 (Transport Layer)

  • Works with IP addresses and ports
  • Cannot inspect application data
  • Faster, less CPU intensive
  • Example: TCP/UDP load balancing

2. Layer 7 (Application Layer)

  • Can inspect HTTP headers, URLs, cookies
  • Content-based routing possible
  • More CPU intensive
  • Example: HTTP/HTTPS load balancing
# NGINX Layer 7 Load Balancing Example
upstream backend {
    server backend1.example.com weight=5;
    server backend2.example.com;
    server backend3.example.com;
}

server {
    location /api/ {
        proxy_pass http://backend;
    }
    
    location /static/ {
        proxy_pass http://static-servers;
    }
}

Health Checks

Active Health Checks

import requests
import time

class HealthChecker:
    def __init__(self, servers, check_interval=30):
        self.servers = servers
        self.check_interval = check_interval
        self.healthy_servers = set(servers)
    
    def check_health(self, server):
        try:
            response = requests.get(f"http://{server}/health", timeout=5)
            return response.status_code == 200
        except:
            return False
    
    def update_healthy_servers(self):
        for server in self.servers:
            if self.check_health(server):
                self.healthy_servers.add(server)
            else:
                self.healthy_servers.discard(server)

Passive Health Checks

  • Monitor actual traffic responses
  • Mark server unhealthy after N failures
  • No additional health check traffic

Session Persistence (Sticky Sessions)

Methods:

  1. Cookie-based: Insert server ID in cookie
  2. IP-based: Route same IP to same server
  3. Session ID: Use application session ID

Drawbacks:

  • Uneven load distribution
  • Complex failover handling
  • Limits horizontal scaling

Advanced Concepts

1. Global Server Load Balancing (GSLB)

  • DNS-based load balancing across data centers
  • Geo-routing capabilities
  • Disaster recovery

2. Auto-scaling Integration

# AWS Auto Scaling with ELB
AutoScalingGroup:
  MinSize: 2
  MaxSize: 10
  TargetGroupARNs:
    - !Ref TargetGroup
  HealthCheckType: ELB
  HealthCheckGracePeriod: 300

3. Circuit Breaker Pattern

class CircuitBreaker:
    def __init__(self, failure_threshold=5, timeout=60):
        self.failure_threshold = failure_threshold
        self.timeout = timeout
        self.failures = 0
        self.last_failure_time = None
        self.state = 'CLOSED'  # CLOSED, OPEN, HALF_OPEN
    
    def call(self, func, *args, **kwargs):
        if self.state == 'OPEN':
            if time.time() - self.last_failure_time > self.timeout:
                self.state = 'HALF_OPEN'
            else:
                raise Exception("Circuit breaker is OPEN")
        
        try:
            result = func(*args, **kwargs)
            if self.state == 'HALF_OPEN':
                self.state = 'CLOSED'
                self.failures = 0
            return result
        except Exception as e:
            self.failures += 1
            self.last_failure_time = time.time()
            if self.failures >= self.failure_threshold:
                self.state = 'OPEN'
            raise e

4. Rate Limiting

from collections import defaultdict
import time

class RateLimiter:
    def __init__(self, max_requests=100, window_seconds=60):
        self.max_requests = max_requests
        self.window_seconds = window_seconds
        self.requests = defaultdict(list)
    
    def is_allowed(self, client_id):
        now = time.time()
        # Remove old requests
        self.requests[client_id] = [
            req_time for req_time in self.requests[client_id]
            if now - req_time < self.window_seconds
        ]
        
        if len(self.requests[client_id]) < self.max_requests:
            self.requests[client_id].append(now)
            return True
        return False

Performance Metrics

Key Metrics to Monitor:

  1. Response Time: Average, P95, P99
  2. Throughput: Requests per second
  3. Error Rate: 4xx, 5xx responses
  4. Connection Count: Active connections per server
  5. CPU/Memory Usage: Resource utilization
  6. Queue Length: Pending requests

Critical Interview Topics

Scaling to Millions of Users

  • Multi-layer Load Balancing: DNS → L4 → L7 hierarchy
  • Geographic Distribution: GSLB with regional load balancers
  • CDN Integration: Offload static content
  • Caching Layers: Application, database, CDN caches
  • Database Strategies: Read replicas, sharding, partitioning

Zero Downtime Deployments

  • Blue-Green Deployment: Instant switchover between environments
  • Rolling Updates: Gradual replacement of instances
  • Canary Releases: Test with small traffic percentage
  • Health Check Gates: Verify before routing traffic

High Availability Design

  • Active-Active Configuration: Multiple active load balancers
  • Floating IP/VRRP: Automatic failover
  • DNS Failover: Multiple A records or health-based routing
  • Cross-Region Redundancy: Geographic distribution

SSL/TLS Strategies

  • SSL Termination: Decrypt at load balancer (performance)
  • SSL Passthrough: End-to-end encryption (security)
  • SSL Bridging: Re-encrypt after inspection

Design Considerations

1. Capacity Planning

  • Peak traffic estimation
  • Growth projections
  • Buffer capacity (N+1 redundancy)

2. Cost Optimization

  • Right-sizing instances
  • Auto-scaling policies
  • Reserved capacity for baseline

3. Security

  • DDoS protection
  • Web Application Firewall (WAF)
  • IP whitelisting/blacklisting
  • SSL/TLS configuration

Best Practices

  1. Start Simple: Begin with round-robin, optimize later
  2. Monitor Everything: You can't improve what you don't measure
  3. Plan for Failure: Design for N+1 redundancy
  4. Test Load Balancing: Regular failover drills
  5. Document Configuration: Keep configs in version control
  6. Consider Caching: Reduce backend load
  7. Use Health Checks: Both active and passive
  8. Implement Graceful Shutdown: Drain connections before removal

Quick Implementation Examples

HAProxy Configuration

global
    maxconn 4096

defaults
    mode http
    timeout connect 5000ms
    timeout client 50000ms
    timeout server 50000ms

backend servers
    balance roundrobin
    option httpchk GET /health
    server server1 192.168.1.10:80 check
    server server2 192.168.1.11:80 check

NGINX Configuration

http {
    upstream myapp {
        least_conn;
        server srv1.example.com;
        server srv2.example.com;
        server srv3.example.com down;
        server srv4.example.com backup;
    }
    
    server {
        listen 80;
        location / {
            proxy_pass http://myapp;
            proxy_set_header Host $host;
            proxy_set_header X-Real-IP $remote_addr;
        }
    }
}

Interview Strategy

  1. Clarify Requirements: Ask about scale, budget, existing infrastructure
  2. Start High-Level: Draw the architecture before diving into details
  3. Consider Trade-offs: Discuss pros/cons of each approach
  4. Think About Edge Cases: What happens during failures?
  5. Mention Monitoring: Show you think about operations
  6. Be Realistic: Don't over-engineer for the given scale

Key Takeaways

  • Algorithm Selection: Match to traffic patterns (RR for uniform, LC for varied)
  • Layer Choice: L4 for speed, L7 for intelligence
  • Health Checks: Critical for reliability, use both active and passive
  • Session Handling: Avoid sticky sessions when possible
  • Monitoring: Track latency percentiles, not just averages
  • Failure Planning: N+1 redundancy minimum
  • Security: Consider DDoS, SSL termination, WAF integration
  • Cloud vs On-Premise: Managed services reduce operational overhead
Found this useful? Pass it on.
Pro · $10/mo

The sheet is free. Pro goes deeper.

Pro opens the full question library behind every sheet, every refresher and a monthly AI allowance. One subscription, all formats.

Full question library All refreshers Cancel anytime