Summary
Redis (Remote Dictionary Server) is an in-memory data structure store used as a database, cache, and message broker. It supports various data structures including strings, hashes, lists, sets, sorted sets, bitmaps, hyperloglogs, and streams. Key features include single-threaded architecture with I/O multiplexing, optional persistence (RDB/AOF), built-in replication, Lua scripting, transactions, and pub/sub messaging. Essential for high-performance caching, session management, real-time analytics, leaderboards, and distributed system coordination.
1. Redis Fundamentals
What is Redis?
- Redis = Remote Dictionary Server
- In-memory data structure store
- Used as: Database, Cache, Message broker
- Single-threaded (but I/O multiplexing)
- Written in C, extremely fast
Key Characteristics
- In-memory: All data stored in RAM
- Persistent: Optional disk persistence
- Atomic operations: All operations are atomic
- Data structures: Rich set of data types
- Replication: Master-slave replication
- Clustering: Automatic partitioning
2. Data Types & Commands
Strings
SET key value # Set key
GET key # Get value
SET key value EX 60 # Set with 60s expiry
INCR counter # Increment by 1
DECR counter # Decrement by 1
APPEND key value # Append to string
STRLEN key # String length
Lists (Linked Lists)
LPUSH list value # Add to left
RPUSH list value # Add to right
LPOP list # Remove from left
RPOP list # Remove from right
LRANGE list 0 -1 # Get all elements
LLEN list # List length
Sets (Unordered Collections)
SADD set member # Add member
SREM set member # Remove member
SMEMBERS set # Get all members
SISMEMBER set member # Check membership
SCARD set # Set cardinality
SUNION set1 set2 # Union of sets
Sorted Sets (Ordered by Score)
ZADD zset 100 member # Add with score
ZRANGE zset 0 -1 # Get by index
ZRANGEBYSCORE zset 0 100 # Get by score range
ZRANK zset member # Get rank
ZSCORE zset member # Get score
ZREM zset member # Remove member
Hashes (Field-Value Pairs)
HSET hash field value # Set field
HGET hash field # Get field value
HMGET hash f1 f2 # Get multiple fields
HGETALL hash # Get all fields
HDEL hash field # Delete field
HEXISTS hash field # Check field exists
Bitmaps (String-based)
SETBIT key offset 1 # Set bit
GETBIT key offset # Get bit
BITCOUNT key # Count set bits
BITOP AND dest key1 key2 # Bitwise operations
HyperLogLog (Cardinality Estimation)
PFADD hll element # Add element
PFCOUNT hll # Estimate count
PFMERGE dest hll1 hll2 # Merge HLLs
Streams (Message Queue)
XADD stream * field value # Add message
XREAD COUNT 2 STREAMS s 0 # Read messages
XRANGE stream - + # Get range
3. Key Management
EXISTS key # Check if key exists
DEL key # Delete key
EXPIRE key 60 # Set TTL (seconds)
TTL key # Get remaining TTL
PERSIST key # Remove expiration
KEYS pattern # Find keys (avoid in prod)
SCAN cursor MATCH pattern # Iterate keys safely
TYPE key # Get key type
4. Transactions
MULTI # Start transaction
SET key1 value1
SET key2 value2
EXEC # Execute transaction
DISCARD # Cancel transaction
WATCH key # Optimistic locking
5. Pub/Sub (Publish/Subscribe)
SUBSCRIBE channel # Subscribe to channel
PUBLISH channel message # Publish message
UNSUBSCRIBE channel # Unsubscribe
PSUBSCRIBE pattern* # Pattern subscribe
6. Persistence Options
RDB (Redis Database)
- Point-in-time snapshots
- Compact, good for backups
- Faster restarts
- Risk of data loss
SAVE # Synchronous save
BGSAVE # Background save
AOF (Append Only File)
- Logs every write operation
- More durable
- Larger files
- Slower restarts
BGREWRITEAOF # Optimize AOF file
Configuration
# redis.conf
save 900 1 # RDB: Save after 900s if 1 key changed
appendonly yes # Enable AOF
appendfsync everysec # AOF sync policy
7. Replication & High Availability
Master-Slave Replication
# On slave
REPLICAOF master_ip port # Make replica of master
REPLICAOF NO ONE # Promote to master
Sentinel (HA)
- Monitors masters and slaves
- Automatic failover
- Configuration provider
# Start sentinel
redis-sentinel sentinel.conf
Cluster
- Automatic sharding
- 16384 hash slots
- Master-slave nodes
CLUSTER NODES # List nodes
CLUSTER INFO # Cluster info
CLUSTER SLOTS # Slot assignments
8. Performance Optimization
Memory Optimization
# Eviction policies
maxmemory-policy noeviction # No eviction (default)
maxmemory-policy lru # Least Recently Used
maxmemory-policy lfu # Least Frequently Used
maxmemory-policy volatile-lru # LRU for keys with TTL
Pipeline (Batch Commands)
# Python example
pipe = redis.pipeline()
pipe.set('key1', 'value1')
pipe.set('key2', 'value2')
pipe.execute()
Lua Scripting
-- Atomic operations
EVAL "return redis.call('get', KEYS[1])" 1 key
9. Common Patterns & Use Cases
1. Caching
def get_user(user_id):
# Check cache
user = redis.get(f"user:{user_id}")
if user:
return json.loads(user)
# Cache miss - get from DB
user = db.get_user(user_id)
redis.setex(f"user:{user_id}", 3600, json.dumps(user))
return user
2. Session Storage
# Store session
redis.setex(f"session:{session_id}", 1800, user_data)
# Get session
session_data = redis.get(f"session:{session_id}")
3. Rate Limiting
def is_rate_limited(user_id, limit=10):
key = f"rate:{user_id}:{int(time.time()/60)}"
count = redis.incr(key)
redis.expire(key, 60)
return count > limit
4. Leaderboards
# Add scores
ZADD leaderboard 100 "player1"
ZADD leaderboard 150 "player2"
# Get top 10
ZREVRANGE leaderboard 0 9 WITHSCORES
5. Real-time Analytics
# Count unique visitors
PFADD visitors:2024-01-01 user123
PFCOUNT visitors:2024-01-01
6. Distributed Locks
# Simple lock implementation
def acquire_lock(lock_name, timeout=10):
identifier = str(uuid.uuid4())
return redis.set(lock_name, identifier, nx=True, ex=timeout)
10. Advanced Patterns & Architecture
Redis Cluster Architecture
# Cluster with 6 nodes (3 masters, 3 replicas)
redis-cli --cluster create \
127.0.0.1:7000 127.0.0.1:7001 127.0.0.1:7002 \
127.0.0.1:7003 127.0.0.1:7004 127.0.0.1:7005 \
--cluster-replicas 1
# Key distribution via hash slots
# CRC16(key) mod 16384 determines slot
# Each master handles subset of 16384 slots
Advanced Data Structures Usage
Geospatial Indexes
# Add locations
GEOADD locations 13.361389 38.115556 "Palermo"
GEOADD locations 15.087269 37.502669 "Catania"
# Find nearby locations
GEORADIUS locations 15 37 200 km WITHDIST
# Get distance between points
GEODIST locations Palermo Catania km
Streams for Event Sourcing
# Add events to stream
XADD events * user_id 123 action login timestamp 1234567890
XADD events * user_id 123 action purchase item_id 456
# Read events
XREAD COUNT 100 STREAMS events 0
# Consumer groups
XGROUP CREATE events mygroup $
XREADGROUP GROUP mygroup consumer1 COUNT 10 STREAMS events >
Production Patterns
Cache-Aside Pattern
def get_user_with_cache(user_id):
# Try cache first
cache_key = f"user:{user_id}"
cached = redis.get(cache_key)
if cached:
# Cache hit
return json.loads(cached)
# Cache miss - get from database
user = database.get_user(user_id)
# Write to cache with TTL
redis.setex(cache_key, 3600, json.dumps(user))
return user
# Cache invalidation
def update_user(user_id, data):
database.update_user(user_id, data)
redis.delete(f"user:{user_id}")
Write-Through Cache
def save_user(user_id, data):
# Write to cache and database
cache_key = f"user:{user_id}"
# Atomic operation
with redis.pipeline() as pipe:
pipe.multi()
pipe.setex(cache_key, 3600, json.dumps(data))
database.save_user(user_id, data)
pipe.execute()
Distributed Rate Limiting
-- Sliding window rate limiter (Lua script)
local key = KEYS[1]
local window = tonumber(ARGV[1])
local limit = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
-- Remove old entries
redis.call('ZREMRANGEBYSCORE', key, 0, now - window)
-- Count current requests
local current = redis.call('ZCARD', key)
if current < limit then
-- Add new request
redis.call('ZADD', key, now, now)
redis.call('EXPIRE', key, window)
return 1
else
return 0
end
Redis Design Patterns
1. Bloom Filter Pattern
# Check membership without storing all data
def might_exist(item):
# Use multiple hash functions
for i in range(3):
bit_position = hash(f"{item}:{i}") % 1000000
if not redis.getbit("bloom:filter", bit_position):
return False
return True # Might exist (false positives possible)
def add_to_bloom(item):
for i in range(3):
bit_position = hash(f"{item}:{i}") % 1000000
redis.setbit("bloom:filter", bit_position, 1)
2. Delayed Queue Pattern
# Add job with delay
def add_delayed_job(job_data, delay_seconds):
score = time.time() + delay_seconds
redis.zadd("delayed:queue", {json.dumps(job_data): score})
# Process delayed jobs
def process_delayed_jobs():
now = time.time()
# Get jobs ready to process
jobs = redis.zrangebyscore("delayed:queue", 0, now, start=0, num=10)
for job_data in jobs:
# Process job
process_job(json.loads(job_data))
# Remove from queue
redis.zrem("delayed:queue", job_data)
3. Distributed Lock with Redlock
import time
import uuid
class Redlock:
def __init__(self, redis_nodes, retry_count=3, retry_delay=0.2):
self.redis_nodes = redis_nodes
self.retry_count = retry_count
self.retry_delay = retry_delay
self.quorum = len(redis_nodes) // 2 + 1
def acquire_lock(self, resource, ttl):
identifier = str(uuid.uuid4())
for attempt in range(self.retry_count):
acquired = 0
start_time = time.time()
# Try to acquire lock on all nodes
for redis_node in self.redis_nodes:
if self._acquire_lock_instance(redis_node, resource, identifier, ttl):
acquired += 1
# Check if we have quorum
elapsed_time = (time.time() - start_time) * 1000
validity_time = ttl - elapsed_time
if acquired >= self.quorum and validity_time > 0:
return identifier, validity_time
else:
# Release partial locks
self._release_all(resource, identifier)
time.sleep(self.retry_delay)
return None, 0
11. Performance Optimization Deep Dive
Memory Optimization Strategies
# Memory analysis
redis-cli --bigkeys # Find large keys
redis-cli --memkeys # Memory by key pattern
# Compression for strings
SET key "compressed_value"
# Use client-side compression for large values
# Hash optimization
# Convert multiple keys to hash fields
# Before: user:1:name, user:1:email
# After: HSET user:1 name "John" email "john@example.com"
Pipeline vs Transaction Performance
# Pipeline - faster for bulk operations
def bulk_insert_pipeline(data):
pipe = redis.pipeline(transaction=False)
for key, value in data.items():
pipe.set(key, value)
return pipe.execute()
# Transaction - guarantees atomicity
def atomic_transfer(from_key, to_key, amount):
with redis.pipeline() as pipe:
while True:
try:
pipe.watch(from_key)
balance = int(pipe.get(from_key) or 0)
if balance < amount:
pipe.unwatch()
return False
pipe.multi()
pipe.decrby(from_key, amount)
pipe.incrby(to_key, amount)
pipe.execute()
return True
except redis.WatchError:
continue
12. Redis vs Other Technologies
Redis vs Traditional Databases
| Feature | Redis | RDBMS |
|---|---|---|
| Storage | In-memory | Disk-based |
| Data Model | Key-Value + structures | Relational |
| ACID | Limited | Full |
| Query Language | Commands | SQL |
| Performance | Microseconds | Milliseconds |
| Persistence | Optional | Always |
Redis vs Message Queues
| Feature | Redis | RabbitMQ/Kafka |
|---|---|---|
| Use Case | Simple queues | Complex routing |
| Persistence | Optional | Built-in |
| Delivery Guarantees | Basic | Advanced |
| Throughput | Very High | High |
| Features | Limited | Rich |
13. Troubleshooting & Best Practices
Common Performance Issues
# Identify slow commands
SLOWLOG GET 10
# Check for blocking operations
CLIENT LIST
# Memory fragmentation
INFO memory
# mem_fragmentation_ratio > 1.5 indicates fragmentation
# Hot key detection
redis-cli --hotkeys
# Monitor real-time commands
MONITOR # Use sparingly in production
Production Checklist
Security
- Bind to private IPs only
- Use strong passwords
- Enable ACL for fine-grained access
- Use TLS for encryption
High Availability
- Configure Redis Sentinel
- Set up proper replication
- Test failover procedures
- Monitor replication lag
Performance
- Set appropriate maxmemory
- Choose correct eviction policy
- Use connection pooling
- Avoid large keys/values
Monitoring
- Track memory usage
- Monitor command latency
- Watch for connection count
- Set up alerts for critical metrics
14. Key Interview Concepts
Architecture Deep Dive
Single-threaded Model
- Uses I/O multiplexing (epoll/kqueue)
- Avoids context switching overhead
- Commands executed sequentially
- Background threads for specific tasks (Redis 6+)
Memory Management
- Uses jemalloc for efficient allocation
- Copy-on-write for RDB snapshots
- Memory fragmentation considerations
- Lazy freeing for large objects
Networking
- TCP server with custom protocol (RESP)
- Pipelining for batching commands
- Connection pooling best practices
- Unix domain sockets for local connections
Common Interview Scenarios
Scenario 1: Design a Distributed Cache
class DistributedCache:
def __init__(self, redis_cluster, local_cache_size=1000):
self.redis = redis_cluster
self.local_cache = LRUCache(local_cache_size)
def get(self, key):
# L1 cache (local)
value = self.local_cache.get(key)
if value:
return value
# L2 cache (Redis)
value = self.redis.get(key)
if value:
self.local_cache.put(key, value)
return value
return None
def set(self, key, value, ttl=3600):
# Write to both caches
self.local_cache.put(key, value)
self.redis.setex(key, ttl, value)
Scenario 2: Implement Leaderboard System
class Leaderboard:
def __init__(self, redis_client, leaderboard_key):
self.redis = redis_client
self.key = leaderboard_key
def add_score(self, user_id, score):
# Add or update score
self.redis.zadd(self.key, {user_id: score})
def get_top_players(self, count=10):
# Get top N players with scores
return self.redis.zrevrange(self.key, 0, count-1, withscores=True)
def get_user_rank(self, user_id):
# Get user's rank (0-based)
rank = self.redis.zrevrank(self.key, user_id)
return rank + 1 if rank is not None else None
def get_surrounding_players(self, user_id, count=5):
# Get players around user
rank = self.redis.zrevrank(self.key, user_id)
if rank is None:
return []
start = max(0, rank - count)
end = rank + count
return self.redis.zrevrange(self.key, start, end, withscores=True)
Performance Considerations
Command Complexity
- O(1): GET, SET, HGET, HSET
- O(log N): ZADD, ZREM
- O(N): KEYS, HGETALL, SMEMBERS
- O(N+M): SUNION, SINTER
Memory Optimization
- Use hashes for small objects
- Enable compression for lists
- Set appropriate TTLs
- Monitor memory fragmentation
Network Optimization
- Use pipelining for bulk operations
- Implement connection pooling
- Consider Unix sockets for local connections
- Use binary protocol when possible
15. Redis Ecosystem & Tools
Monitoring Tools
- Redis Insight: Official GUI
- redis-stat: Real-time stats
- Prometheus + Grafana: Metrics & visualization
- redis-cli --stat: Built-in monitoring
Client Libraries Best Practices
# Python (redis-py)
import redis
from redis.sentinel import Sentinel
# Connection pool
pool = redis.ConnectionPool(
host='localhost',
port=6379,
max_connections=50,
decode_responses=True
)
r = redis.Redis(connection_pool=pool)
# Sentinel for HA
sentinel = Sentinel([('localhost', 26379)])
master = sentinel.master_for('mymaster', socket_timeout=0.1)
Deployment Patterns
- Standalone: Single instance
- Master-Replica: Read scaling
- Sentinel: Automatic failover
- Cluster: Horizontal scaling
- Redis on Flash: SSD extension
Interview Quick Reference
Do's and Don'ts
Do's:
- Explain time complexity of operations
- Discuss memory vs disk trade-offs
- Mention real-world use cases
- Consider data consistency requirements
- Think about scaling strategies
Don'ts:
- Don't use KEYS in production scenarios
- Don't ignore persistence options
- Don't overlook security considerations
- Don't assume Redis solves all problems
- Don't forget about data expiration
Key Takeaways
- Redis = Speed + Simplicity: In-memory, single-threaded, rich data structures
- Use Cases: Caching, sessions, real-time analytics, queues, pub/sub
- Trade-offs: Memory limitations, persistence overhead, eventual consistency
- Scaling: Replication for reads, clustering for writes, proper key design
- Best Practices: Monitor memory, use appropriate data types, implement proper error handling
Remember: Redis excels at speed and flexibility. Always consider whether Redis is the right tool for the specific problem at hand.