Summary
Amazon DynamoDB is a fully managed NoSQL database service that provides fast and predictable performance with seamless scalability. It supports both key-value and document data models, offers single-digit millisecond latency, and provides features like Global Secondary Indexes, DynamoDB Streams, and ACID transactions. Key concepts include partition and sort keys, capacity modes (on-demand vs provisioned), consistency models (eventual vs strong), and advanced patterns like single-table design. Essential for cloud-native applications requiring high availability and massive scale.
1. Core Concepts
What is DynamoDB?
- Fully managed NoSQL database by AWS
- Key-value and document database
- Serverless - no servers to manage
- Multi-region, multi-active database
- Built-in security, backup, and restore
Key Characteristics
- Performance: Single-digit millisecond latency
- Scalability: Virtually unlimited throughput and storage
- Availability: 99.999% availability SLA
- Consistency: Supports both eventual and strong consistency
2. Data Model
Primary Key Types
1. Partition Key (Simple Primary Key)
{
"UserId": "12345", # Partition Key
"Name": "John",
"Email": "john@example.com"
}
2. Composite Primary Key (Partition + Sort Key)
{
"UserId": "12345", # Partition Key
"OrderId": "ORDER001", # Sort Key
"Amount": 100.00,
"Date": "2024-01-15"
}
Data Types
- Scalar: String, Number, Binary, Boolean, Null
- Document: List, Map
- Set: String Set, Number Set, Binary Set
3. Key Components
1. Tables
- Container for items
- Schemaless (except primary key)
- Must define primary key at creation
2. Items
- Collection of attributes
- Max size: 400 KB
- Uniquely identified by primary key
3. Attributes
- Fundamental data element
- Can be nested (up to 32 levels)
4. Basic Operations
Create Table
import boto3
dynamodb = boto3.resource('dynamodb')
table = dynamodb.create_table(
TableName='Users',
KeySchema=[
{
'AttributeName': 'UserId',
'KeyType': 'HASH' # Partition key
}
],
AttributeDefinitions=[
{
'AttributeName': 'UserId',
'AttributeType': 'S' # String
}
],
BillingMode='PAY_PER_REQUEST'
)
Put Item
table.put_item(
Item={
'UserId': '123',
'Name': 'Alice',
'Age': 30
}
)
Get Item
response = table.get_item(
Key={'UserId': '123'}
)
item = response.get('Item')
Update Item
table.update_item(
Key={'UserId': '123'},
UpdateExpression='SET Age = :age',
ExpressionAttributeValues={
':age': 31
}
)
Delete Item
table.delete_item(
Key={'UserId': '123'}
)
5. Querying & Scanning
Query
- Efficient: Uses primary key
- Requires: Partition key value
- Optional: Sort key conditions
response = table.query(
KeyConditionExpression=Key('UserId').eq('123')
)
Scan
- Inefficient: Reads entire table
- Use sparingly: Expensive operation
- Supports: Filtering
response = table.scan(
FilterExpression=Attr('Age').gt(25)
)
6. Secondary Indexes
Global Secondary Index (GSI)
- Different partition and sort key
- Eventually consistent only
- Can be added after table creation
table.update(
AttributeDefinitions=[
{'AttributeName': 'Email', 'AttributeType': 'S'}
],
GlobalSecondaryIndexUpdates=[{
'Create': {
'IndexName': 'EmailIndex',
'Keys': [
{'AttributeName': 'Email', 'KeyType': 'HASH'}
],
'Projection': {'ProjectionType': 'ALL'}
}
}]
)
Local Secondary Index (LSI)
- Same partition key, different sort key
- Strong consistency available
- Must be created with table
7. Performance Optimization
1. Capacity Modes
On-Demand
- Pay per request
- No capacity planning
- Good for unpredictable workloads
Provisioned
- Specify RCU/WCU
- Auto-scaling available
- Cost-effective for predictable workloads
2. Read/Write Capacity Units
- RCU: 1 RCU = 1 strongly consistent read/sec (up to 4KB)
- WCU: 1 WCU = 1 write/sec (up to 1KB)
3. Batch Operations
# Batch write (up to 25 items)
with table.batch_writer() as batch:
for i in range(25):
batch.put_item(Item={'id': str(i), 'data': 'value'})
# Batch get
response = dynamodb.batch_get_item(
RequestItems={
'Users': {
'Keys': [{'UserId': '1'}, {'UserId': '2'}]
}
}
)
8. Data Modeling Best Practices
1. Single Table Design
- Store multiple entity types in one table
- Use generic attribute names
- Leverage sparse indexes
# Example: Orders and Products in same table
{
"PK": "USER#123",
"SK": "ORDER#456",
"Type": "Order",
"Amount": 100
}
{
"PK": "PRODUCT#789",
"SK": "PRODUCT#789",
"Type": "Product",
"Name": "Widget"
}
2. Access Patterns First
- Design table based on queries
- Denormalize data
- Duplicate data when needed
3. Hot Partition Prevention
- Add random suffix to partition keys
- Use write sharding
- Distribute load evenly
9. Advanced Features
1. DynamoDB Streams
- Captures data modifications
- Triggers Lambda functions
- Use cases: Replication, analytics
table = dynamodb.create_table(
TableName='MyTable',
StreamSpecification={
'StreamEnabled': True,
'StreamViewType': 'NEW_AND_OLD_IMAGES'
},
# ... other parameters
)
2. Transactions
- ACID transactions
- Up to 25 items
- All-or-nothing execution
client = boto3.client('dynamodb')
client.transact_write_items(
TransactItems=[
{
'Put': {
'TableName': 'Users',
'Item': {'UserId': {'S': '123'}}
}
},
{
'Update': {
'TableName': 'Accounts',
'Key': {'AccountId': {'S': '456'}},
'UpdateExpression': 'SET balance = balance - :amt',
'ExpressionAttributeValues': {':amt': {'N': '50'}}
}
}
]
)
3. Time to Live (TTL)
- Automatic item deletion
- No additional cost
- Background process
# Set TTL attribute
table.put_item(
Item={
'UserId': '123',
'ExpirationTime': 1640995200 # Unix timestamp
}
)
10. Use Cases & Design Patterns
When to Use DynamoDB
Good for:
- High-scale web applications (millions of requests/sec)
- Gaming leaderboards and session storage
- IoT device data collection
- Real-time bidding platforms
- Mobile app backends
- Content management and catalogs
Not ideal for:
- Complex relational queries with joins
- Analytics and reporting workloads
- Applications requiring strong ACID guarantees
- Traditional OLAP operations
DynamoDB vs Other Databases
vs Amazon RDS
| Feature | DynamoDB | RDS |
|---|---|---|
| Type | NoSQL | SQL |
| Schema | Flexible | Fixed |
| Scaling | Horizontal | Vertical |
| Maintenance | Fully managed | Some management |
| Queries | Key-based | SQL with joins |
| Consistency | Eventual/Strong | Strong |
vs MongoDB
| Feature | DynamoDB | MongoDB |
|---|---|---|
| Hosting | AWS managed | Self-managed or Atlas |
| Scaling | Automatic | Manual sharding |
| Query Language | Boto3/CLI | MongoDB Query Language |
| Transactions | Limited (25 items) | Full ACID |
Common Design Patterns
Single Table Design
# User profile and orders in same table
{
"PK": "USER#123",
"SK": "PROFILE",
"UserName": "john_doe",
"Email": "john@example.com"
}
{
"PK": "USER#123",
"SK": "ORDER#2024-01-15#001",
"Amount": 99.99,
"Status": "shipped"
}
Time Series Data
# IoT sensor data with time-based partition
{
"PK": "SENSOR#ABC123",
"SK": "2024-01-15T10:30:00Z",
"Temperature": 23.5,
"Humidity": 65.2,
"TTL": 1640995200 # Auto-delete old data
}
Leaderboard Pattern
# Gaming leaderboard with GSI
{
"PK": "GAME#racing",
"SK": "USER#player123",
"Score": 15420,
"GSI1PK": "LEADERBOARD#racing",
"GSI1SK": "SCORE#000015420" # Zero-padded for sorting
}
11. Performance & Cost Optimization
Capacity Planning
- On-Demand: Good for unpredictable or new workloads
- Provisioned: Better cost for steady, predictable traffic
- Auto Scaling: Adjusts provisioned capacity automatically
Hot Partition Prevention
# Bad: All items go to same partition
PK = "LOGS"
# Good: Distribute across multiple partitions
PK = f"LOGS#{random.randint(1, 10)}"
PK = f"LOGS#{datetime.now().strftime('%Y-%m-%d-%H')}"
Query Optimization
# Efficient: Query with partition key
response = table.query(
KeyConditionExpression=Key('PK').eq('USER#123')
)
# Less efficient: Scan with filter
response = table.scan(
FilterExpression=Attr('UserName').eq('john_doe')
)
Error Handling & Retry Logic
import time
import random
from botocore.exceptions import ClientError
def exponential_backoff_retry(func, max_retries=3):
for attempt in range(max_retries):
try:
return func()
except ClientError as e:
if e.response['Error']['Code'] == 'ProvisionedThroughputExceededException':
if attempt < max_retries - 1:
wait_time = (2 ** attempt) + random.uniform(0, 1)
time.sleep(wait_time)
else:
raise
else:
raise
12. Key Limits & Best Practices
Important Limits
- Item Size: 400 KB maximum
- Batch Operations: 25 items max
- Transaction Items: 25 items max
- Query Result: 1 MB max (pagination required)
- Attribute Name: 255 characters max
- GSI per Table: 20 maximum
Best Practices
- Design for Access Patterns: Know your queries before designing
- Use Sparse Indexes: GSI only includes items with the index key
- Implement Pagination: Handle large result sets properly
- Monitor Metrics: Watch consumed capacity and throttles
- Use Projections: Only include needed attributes in GSI
- Enable Point-in-Time Recovery: For production tables
- Set Up Alarms: Monitor throttling and error rates
13. Quick Reference
CLI Commands
# List tables
aws dynamodb list-tables
# Describe table
aws dynamodb describe-table --table-name Users
# Get item
aws dynamodb get-item --table-name Users --key '{"UserId":{"S":"123"}}'
# Put item
aws dynamodb put-item --table-name Users --item '{"UserId":{"S":"123"},"Name":{"S":"John"}}'
Expression Examples
# Condition Expression
ConditionExpression='attribute_exists(UserId)'
# Update Expression
UpdateExpression='SET #n = :name, Age = Age + :inc'
# Filter Expression
FilterExpression='Age > :age AND #s = :status'
# Projection Expression
ProjectionExpression='UserId, #n, Email'
Remember: DynamoDB is about predictable performance at any scale!