LearnThatStack Ace your next interview

DynamoDB.
Interview cheat sheet.

Quick reference for DynamoDB - sectioned for fast scanning. Skim the part you're shaky on, walk in confident.

Database Technologies 14-section reference ~5 min read

Summary

Amazon DynamoDB is a fully managed NoSQL database service that provides fast and predictable performance with seamless scalability. It supports both key-value and document data models, offers single-digit millisecond latency, and provides features like Global Secondary Indexes, DynamoDB Streams, and ACID transactions. Key concepts include partition and sort keys, capacity modes (on-demand vs provisioned), consistency models (eventual vs strong), and advanced patterns like single-table design. Essential for cloud-native applications requiring high availability and massive scale.

1. Core Concepts

What is DynamoDB?

  • Fully managed NoSQL database by AWS
  • Key-value and document database
  • Serverless - no servers to manage
  • Multi-region, multi-active database
  • Built-in security, backup, and restore

Key Characteristics

  • Performance: Single-digit millisecond latency
  • Scalability: Virtually unlimited throughput and storage
  • Availability: 99.999% availability SLA
  • Consistency: Supports both eventual and strong consistency

2. Data Model

Primary Key Types

1. Partition Key (Simple Primary Key)

{
    "UserId": "12345",  # Partition Key
    "Name": "John",
    "Email": "john@example.com"
}

2. Composite Primary Key (Partition + Sort Key)

{
    "UserId": "12345",      # Partition Key
    "OrderId": "ORDER001",  # Sort Key
    "Amount": 100.00,
    "Date": "2024-01-15"
}

Data Types

  • Scalar: String, Number, Binary, Boolean, Null
  • Document: List, Map
  • Set: String Set, Number Set, Binary Set

3. Key Components

1. Tables

  • Container for items
  • Schemaless (except primary key)
  • Must define primary key at creation

2. Items

  • Collection of attributes
  • Max size: 400 KB
  • Uniquely identified by primary key

3. Attributes

  • Fundamental data element
  • Can be nested (up to 32 levels)

4. Basic Operations

Create Table

import boto3

dynamodb = boto3.resource('dynamodb')

table = dynamodb.create_table(
    TableName='Users',
    KeySchema=[
        {
            'AttributeName': 'UserId',
            'KeyType': 'HASH'  # Partition key
        }
    ],
    AttributeDefinitions=[
        {
            'AttributeName': 'UserId',
            'AttributeType': 'S'  # String
        }
    ],
    BillingMode='PAY_PER_REQUEST'
)

Put Item

table.put_item(
    Item={
        'UserId': '123',
        'Name': 'Alice',
        'Age': 30
    }
)

Get Item

response = table.get_item(
    Key={'UserId': '123'}
)
item = response.get('Item')

Update Item

table.update_item(
    Key={'UserId': '123'},
    UpdateExpression='SET Age = :age',
    ExpressionAttributeValues={
        ':age': 31
    }
)

Delete Item

table.delete_item(
    Key={'UserId': '123'}
)

5. Querying & Scanning

Query

  • Efficient: Uses primary key
  • Requires: Partition key value
  • Optional: Sort key conditions
response = table.query(
    KeyConditionExpression=Key('UserId').eq('123')
)

Scan

  • Inefficient: Reads entire table
  • Use sparingly: Expensive operation
  • Supports: Filtering
response = table.scan(
    FilterExpression=Attr('Age').gt(25)
)

6. Secondary Indexes

Global Secondary Index (GSI)

  • Different partition and sort key
  • Eventually consistent only
  • Can be added after table creation
table.update(
    AttributeDefinitions=[
        {'AttributeName': 'Email', 'AttributeType': 'S'}
    ],
    GlobalSecondaryIndexUpdates=[{
        'Create': {
            'IndexName': 'EmailIndex',
            'Keys': [
                {'AttributeName': 'Email', 'KeyType': 'HASH'}
            ],
            'Projection': {'ProjectionType': 'ALL'}
        }
    }]
)

Local Secondary Index (LSI)

  • Same partition key, different sort key
  • Strong consistency available
  • Must be created with table

7. Performance Optimization

1. Capacity Modes

On-Demand

  • Pay per request
  • No capacity planning
  • Good for unpredictable workloads

Provisioned

  • Specify RCU/WCU
  • Auto-scaling available
  • Cost-effective for predictable workloads

2. Read/Write Capacity Units

  • RCU: 1 RCU = 1 strongly consistent read/sec (up to 4KB)
  • WCU: 1 WCU = 1 write/sec (up to 1KB)

3. Batch Operations

# Batch write (up to 25 items)
with table.batch_writer() as batch:
    for i in range(25):
        batch.put_item(Item={'id': str(i), 'data': 'value'})

# Batch get
response = dynamodb.batch_get_item(
    RequestItems={
        'Users': {
            'Keys': [{'UserId': '1'}, {'UserId': '2'}]
        }
    }
)

8. Data Modeling Best Practices

1. Single Table Design

  • Store multiple entity types in one table
  • Use generic attribute names
  • Leverage sparse indexes
# Example: Orders and Products in same table
{
    "PK": "USER#123",
    "SK": "ORDER#456",
    "Type": "Order",
    "Amount": 100
}
{
    "PK": "PRODUCT#789",
    "SK": "PRODUCT#789",
    "Type": "Product",
    "Name": "Widget"
}

2. Access Patterns First

  • Design table based on queries
  • Denormalize data
  • Duplicate data when needed

3. Hot Partition Prevention

  • Add random suffix to partition keys
  • Use write sharding
  • Distribute load evenly

9. Advanced Features

1. DynamoDB Streams

  • Captures data modifications
  • Triggers Lambda functions
  • Use cases: Replication, analytics
table = dynamodb.create_table(
    TableName='MyTable',
    StreamSpecification={
        'StreamEnabled': True,
        'StreamViewType': 'NEW_AND_OLD_IMAGES'
    },
    # ... other parameters
)

2. Transactions

  • ACID transactions
  • Up to 25 items
  • All-or-nothing execution
client = boto3.client('dynamodb')
client.transact_write_items(
    TransactItems=[
        {
            'Put': {
                'TableName': 'Users',
                'Item': {'UserId': {'S': '123'}}
            }
        },
        {
            'Update': {
                'TableName': 'Accounts',
                'Key': {'AccountId': {'S': '456'}},
                'UpdateExpression': 'SET balance = balance - :amt',
                'ExpressionAttributeValues': {':amt': {'N': '50'}}
            }
        }
    ]
)

3. Time to Live (TTL)

  • Automatic item deletion
  • No additional cost
  • Background process
# Set TTL attribute
table.put_item(
    Item={
        'UserId': '123',
        'ExpirationTime': 1640995200  # Unix timestamp
    }
)

10. Use Cases & Design Patterns

When to Use DynamoDB

Good for:

  • High-scale web applications (millions of requests/sec)
  • Gaming leaderboards and session storage
  • IoT device data collection
  • Real-time bidding platforms
  • Mobile app backends
  • Content management and catalogs

Not ideal for:

  • Complex relational queries with joins
  • Analytics and reporting workloads
  • Applications requiring strong ACID guarantees
  • Traditional OLAP operations

DynamoDB vs Other Databases

vs Amazon RDS

Feature DynamoDB RDS
Type NoSQL SQL
Schema Flexible Fixed
Scaling Horizontal Vertical
Maintenance Fully managed Some management
Queries Key-based SQL with joins
Consistency Eventual/Strong Strong

vs MongoDB

Feature DynamoDB MongoDB
Hosting AWS managed Self-managed or Atlas
Scaling Automatic Manual sharding
Query Language Boto3/CLI MongoDB Query Language
Transactions Limited (25 items) Full ACID

Common Design Patterns

Single Table Design

# User profile and orders in same table
{
    "PK": "USER#123",
    "SK": "PROFILE",
    "UserName": "john_doe",
    "Email": "john@example.com"
}
{
    "PK": "USER#123", 
    "SK": "ORDER#2024-01-15#001",
    "Amount": 99.99,
    "Status": "shipped"
}

Time Series Data

# IoT sensor data with time-based partition
{
    "PK": "SENSOR#ABC123",
    "SK": "2024-01-15T10:30:00Z",
    "Temperature": 23.5,
    "Humidity": 65.2,
    "TTL": 1640995200  # Auto-delete old data
}

Leaderboard Pattern

# Gaming leaderboard with GSI
{
    "PK": "GAME#racing",
    "SK": "USER#player123", 
    "Score": 15420,
    "GSI1PK": "LEADERBOARD#racing",
    "GSI1SK": "SCORE#000015420"  # Zero-padded for sorting
}

11. Performance & Cost Optimization

Capacity Planning

  • On-Demand: Good for unpredictable or new workloads
  • Provisioned: Better cost for steady, predictable traffic
  • Auto Scaling: Adjusts provisioned capacity automatically

Hot Partition Prevention

# Bad: All items go to same partition
PK = "LOGS"

# Good: Distribute across multiple partitions
PK = f"LOGS#{random.randint(1, 10)}"
PK = f"LOGS#{datetime.now().strftime('%Y-%m-%d-%H')}"

Query Optimization

# Efficient: Query with partition key
response = table.query(
    KeyConditionExpression=Key('PK').eq('USER#123')
)

# Less efficient: Scan with filter
response = table.scan(
    FilterExpression=Attr('UserName').eq('john_doe')
)

Error Handling & Retry Logic

import time
import random
from botocore.exceptions import ClientError

def exponential_backoff_retry(func, max_retries=3):
    for attempt in range(max_retries):
        try:
            return func()
        except ClientError as e:
            if e.response['Error']['Code'] == 'ProvisionedThroughputExceededException':
                if attempt < max_retries - 1:
                    wait_time = (2 ** attempt) + random.uniform(0, 1)
                    time.sleep(wait_time)
                else:
                    raise
            else:
                raise

12. Key Limits & Best Practices

Important Limits

  • Item Size: 400 KB maximum
  • Batch Operations: 25 items max
  • Transaction Items: 25 items max
  • Query Result: 1 MB max (pagination required)
  • Attribute Name: 255 characters max
  • GSI per Table: 20 maximum

Best Practices

  1. Design for Access Patterns: Know your queries before designing
  2. Use Sparse Indexes: GSI only includes items with the index key
  3. Implement Pagination: Handle large result sets properly
  4. Monitor Metrics: Watch consumed capacity and throttles
  5. Use Projections: Only include needed attributes in GSI
  6. Enable Point-in-Time Recovery: For production tables
  7. Set Up Alarms: Monitor throttling and error rates

13. Quick Reference

CLI Commands

# List tables
aws dynamodb list-tables

# Describe table
aws dynamodb describe-table --table-name Users

# Get item
aws dynamodb get-item --table-name Users --key '{"UserId":{"S":"123"}}'

# Put item
aws dynamodb put-item --table-name Users --item '{"UserId":{"S":"123"},"Name":{"S":"John"}}'

Expression Examples

# Condition Expression
ConditionExpression='attribute_exists(UserId)'

# Update Expression
UpdateExpression='SET #n = :name, Age = Age + :inc'

# Filter Expression
FilterExpression='Age > :age AND #s = :status'

# Projection Expression
ProjectionExpression='UserId, #n, Email'

Remember: DynamoDB is about predictable performance at any scale!

Found this useful? Pass it on.
Pro · $10/mo

The sheet is free. Pro goes deeper.

Pro opens the full question library behind every sheet, every refresher and a monthly AI allowance. One subscription, all formats.

Full question library All refreshers Cancel anytime