LearnThatStack Ace your next interview

Elasticsearch.
Interview cheat sheet.

Quick reference for Elasticsearch - sectioned for fast scanning. Skim the part you're shaky on, walk in confident.

Database Technologies 14-section reference ~6 min read

Summary

Elasticsearch is a distributed, RESTful search and analytics engine built on Apache Lucene. It provides near real-time search capabilities with horizontal scalability and schema-free JSON document storage. Key concepts include clusters, nodes, indices, shards, and documents. Features advanced search capabilities through Query DSL, aggregations for analytics, and performance optimization through proper indexing and sharding strategies. Essential for applications requiring full-text search, log analysis, and real-time data analytics.

1. Core Concepts

What is Elasticsearch?

  • Distributed, RESTful search and analytics engine
  • Built on Apache Lucene
  • Near real-time search capabilities
  • Schema-free JSON documents
  • Horizontally scalable

Key Components

Cluster → Nodes → Indices → Shards → Documents
  • Cluster: Collection of nodes
  • Node: Single server instance
  • Index: Collection of documents (like database)
  • Shard: Subdivision of index (primary/replica)
  • Document: Basic unit of information (JSON)

2. Architecture

Node Types

  1. Master Node: Cluster management
  2. Data Node: Stores data, executes queries
  3. Ingest Node: Pre-processes documents
  4. Coordinating Node: Routes requests

Sharding

  • Primary Shard: Original data
  • Replica Shard: Copy for fault tolerance
  • Default: 1 primary, 1 replica per index

3. Basic Operations (CRUD)

Create/Index Document

PUT /products/_doc/1
{
  "name": "Laptop",
  "price": 999,
  "brand": "Dell"
}

Read/Get Document

GET /products/_doc/1

Update Document

POST /products/_update/1
{
  "doc": {
    "price": 899
  }
}

Delete Document

DELETE /products/_doc/1

Bulk Operations

POST /_bulk
{"index": {"_index": "products", "_id": "2"}}
{"name": "Mouse", "price": 25}
{"delete": {"_index": "products", "_id": "1"}}

4. Search & Query DSL

Match Query (Full-text)

GET /products/_search
{
  "query": {
    "match": {
      "name": "laptop"
    }
  }
}

Term Query (Exact match)

GET /products/_search
{
  "query": {
    "term": {
      "brand.keyword": "Dell"
    }
  }
}

Bool Query (Compound)

GET /products/_search
{
  "query": {
    "bool": {
      "must": [
        {"match": {"name": "laptop"}}
      ],
      "filter": [
        {"range": {"price": {"gte": 500, "lte": 1500}}}
      ],
      "must_not": [
        {"term": {"brand.keyword": "HP"}}
      ]
    }
  }
}

Common Query Types

  • match: Full-text search
  • term: Exact value
  • range: Numeric/date ranges
  • exists: Field existence
  • prefix: Prefix matching
  • wildcard: Pattern matching
  • fuzzy: Approximate matching

5. Aggregations

Metrics Aggregations

GET /products/_search
{
  "aggs": {
    "avg_price": {
      "avg": {
        "field": "price"
      }
    },
    "max_price": {
      "max": {
        "field": "price"
      }
    }
  }
}

Bucket Aggregations

GET /products/_search
{
  "aggs": {
    "brands": {
      "terms": {
        "field": "brand.keyword"
      },
      "aggs": {
        "avg_price": {
          "avg": {
            "field": "price"
          }
        }
      }
    }
  }
}

Common Aggregations

  • Metrics: sum, avg, min, max, stats
  • Bucket: terms, range, date_histogram
  • Pipeline: moving_avg, derivative

6. Mapping & Index Management

Create Index with Mapping

PUT /products
{
  "mappings": {
    "properties": {
      "name": {
        "type": "text",
        "fields": {
          "keyword": {
            "type": "keyword"
          }
        }
      },
      "price": {
        "type": "float"
      },
      "created_at": {
        "type": "date"
      },
      "tags": {
        "type": "keyword"
      }
    }
  }
}

Field Data Types

  • Text: Full-text searchable
  • Keyword: Exact matches, sorting, aggregations
  • Numeric: long, integer, short, byte, double, float
  • Date: Date/time values
  • Boolean: true/false
  • Object: JSON objects
  • Nested: Arrays of objects

Index Settings

PUT /products/_settings
{
  "number_of_replicas": 2,
  "refresh_interval": "30s"
}

7. Performance Optimization

Query Performance

  1. Use filters instead of queries when possible
  2. Limit result size with size parameter
  3. Use source filtering to return only needed fields
  4. Avoid wildcards at the beginning of terms

Example: Optimized Query

GET /products/_search
{
  "_source": ["name", "price"],
  "size": 10,
  "query": {
    "bool": {
      "filter": [
        {"term": {"brand.keyword": "Dell"}},
        {"range": {"price": {"gte": 500}}}
      ]
    }
  }
}

Indexing Performance

  1. Bulk indexing for multiple documents
  2. Disable refresh during bulk operations
  3. Use appropriate shard count
  4. Configure proper heap size (50% of RAM, max 32GB)

8. Cluster Management

Cluster Health

GET /_cluster/health

States:

  • Green: All shards allocated
  • Yellow: Primary allocated, replicas not
  • Red: Some primaries not allocated

Node Stats

GET /_nodes/stats

Index Stats

GET /products/_stats

9. Advanced Features

Analyzers

PUT /products
{
  "settings": {
    "analysis": {
      "analyzer": {
        "custom_analyzer": {
          "tokenizer": "standard",
          "filter": ["lowercase", "stop"]
        }
      }
    }
  }
}

Highlighting

GET /products/_search
{
  "query": {
    "match": {"name": "laptop"}
  },
  "highlight": {
    "fields": {
      "name": {}
    }
  }
}

Scroll API (Large Result Sets)

POST /products/_search?scroll=1m
{
  "size": 100,
  "query": {"match_all": {}}
}

Reindex API

POST /_reindex
{
  "source": {
    "index": "old_products"
  },
  "dest": {
    "index": "new_products"
  }
}

10. Use Cases & Design Patterns

When to Use Elasticsearch

Good for:

  • Full-text search applications
  • Log and event data analysis (ELK stack)
  • Real-time analytics dashboards
  • Content discovery and recommendations
  • Application performance monitoring (APM)
  • Security information and event management (SIEM)

Not ideal for:

  • Primary transactional data storage
  • Complex relational queries with joins
  • Strong consistency requirements
  • Frequent updates to individual documents

Common Design Patterns

Log Aggregation (ELK Stack)

// Logstash configuration for log parsing
{
  "index": "logs-2024.01.15",
  "document": {
    "@timestamp": "2024-01-15T10:30:00Z",
    "level": "ERROR",
    "message": "Database connection failed",
    "service": "api-gateway",
    "host": "server-01",
    "tags": ["error", "database"]
  }
}
// Product search with facets
GET /products/_search
{
  "query": {
    "bool": {
      "must": [
        {"match": {"name": "laptop"}}
      ],
      "filter": [
        {"term": {"category.keyword": "electronics"}},
        {"range": {"price": {"gte": 500, "lte": 2000}}}
      ]
    }
  },
  "aggs": {
    "brands": {
      "terms": {"field": "brand.keyword"}
    },
    "price_ranges": {
      "range": {
        "field": "price",
        "ranges": [
          {"to": 500},
          {"from": 500, "to": 1000},
          {"from": 1000}
        ]
      }
    }
  }
}

Time Series Analytics

// Website analytics dashboard
{
  "query": {
    "range": {
      "@timestamp": {
        "gte": "now-24h"
      }
    }
  },
  "aggs": {
    "page_views_over_time": {
      "date_histogram": {
        "field": "@timestamp",
        "interval": "1h"
      },
      "aggs": {
        "unique_visitors": {
          "cardinality": {
            "field": "user_id"
          }
        }
      }
    }
  }
}

Architecture Patterns

Hot-Warm-Cold Architecture

// Hot nodes (recent data, fast hardware)
PUT /logs-2024.01.15
{
  "settings": {
    "index.routing.allocation.require.data": "hot",
    "number_of_shards": 3,
    "number_of_replicas": 1
  }
}

// Warm nodes (older data, slower hardware)
PUT /logs-2024.01.01/_settings
{
  "index.routing.allocation.require.data": "warm",
  "number_of_replicas": 0
}

Index Templates

PUT /_template/logs_template
{
  "index_patterns": ["logs-*"],
  "settings": {
    "number_of_shards": 1,
    "number_of_replicas": 1
  },
  "mappings": {
    "properties": {
      "@timestamp": {"type": "date"},
      "message": {"type": "text"},
      "level": {"type": "keyword"}
    }
  }
}

Performance Optimization Strategies

Sharding Strategy

  • Small indices: 1 shard
  • Medium indices: 3-5 shards
  • Large indices: Number of shards = (data size in GB) / 30-50
  • Rule of thumb: Shard size between 10-50 GB

Memory Management

// JVM heap configuration
-Xms16g -Xmx16g  // Set min = max heap
// File system cache gets remaining memory

Query Optimization

// Use filters for exact matches (cached)
"filter": [
  {"term": {"status": "published"}},
  {"range": {"date": {"gte": "2024-01-01"}}}
]

// Use queries for relevance scoring
"query": {
  "match": {"content": "search terms"}
}

Troubleshooting Common Issues

Split Brain Prevention

# elasticsearch.yml
discovery.zen.minimum_master_nodes: 2  # (master_eligible_nodes / 2) + 1

Circuit Breaker Settings

PUT /_cluster/settings
{
  "transient": {
    "indices.breaker.fielddata.limit": "40%",
    "indices.breaker.request.limit": "60%"
  }
}

Slow Query Analysis

// Enable slow log
PUT /index/_settings
{
  "index.search.slowlog.threshold.query.warn": "10s",
  "index.search.slowlog.threshold.fetch.warn": "1s"
}

11. Security Features

Basic Authentication

curl -u elastic:password -X GET "localhost:9200/"

Role-Based Access Control

POST /_security/role/read_only
{
  "indices": [{
    "names": ["products"],
    "privileges": ["read"]
  }]
}

12. Monitoring & Debugging

Explain API

GET /products/_explain/1
{
  "query": {
    "match": {"name": "laptop"}
  }
}

Profile API

GET /products/_search
{
  "profile": true,
  "query": {
    "match": {"name": "laptop"}
  }
}

Slow Log Configuration

PUT /products/_settings
{
  "index.search.slowlog.threshold.query.warn": "10s",
  "index.search.slowlog.threshold.fetch.warn": "1s"
}

13. Quick Reference

REST API Patterns

GET    /_search                 # Search all indices
GET    /index/_search          # Search specific index
GET    /index/_doc/id          # Get document
PUT    /index/_doc/id          # Index document
POST   /index/_update/id       # Update document
DELETE /index/_doc/id          # Delete document
GET    /_cat/indices           # List indices
GET    /_cat/nodes             # List nodes
GET    /_cluster/health        # Cluster health

Query Context vs Filter Context

  • Query Context: Calculates relevance score
  • Filter Context: Yes/no match, cached, no scoring

Memory Guidelines

  • Heap: 50% of available RAM, max 32GB
  • File System Cache: Remaining 50% for Lucene
  • min_heap = max_heap for production

Remember: Elasticsearch is about trade-offs between speed, accuracy, and resource usage. Always consider your specific use case when making design decisions.

Found this useful? Pass it on.
Pro · $10/mo

The sheet is free. Pro goes deeper.

Pro opens the full question library behind every sheet, every refresher and a monthly AI allowance. One subscription, all formats.

Full question library All refreshers Cancel anytime