Summary
Elasticsearch is a distributed, RESTful search and analytics engine built on Apache Lucene. It provides near real-time search capabilities with horizontal scalability and schema-free JSON document storage. Key concepts include clusters, nodes, indices, shards, and documents. Features advanced search capabilities through Query DSL, aggregations for analytics, and performance optimization through proper indexing and sharding strategies. Essential for applications requiring full-text search, log analysis, and real-time data analytics.
1. Core Concepts
What is Elasticsearch?
- Distributed, RESTful search and analytics engine
- Built on Apache Lucene
- Near real-time search capabilities
- Schema-free JSON documents
- Horizontally scalable
Key Components
Cluster → Nodes → Indices → Shards → Documents
- Cluster: Collection of nodes
- Node: Single server instance
- Index: Collection of documents (like database)
- Shard: Subdivision of index (primary/replica)
- Document: Basic unit of information (JSON)
2. Architecture
Node Types
- Master Node: Cluster management
- Data Node: Stores data, executes queries
- Ingest Node: Pre-processes documents
- Coordinating Node: Routes requests
Sharding
- Primary Shard: Original data
- Replica Shard: Copy for fault tolerance
- Default: 1 primary, 1 replica per index
3. Basic Operations (CRUD)
Create/Index Document
PUT /products/_doc/1
{
"name": "Laptop",
"price": 999,
"brand": "Dell"
}
Read/Get Document
GET /products/_doc/1
Update Document
POST /products/_update/1
{
"doc": {
"price": 899
}
}
Delete Document
DELETE /products/_doc/1
Bulk Operations
POST /_bulk
{"index": {"_index": "products", "_id": "2"}}
{"name": "Mouse", "price": 25}
{"delete": {"_index": "products", "_id": "1"}}
4. Search & Query DSL
Match Query (Full-text)
GET /products/_search
{
"query": {
"match": {
"name": "laptop"
}
}
}
Term Query (Exact match)
GET /products/_search
{
"query": {
"term": {
"brand.keyword": "Dell"
}
}
}
Bool Query (Compound)
GET /products/_search
{
"query": {
"bool": {
"must": [
{"match": {"name": "laptop"}}
],
"filter": [
{"range": {"price": {"gte": 500, "lte": 1500}}}
],
"must_not": [
{"term": {"brand.keyword": "HP"}}
]
}
}
}
Common Query Types
- match: Full-text search
- term: Exact value
- range: Numeric/date ranges
- exists: Field existence
- prefix: Prefix matching
- wildcard: Pattern matching
- fuzzy: Approximate matching
5. Aggregations
Metrics Aggregations
GET /products/_search
{
"aggs": {
"avg_price": {
"avg": {
"field": "price"
}
},
"max_price": {
"max": {
"field": "price"
}
}
}
}
Bucket Aggregations
GET /products/_search
{
"aggs": {
"brands": {
"terms": {
"field": "brand.keyword"
},
"aggs": {
"avg_price": {
"avg": {
"field": "price"
}
}
}
}
}
}
Common Aggregations
- Metrics: sum, avg, min, max, stats
- Bucket: terms, range, date_histogram
- Pipeline: moving_avg, derivative
6. Mapping & Index Management
Create Index with Mapping
PUT /products
{
"mappings": {
"properties": {
"name": {
"type": "text",
"fields": {
"keyword": {
"type": "keyword"
}
}
},
"price": {
"type": "float"
},
"created_at": {
"type": "date"
},
"tags": {
"type": "keyword"
}
}
}
}
Field Data Types
- Text: Full-text searchable
- Keyword: Exact matches, sorting, aggregations
- Numeric: long, integer, short, byte, double, float
- Date: Date/time values
- Boolean: true/false
- Object: JSON objects
- Nested: Arrays of objects
Index Settings
PUT /products/_settings
{
"number_of_replicas": 2,
"refresh_interval": "30s"
}
7. Performance Optimization
Query Performance
- Use filters instead of queries when possible
- Limit result size with
sizeparameter - Use
source filteringto return only needed fields - Avoid wildcards at the beginning of terms
Example: Optimized Query
GET /products/_search
{
"_source": ["name", "price"],
"size": 10,
"query": {
"bool": {
"filter": [
{"term": {"brand.keyword": "Dell"}},
{"range": {"price": {"gte": 500}}}
]
}
}
}
Indexing Performance
- Bulk indexing for multiple documents
- Disable refresh during bulk operations
- Use appropriate shard count
- Configure proper heap size (50% of RAM, max 32GB)
8. Cluster Management
Cluster Health
GET /_cluster/health
States:
- Green: All shards allocated
- Yellow: Primary allocated, replicas not
- Red: Some primaries not allocated
Node Stats
GET /_nodes/stats
Index Stats
GET /products/_stats
9. Advanced Features
Analyzers
PUT /products
{
"settings": {
"analysis": {
"analyzer": {
"custom_analyzer": {
"tokenizer": "standard",
"filter": ["lowercase", "stop"]
}
}
}
}
}
Highlighting
GET /products/_search
{
"query": {
"match": {"name": "laptop"}
},
"highlight": {
"fields": {
"name": {}
}
}
}
Scroll API (Large Result Sets)
POST /products/_search?scroll=1m
{
"size": 100,
"query": {"match_all": {}}
}
Reindex API
POST /_reindex
{
"source": {
"index": "old_products"
},
"dest": {
"index": "new_products"
}
}
10. Use Cases & Design Patterns
When to Use Elasticsearch
Good for:
- Full-text search applications
- Log and event data analysis (ELK stack)
- Real-time analytics dashboards
- Content discovery and recommendations
- Application performance monitoring (APM)
- Security information and event management (SIEM)
Not ideal for:
- Primary transactional data storage
- Complex relational queries with joins
- Strong consistency requirements
- Frequent updates to individual documents
Common Design Patterns
Log Aggregation (ELK Stack)
// Logstash configuration for log parsing
{
"index": "logs-2024.01.15",
"document": {
"@timestamp": "2024-01-15T10:30:00Z",
"level": "ERROR",
"message": "Database connection failed",
"service": "api-gateway",
"host": "server-01",
"tags": ["error", "database"]
}
}
E-commerce Search
// Product search with facets
GET /products/_search
{
"query": {
"bool": {
"must": [
{"match": {"name": "laptop"}}
],
"filter": [
{"term": {"category.keyword": "electronics"}},
{"range": {"price": {"gte": 500, "lte": 2000}}}
]
}
},
"aggs": {
"brands": {
"terms": {"field": "brand.keyword"}
},
"price_ranges": {
"range": {
"field": "price",
"ranges": [
{"to": 500},
{"from": 500, "to": 1000},
{"from": 1000}
]
}
}
}
}
Time Series Analytics
// Website analytics dashboard
{
"query": {
"range": {
"@timestamp": {
"gte": "now-24h"
}
}
},
"aggs": {
"page_views_over_time": {
"date_histogram": {
"field": "@timestamp",
"interval": "1h"
},
"aggs": {
"unique_visitors": {
"cardinality": {
"field": "user_id"
}
}
}
}
}
}
Architecture Patterns
Hot-Warm-Cold Architecture
// Hot nodes (recent data, fast hardware)
PUT /logs-2024.01.15
{
"settings": {
"index.routing.allocation.require.data": "hot",
"number_of_shards": 3,
"number_of_replicas": 1
}
}
// Warm nodes (older data, slower hardware)
PUT /logs-2024.01.01/_settings
{
"index.routing.allocation.require.data": "warm",
"number_of_replicas": 0
}
Index Templates
PUT /_template/logs_template
{
"index_patterns": ["logs-*"],
"settings": {
"number_of_shards": 1,
"number_of_replicas": 1
},
"mappings": {
"properties": {
"@timestamp": {"type": "date"},
"message": {"type": "text"},
"level": {"type": "keyword"}
}
}
}
Performance Optimization Strategies
Sharding Strategy
- Small indices: 1 shard
- Medium indices: 3-5 shards
- Large indices: Number of shards = (data size in GB) / 30-50
- Rule of thumb: Shard size between 10-50 GB
Memory Management
// JVM heap configuration
-Xms16g -Xmx16g // Set min = max heap
// File system cache gets remaining memory
Query Optimization
// Use filters for exact matches (cached)
"filter": [
{"term": {"status": "published"}},
{"range": {"date": {"gte": "2024-01-01"}}}
]
// Use queries for relevance scoring
"query": {
"match": {"content": "search terms"}
}
Troubleshooting Common Issues
Split Brain Prevention
# elasticsearch.yml
discovery.zen.minimum_master_nodes: 2 # (master_eligible_nodes / 2) + 1
Circuit Breaker Settings
PUT /_cluster/settings
{
"transient": {
"indices.breaker.fielddata.limit": "40%",
"indices.breaker.request.limit": "60%"
}
}
Slow Query Analysis
// Enable slow log
PUT /index/_settings
{
"index.search.slowlog.threshold.query.warn": "10s",
"index.search.slowlog.threshold.fetch.warn": "1s"
}
11. Security Features
Basic Authentication
curl -u elastic:password -X GET "localhost:9200/"
Role-Based Access Control
POST /_security/role/read_only
{
"indices": [{
"names": ["products"],
"privileges": ["read"]
}]
}
12. Monitoring & Debugging
Explain API
GET /products/_explain/1
{
"query": {
"match": {"name": "laptop"}
}
}
Profile API
GET /products/_search
{
"profile": true,
"query": {
"match": {"name": "laptop"}
}
}
Slow Log Configuration
PUT /products/_settings
{
"index.search.slowlog.threshold.query.warn": "10s",
"index.search.slowlog.threshold.fetch.warn": "1s"
}
13. Quick Reference
REST API Patterns
GET /_search # Search all indices
GET /index/_search # Search specific index
GET /index/_doc/id # Get document
PUT /index/_doc/id # Index document
POST /index/_update/id # Update document
DELETE /index/_doc/id # Delete document
GET /_cat/indices # List indices
GET /_cat/nodes # List nodes
GET /_cluster/health # Cluster health
Query Context vs Filter Context
- Query Context: Calculates relevance score
- Filter Context: Yes/no match, cached, no scoring
Memory Guidelines
- Heap: 50% of available RAM, max 32GB
- File System Cache: Remaining 50% for Lucene
- min_heap = max_heap for production
Remember: Elasticsearch is about trade-offs between speed, accuracy, and resource usage. Always consider your specific use case when making design decisions.