Summary
Load testing evaluates system performance under expected and peak user loads, identifying bottlenecks before production. Master performance metrics, testing tools (JMeter, K6, Locust), load patterns, bottleneck analysis, and result interpretation to ensure applications meet performance SLAs and scale effectively.
1. What is Load Testing?
Definition: A type of performance testing that simulates real-world load on any software, application, or website to determine how it behaves under expected and peak load conditions.
Key Purpose: Identify performance bottlenecks before production deployment.
2. Load Testing vs Other Performance Tests
| Test Type | Purpose | Load Level |
|---|---|---|
| Load Testing | Test expected user load | Normal to peak expected |
| Stress Testing | Find breaking point | Beyond normal capacity |
| Spike Testing | Test sudden load increase | Sudden spike then normal |
| Volume Testing | Test with large data | Normal load, high data |
| Soak Testing | Test sustained load | Normal load, extended time |
3. Key Metrics to Monitor
Response Time Metrics
- Average Response Time: Mean time for all requests
- Peak Response Time: Maximum response time recorded
- Percentiles (P90, P95, P99): Response time for 90%, 95%, 99% of requests
Throughput Metrics
- Requests per Second (RPS): Number of requests handled per second
- Transactions per Second (TPS): Business transactions completed per second
- Bandwidth: Data transferred per second
Resource Metrics
- CPU Utilization: Percentage of CPU used
- Memory Usage: RAM consumption
- Disk I/O: Read/write operations
- Network I/O: Network bandwidth usage
Error Metrics
- Error Rate: Percentage of failed requests
- Error Types: 4xx, 5xx HTTP errors
- Timeout Rate: Requests that exceeded time limit
4. Load Testing Process
graph LR
A[1. Define Goals] --> B[2. Create Test Plan]
B --> C[3. Setup Test Environment]
C --> D[4. Create Test Scripts]
D --> E[5. Execute Tests]
E --> F[6. Monitor & Collect Data]
F --> G[7. Analyze Results]
G --> H[8. Report Findings]
Detailed Steps:
Define Goals
- Expected concurrent users
- Response time SLAs
- Throughput requirements
Create Test Plan
- Test scenarios
- Load patterns
- Success criteria
Setup Environment
- Production-like setup
- Monitoring tools
- Test data preparation
Create Scripts
- User workflows
- Think time
- Data parameterization
Execute Tests
- Gradual ramp-up
- Sustained load
- Gradual ramp-down
5. Popular Load Testing Tools
Open Source
- JMeter: Java-based, GUI and CLI support
- Gatling: Scala-based, code-as-configuration
- Locust: Python-based, distributed testing
- K6: JavaScript-based, developer-centric
Commercial
- LoadRunner: Enterprise solution by Micro Focus
- NeoLoad: Continuous performance testing
- BlazeMeter: Cloud-based, CI/CD integration
6. JMeter Basic Example
<!-- Simple HTTP Test Plan -->
<TestPlan>
<ThreadGroup>
<numThreads>100</numThreads>
<rampUp>10</rampUp>
<duration>300</duration>
<HTTPSampler>
<domain>api.example.com</domain>
<path>/users</path>
<method>GET</method>
</HTTPSampler>
<ResponseAssertion>
<responseCode>200</responseCode>
<responseTime>1000</responseTime>
</ResponseAssertion>
</ThreadGroup>
</TestPlan>
7. Locust Example Script
from locust import HttpUser, task, between
class WebsiteUser(HttpUser):
wait_time = between(1, 3) # Think time
@task(3)
def view_products(self):
self.client.get("/products")
@task(1)
def view_product_details(self):
product_id = random.randint(1, 1000)
self.client.get(f"/products/{product_id}")
def on_start(self):
# Login once per user
self.client.post("/login", {
"username": "testuser",
"password": "testpass"
})
8. K6 Example Script
import http from 'k6/http';
import { check, sleep } from 'k6';
export let options = {
stages: [
{ duration: '2m', target: 100 }, // Ramp up
{ duration: '5m', target: 100 }, // Stay at 100
{ duration: '2m', target: 0 }, // Ramp down
],
thresholds: {
http_req_duration: ['p(95)<500'], // 95% under 500ms
http_req_failed: ['rate<0.1'], // Error rate < 10%
},
};
export default function() {
let response = http.get('https://api.example.com/users');
check(response, {
'status is 200': (r) => r.status === 200,
'response time < 500ms': (r) => r.timings.duration < 500,
});
sleep(1); // Think time
}
9. Common Bottlenecks & Solutions
| Bottleneck | Symptoms | Solutions |
|---|---|---|
| Database | Slow queries, locks | Query optimization, indexing, caching |
| Application Server | High CPU, memory | Code optimization, horizontal scaling |
| Network | High latency, packet loss | CDN, compression, connection pooling |
| Third-party APIs | Timeouts, rate limits | Caching, circuit breakers, async processing |
10. Best Practices
Test Design
- ✅ Use realistic test data
- ✅ Include think time between requests
- ✅ Simulate realistic user behavior
- ✅ Test with production-like environment
Execution
- ✅ Start with small load, gradually increase
- ✅ Run tests multiple times for consistency
- ✅ Monitor both client and server metrics
- ✅ Test during different times/conditions
Analysis
- ✅ Look for trends, not just averages
- ✅ Correlate metrics (CPU vs response time)
- ✅ Identify bottlenecks systematically
- ✅ Compare with baseline results
11. Critical Interview Topics
Virtual User Calculation
Formula: Virtual Users = (Hourly Sessions × Average Session Duration) / 3600
Example Calculation:
- 10,000 sessions/hour
- 5 minutes average session
- VU = (10,000 × 300) / 3600 = 833 concurrent users
Considerations:
- Peak hour traffic patterns
- Geographic distribution
- User behavior variations
Concurrent vs Simultaneous Users
- Concurrent Users: Total users with active sessions (including think time)
- Simultaneous Users: Users making requests at exact same moment
- Conversion Rule: Simultaneous ≈ 10-20% of concurrent
- Impact: Server resources, connection pools, database locks
Dynamic Data Handling Strategies
Parameterization
- CSV data files for user credentials
- Database queries for test data
- Random data generation functions
Correlation
// JMeter Example
Regular Expression Extractor:
Reference Name: sessionId
RegEx: sessionId=([^&]+)
Template: $1$
// Usage: ${sessionId}
Data Management
- Unique data per virtual user
- Data recycling strategies
- Cache simulation considerations
Little's Law Application
Formula: L = λ × W
- L: Average number of users in system
- λ: Average arrival rate (users/second)
- W: Average time in system (seconds)
Practical Uses:
- Validate test results consistency
- Capacity planning calculations
- Queue theory applications
- Performance baseline establishment
Memory Leak Detection
Detection Techniques
- Heap Monitoring: Track heap usage trends over time
- GC Analysis: Frequent full GC indicates issues
- Object Growth: Monitor object creation/destruction rates
- Profiler Integration: Use during load tests for deep analysis
Warning Signs
- Continuously increasing memory usage
- Declining throughput over time
- Increasing response times
- OutOfMemoryError occurrences
12. Load Test Report Components
Executive Summary
- Pass/Fail status
- Key findings
- Recommendations
Test Configuration
- Environment details
- Test scenarios
- Load pattern
Results
- Response times (with percentiles)
- Throughput achieved
- Error rates
- Resource utilization
Analysis
- Bottlenecks identified
- Performance vs SLAs
- Scalability insights
Recommendations
- Short-term fixes
- Long-term improvements
- Capacity planning
13. CI/CD Integration
# Example: GitLab CI with K6
load_test:
stage: performance
script:
- k6 run --out cloud script.js
rules:
- if: '$CI_COMMIT_BRANCH == "main"'
artifacts:
reports:
performance: k6-results.json
14. Cloud vs On-Premise Load Testing
| Aspect | Cloud | On-Premise |
|---|---|---|
| Scalability | Unlimited | Limited by hardware |
| Cost | Pay-per-use | Fixed infrastructure |
| Geographic Distribution | Easy | Complex |
| Security | Shared responsibility | Full control |
| Setup Time | Minutes | Days/Weeks |
15. Quick Reference Formulas
Think Time = (Total Test Duration × VUsers - Total Requests Time) / Total Requests
Pacing = (Test Duration × VUsers) / Total Transactions Required
Throughput = Total Requests / Test Duration
Error Rate = (Failed Requests / Total Requests) × 100
Apdex Score = (Satisfied + Tolerating/2) / Total Samples
16. Red Flags in Load Test Results
⚠️ Response time increases linearly with load → Synchronization issues
⚠️ Sudden spike in errors at specific load → Resource limit reached
⚠️ Memory continuously increasing → Memory leak
⚠️ CPU 100% but low throughput → Inefficient code
⚠️ Database connections maxed out → Connection pool issues
Interview Success Strategies
Technical Excellence
- Metrics Mastery: Understand percentiles (P50, P90, P99) over averages
- Tool Proficiency: Hands-on experience with at least 2 tools
- Scripting Skills: Demonstrate correlation, parameterization
- Analysis Capability: Identify bottlenecks from metrics
- Cloud Experience: Distributed load generation, auto-scaling
Business Alignment
- Impact Communication: "20% slower checkout = 5% revenue loss"
- Cost-Benefit Analysis: Balance performance vs infrastructure costs
- Risk Assessment: Identify critical user journeys
- SLA Understanding: Response time, availability targets
- Stakeholder Language: Translate technical metrics to business impact
Problem-Solving Approach
- Systematic Investigation: Isolate → Investigate → Resolve
- Baseline Establishment: Always compare against known good state
- Incremental Testing: Start small, gradually increase load
- Root Cause Analysis: Don't just treat symptoms
- Documentation: Clear reports with actionable recommendations
Key Discussion Points
- Production Readiness: How to validate before release
- Continuous Monitoring: APM tools, alerting strategies
- Capacity Planning: Growth projections, scaling strategies
- Performance Budget: Setting and maintaining limits
- DevOps Integration: Shift-left performance testing
Questions to Ask Interviewers
- What are the current performance SLAs?
- What is the expected user growth?
- Are there specific performance pain points?
- What monitoring tools are in place?
- How is performance testing integrated in CI/CD?