Definition: A load balancer is a component that receives incoming traffic and distributes it across multiple servers/instances instead of sending everything to one server.
It helps applications handle more traffic, improve performance, and remain available when one server becomes unhealthy. (Microsoft Learn)
Simple example
Without a load balancer:
Users
│
▼
Server 1If 10,000 users arrive, Server 1 has to handle everyone.
With a load balancer:
┌──→ Server 1
Users → Load ────┼──→ Server 2
Balancer └──→ Server 3The load balancer decides where each incoming request should go. It can also check server health and stop sending traffic to an unhealthy server. (AWS Documentation)
3 examples
1. E-commerce website
1000 requests
↓
Load Balancer
┌────┼────┐
↓ ↓ ↓
S1 S2 S3Instead of one server becoming overloaded during a sale, requests are distributed across several servers.
2. Server failure
Load Balancer
├── Server 1 ✓
├── Server 2 ❌
└── Server 3 ✓If Server 2 becomes unhealthy, the load balancer can stop sending requests to it and continue using the healthy servers. (AWS Documentation)
3. Scaling
You have:
Load Balancer
├── Server 1
└── Server 2Traffic increases, so you add Server 3:
Load Balancer
├── Server 1
├── Server 2
└── Server 3You can add/remove backend resources as demand changes without changing how clients reach the application. (AWS Documentation)
How does it decide where to send requests?
Common strategies include:
-
Round Robin: Server 1 → Server 2 → Server 3 → Server 1…
-
Least Connections: Send the request to the server currently handling the fewest connections.
-
Health-based: Don’t send traffic to unhealthy servers.
-
Weighted: Give more traffic to more powerful servers. (Amazon Web Services, Inc.)
Easy way to remember:
Load balancer = traffic manager for your servers.
Instead of 1000 users → 1 server, it makes 1000 users → multiple servers.