Imagine you run a website that suddenly gets busy. Hundreds of people arrive at once. Then thousands. One server has to deal with all those requests, and pretty soon it starts struggling. Pages slow down. Some requests fail. Nobody’s having a great time.

Load balancing steps in before things get that messy. It spreads incoming traffic across several servers instead of sending everything to one machine. The result feels simple from the outside. You ask for a webpage, and it loads. You don’t need to know which server handled your request.

What Happens When a Request Arrives?

A load balancer usually sits between users and the servers running an application. So when your browser sends a request, it reaches the load balancer first. The load balancer then checks which server should handle it and forwards the request there.

The trick is deciding where to send it. A basic setup might rotate requests between servers. The first request goes to Server A. The next goes to Server B. Then Server C gets a turn. This approach is called round robin.

But real systems often need more than taking turns. One server could be busy while another is sitting around with plenty of capacity. A smarter load balancer looks at current server conditions before making its choice.

Checking Server Health

There’s another important job happening quietly in the background. The load balancer checks whether servers are actually working.

If one server stops responding, the load balancer can stop sending new requests to it. Traffic moves toward healthy servers instead. Users might never notice that something went wrong, which is exactly what you want.

• A dead server gets taken out of rotation, so new requests aren’t pushed into a black hole.

• Round robin is simple and predictable, although it doesn’t know whether one server is already having a rough afternoon.

• Health checks keep traffic away from failed machines, and that quiet little check matters more than it sounds.

How Does It Choose a Server?

Different load balancing methods make different choices. Round robin gives each server a turn. Another method sends traffic toward the server handling fewer active connections. There are also methods that consider response time or server capacity.

And some setups use a user’s location when deciding where traffic should go. Someone in Mumbai might reach a nearby data centre while another person connects somewhere else. That reduces the distance data has to travel, which can make the site feel quicker.

Keeping Sessions Consistent

Load balancers can deal with this using session persistence, often called sticky sessions. Requests from the same user are directed back to the same server for a period of time. It’s useful, though I prefer systems designed so any healthy server can handle a request. They’re usually easier to scale and maintain.

Why Load Balancing Matters

A good load balancing setup makes an application feel almost boring. Requests arrive. Servers handle them. Failed machines get avoided. Users keep clicking around without thinking about what’s happening underneath.