Load balancing, and keeping a service up
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
The previous lesson moved a service to a surviving server when one failed. A load balancer goes further: it spreads work across several servers all the time, so no single server carries everything, and a failed server simply stops receiving requests. Most websites and many business applications stay available this way.
This is the capstone of the administration domain because it pulls the earlier lessons together. A load balancer is only as good as its health checks, the servers behind it are only useful if they share state sensibly, and none of it matters if the power, network or storage underneath is a single point of failure. The lesson ends with the language used to promise availability, and what those promises actually allow.
The lesson
Round robin, least connections and weighted balancing
A load balancer sits in front of a group of servers, often called a pool or server farm. Clients connect to one address, a virtual IP, and the load balancer passes each request to one of the servers behind it. Load balancers can be dedicated appliances, software such as HAProxy or nginx, or a service from a cloud provider.
It decides where each request goes using an algorithm:
- Round robin sends requests to each server in turn. It is simple and works well when servers are identical and requests are similar in cost.
- Least connections sends each new request to the server with the fewest active connections. It suits workloads where some requests take far longer than others, because a server tied up with slow requests stops receiving new ones.
- Weighted versions of either give some servers a larger share, in proportion to a weight the administrator sets. A new server with twice the capacity of the old ones can be given twice the weight.
A cruder approach is DNS round robin, publishing several A records for one name so clients spread themselves across the addresses. It costs nothing, but DNS knows nothing about whether a server is working, and caching means clients keep using a dead address until their cached answer expires.
Health checks that detect a real failure
A load balancer must stop sending traffic to a server that has failed. It finds out using health checks, regular tests of each server in the pool. A server that fails enough checks in a row is marked down and taken out of rotation; when it passes again, it is returned.
Health checks come in levels, and the level decides what they can detect:
- A ping proves only that the server's network stack responds.
- A TCP port check proves that something is listening on the port.
- An application check requests a real page or endpoint and examines the response, the HTTP status code and ideally the content.
A web server whose application has crashed, or which cannot reach its database, can still answer pings and accept connections on port 443 while returning nothing but errors. Only an application-level check catches that. The best checks request a dedicated health endpoint that tests the things the application depends on, and reports failure if any of them is broken.
Set the check interval and failure threshold with care. Too sensitive, and a brief slowdown removes a healthy server; too lenient, and users are sent to a dead server for minutes.
Session persistence, and what it costs
Many applications remember things about a user between requests: that they are logged in, or what is in their basket. If that session is stored on one web server and the next request goes to another, the user appears to have been logged out.
Session persistence, also called sticky sessions or affinity, fixes this by sending all of a user's requests to the same server, identified by a cookie the load balancer sets or by the client's IP address.
It has costs:
- Uneven load. Users stay attached to their server however busy it becomes, so the load balancer can no longer spread work freely.
- Lost sessions on failure. When a server fails, every session stored on it is lost, and those users are logged out or lose their work.
- Harder maintenance. A server cannot be removed until its sessions end, so maintenance involves draining it: stopping new sessions while existing ones finish.
- IP affinity is unreliable when many users share one public address behind NAT, which sends them all to the same server.
The better design is to make servers stateless, storing session data in a shared database or cache that every server can reach. Then any server can handle any request, and persistence is not needed.
Redundancy at every layer: power, network, storage and server
A load-balanced pool of servers removes the server as a single point of failure. It does nothing about everything else. A single point of failure is any component whose failure stops the service, and high availability means finding and removing them at every layer:
- Power: redundant power supplies in each server, fed from separate circuits and PDUs, backed by a UPS and ideally a generator.
- Network: teamed NICs connected to two different switches, redundant uplinks and routers, and, if the service is public, more than one internet connection.
- Storage: RAID within servers, and shared storage with redundant controllers and paths.
- Server: clustering or a load-balanced pool with enough spare capacity to lose a member.
- The load balancer itself: a single load balancer is a single point of failure in front of an otherwise redundant pool, so they are deployed in pairs, typically active-passive with a shared virtual IP.
The usual way to find the gaps is to trace one request from the client to the data and back, asking of every component on the way: if this fails, does the service stop? Often the surprise is something outside the servers entirely, such as a single DNS server, a certificate that is about to expire, or a database every server depends on.
Availability in nines, and the downtime each one allows
Availability is expressed as a percentage of time a service is up, usually as a number of nines. Each extra nine cuts the permitted downtime to a tenth:
| Availability | Downtime per year | Downtime per month (approx.) |
|---|---|---|
| 99% (two nines) | 3.65 days | 7.3 hours |
| 99.9% (three nines) | 8.76 hours | 43.8 minutes |
| 99.99% (four nines) | 52.6 minutes | 4.4 minutes |
| 99.999% (five nines) | 5.26 minutes | 26 seconds |
These figures appear in service level agreements (SLAs), which set the availability a provider promises and what happens if it falls short. Two points matter when reading them. First, check whether planned maintenance counts: many SLAs exclude it, so a service can be down for hours and still meet its target. Second, remember that each nine costs more than the last. Four nines leaves under an hour a year, which rules out any maintenance that stops the service and demands redundancy at every layer.
Availability also combines. A service that depends on several components in sequence can be no more available than their availabilities multiplied together: two components at 99.9 per cent each give about 99.8 per cent overall. Redundancy works the other way, which is why the previous topic matters as much as any single server.
Practise what you just read
1. A load balancer uses a TCP port check. A web server's application crashes but its port stays open. What happens?
Select one
Show answer
C. A port check only proves something is listening. An application check that requests a page and examines the response would detect the failure and remove the server from rotation.
2. Which balancing method suits a workload where some requests take far longer than others?
Select one
Show answer
A. Least connections sends new requests to the server with the fewest active ones, so a server busy with slow requests receives fewer new ones. Round robin ignores how busy each server is.
3. Roughly how much downtime per year does 99.99 per cent availability allow?
Select one
Show answer
D. Four nines allows 0.01 per cent of a year, about 52.6 minutes. Three nines allows about 8.76 hours and five nines about 5.26 minutes; each nine divides downtime by ten.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.