Domain 2 capstone: a highly available web service, measured in nines

capstone · 120 min · Objective 2.4

Task

Build a small service that survives the loss of any one of its parts, and measure it. Two web servers behind a load balancer, the load balancer itself on a keepalived virtual IP, a DNS name for the service, a baseline, documentation, and a test that fails each component in turn while a probe measures what users would have seen. Finish by turning the measured downtime into an availability figure.

Steps

  1. Configure HAProxy on both nodes with both web servers as backends and an HTTP health check, and keepalived with the virtual IP, tracking the HAProxy process so the address moves if HAProxy stops. Save both configurations to lab/capstone2/.
  2. Create www.lab.internal on win-srv pointing at 192.168.56.100 and save the dig answer to lab/capstone2/dns.txt.
  3. Write lab/capstone2/spof.csv with header component,redundant,how, where redundant is yes or no, listing every component a request passes through, including DNS.
  4. Start a probe from win-srv or the host that requests http://www.lab.internal/ once a second and appends one line to lab/capstone2/probe.log: the Unix time, a space, and the HTTP status, 000 when nothing answered.
  5. Fail each component in turn, restoring it before the next: stop nginx on lin-b; stop HAProxy on the node holding the virtual IP; shut down that node entirely. Record each event's start and end in lab/capstone2/events.csv with header event,start,end.
  6. Stop the probe after ten minutes. Compute availability as successful requests divided by total, and record requests:, failed:, availability_pct: and nines: in lab/capstone2/result.txt.

Verify

These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.

grep -Ec 'track_script|vrrp_script' lab/capstone2/*
grep -c '192.168.56.100' lab/capstone2/dns.txt
awk -F, 'NR>1 && $2=="no" {print "single point of failure: "$1}' lab/capstone2/spof.csv
wc -l < lab/capstone2/probe.log
awk '$2 != 200 {f++} END {print f+0" failed request(s) of "NR; printf "%.3f%%%s", 100*(NR-f)/NR, ORS}' lab/capstone2/probe.log
awk -F, 'NR>1 {n++} END {print n" failure event(s)"}' lab/capstone2/events.csv
grep -E '^availability_pct: [0-9.]+' lab/capstone2/result.txt

The probe log has about 600 lines, three failure events are recorded, and only a few requests failed -- the seconds while the virtual IP moved -- so availability is high but not 100 per cent. The single-point-of-failure list should name win-srv's DNS, since the lab has only one DNS server: exactly the kind of gap the lesson says tracing a request end to end will find.

Notes

Ten minutes of measurement cannot establish a real availability figure; it shows the method. Four nines allows about 52 minutes of downtime a year, so a service that loses three seconds per failover can afford many failovers -- and very few of the untested kind that take an hour.

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.