Domain 2 capstone: a highly available web service, measured in nines
Task
Build a small service that survives the loss of any one of its parts, and measure it. Two web servers behind a load balancer, the load balancer itself on a keepalived virtual IP, a DNS name for the service, a baseline, documentation, and a test that fails each component in turn while a probe measures what users would have seen. Finish by turning the measured downtime into an availability figure.
Steps
- Configure HAProxy on both nodes with both web servers as backends and an HTTP health check, and keepalived with the virtual IP, tracking the HAProxy process so the address moves if HAProxy stops. Save both configurations to
lab/capstone2/. - Create
www.lab.internalon win-srv pointing at 192.168.56.100 and save thediganswer tolab/capstone2/dns.txt. - Write
lab/capstone2/spof.csvwith headercomponent,redundant,how, where redundant isyesorno, listing every component a request passes through, including DNS. - Start a probe from win-srv or the host that requests
http://www.lab.internal/once a second and appends one line tolab/capstone2/probe.log: the Unix time, a space, and the HTTP status,000when nothing answered. - Fail each component in turn, restoring it before the next: stop nginx on lin-b; stop HAProxy on the node holding the virtual IP; shut down that node entirely. Record each event's start and end in
lab/capstone2/events.csvwith headerevent,start,end. - Stop the probe after ten minutes. Compute availability as successful requests divided by total, and record
requests:,failed:,availability_pct:andnines:inlab/capstone2/result.txt.
Verify
These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.
grep -Ec 'track_script|vrrp_script' lab/capstone2/*
grep -c '192.168.56.100' lab/capstone2/dns.txt
awk -F, 'NR>1 && $2=="no" {print "single point of failure: "$1}' lab/capstone2/spof.csv
wc -l < lab/capstone2/probe.log
awk '$2 != 200 {f++} END {print f+0" failed request(s) of "NR; printf "%.3f%%%s", 100*(NR-f)/NR, ORS}' lab/capstone2/probe.log
awk -F, 'NR>1 {n++} END {print n" failure event(s)"}' lab/capstone2/events.csv
grep -E '^availability_pct: [0-9.]+' lab/capstone2/result.txt
The probe log has about 600 lines, three failure events are recorded, and only a few requests failed -- the seconds while the virtual IP moved -- so availability is high but not 100 per cent. The single-point-of-failure list should name win-srv's DNS, since the lab has only one DNS server: exactly the kind of gap the lesson says tracing a request end to end will find.
Notes
Ten minutes of measurement cannot establish a real availability figure; it shows the method. Four nines allows about 52 minutes of downtime a year, so a service that loses three seconds per failover can afford many failovers -- and very few of the untested kind that take an hour.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.