Course capstone: declare a disaster, fail over, measure it, and fail back
Task
Run a complete disaster recovery exercise on the lab. A web service and its data run at the primary site with replication to a warm recovery site. Declare a disaster, fail over using the DR runbook, measure RTO and RPO against targets set beforehand, run from the recovery site while new data arrives, then fail back to the repaired primary without losing that data. Finish with the exercise report a business continuity plan would receive.
Steps
- Before starting, write
lab/dr/targets.txtwithrto_minutes:andrpo_minutes:, andlab/dr/runbook.mdcovering declaration, failover steps in dependency order, DNS change, verification, and failback. - Set up the service: nginx on both servers, content and an application data folder written to continuously on the primary, replicated to the recovery site every minute, and
www.lab.internalpointing at the primary with a 60-second TTL. Run a probe from the host that requests the site every five seconds and logs one line per request tolab/dr/probe.log: the Unix time, a space, and the HTTP status,000when nothing answered (curl -s -o /dev/null -w '%{http_code}'prints it that way). - Power off lin-srv and log
disasterinlab/dr/log.csvwith headertime,event. Declare the disaster, follow the runbook to fail over, and logdeclared,dns_changedandverifiedas they happen. - Run from the recovery site for ten minutes while new data keeps arriving there. Then start lin-srv, reverse replication to bring it up to date, cut back over in a planned window, and log
failback_startandfailback_verified. - Count the files written at the recovery site during the outage and confirm all of them exist on the primary after failback. Record the count and the check in
lab/dr/failback-data.txt, including the wordsevery fileor0 missingif none were lost. - Write
lab/dr/report.txtwithrto_actual_minutes:,rpo_actual_seconds:,targets_met:, the number of runbook gaps, and three actions for the business continuity plan, each on its own line starting1.,2.and3..
Verify
These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.
grep -Ec '^(rto_minutes|rpo_minutes): [0-9]+' lab/dr/targets.txt
awk -F, 'NR>1 {print $2}' lab/dr/log.csv | paste -sd' ' -
awk '$2 != 200 {f++} END {print f+0" failed probe(s) of "NR}' lab/dr/probe.log
grep -Eic 'match|every file|0 missing|none missing' lab/dr/failback-data.txt
grep -Ec '^(rto_actual_minutes|rpo_actual_seconds|targets_met):' lab/dr/report.txt
grep -Ec '^[0-9-]+[.)]? ' lab/dr/report.txt
The log records disaster, declaration, DNS change, verification and both failback events in order; the probe shows an outage window matching your measured RTO; every file written at the recovery site survived failback; and the report states the actual figures against the targets with at least three actions. If data written during the outage was lost on failback, the reverse replication ran the wrong way -- the failback failure the lesson warns is more common than a failed failover.
Notes
This exercise is the one this whole course has been building toward: hardware redundancy, storage, networking, DNS, security, documentation and backups all had to work for it to pass. Repeat it after any significant change to the lab, because a recovery plan is only as current as its last successful test.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.