Course capstone: declare a disaster, fail over, measure it, and fail back

capstone · 150 min · Objective 4.4

Task

Run a complete disaster recovery exercise on the lab. A web service and its data run at the primary site with replication to a warm recovery site. Declare a disaster, fail over using the DR runbook, measure RTO and RPO against targets set beforehand, run from the recovery site while new data arrives, then fail back to the repaired primary without losing that data. Finish with the exercise report a business continuity plan would receive.

Steps

  1. Before starting, write lab/dr/targets.txt with rto_minutes: and rpo_minutes:, and lab/dr/runbook.md covering declaration, failover steps in dependency order, DNS change, verification, and failback.
  2. Set up the service: nginx on both servers, content and an application data folder written to continuously on the primary, replicated to the recovery site every minute, and www.lab.internal pointing at the primary with a 60-second TTL. Run a probe from the host that requests the site every five seconds and logs one line per request to lab/dr/probe.log: the Unix time, a space, and the HTTP status, 000 when nothing answered (curl -s -o /dev/null -w '%{http_code}' prints it that way).
  3. Power off lin-srv and log disaster in lab/dr/log.csv with header time,event. Declare the disaster, follow the runbook to fail over, and log declared, dns_changed and verified as they happen.
  4. Run from the recovery site for ten minutes while new data keeps arriving there. Then start lin-srv, reverse replication to bring it up to date, cut back over in a planned window, and log failback_start and failback_verified.
  5. Count the files written at the recovery site during the outage and confirm all of them exist on the primary after failback. Record the count and the check in lab/dr/failback-data.txt, including the words every file or 0 missing if none were lost.
  6. Write lab/dr/report.txt with rto_actual_minutes:, rpo_actual_seconds:, targets_met:, the number of runbook gaps, and three actions for the business continuity plan, each on its own line starting 1., 2. and 3..

Verify

These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.

grep -Ec '^(rto_minutes|rpo_minutes): [0-9]+' lab/dr/targets.txt
awk -F, 'NR>1 {print $2}' lab/dr/log.csv | paste -sd' ' -
awk '$2 != 200 {f++} END {print f+0" failed probe(s) of "NR}' lab/dr/probe.log
grep -Eic 'match|every file|0 missing|none missing' lab/dr/failback-data.txt
grep -Ec '^(rto_actual_minutes|rpo_actual_seconds|targets_met):' lab/dr/report.txt
grep -Ec '^[0-9-]+[.)]? ' lab/dr/report.txt

The log records disaster, declaration, DNS change, verification and both failback events in order; the probe shows an outage window matching your measured RTO; every file written at the recovery site survived failback; and the report states the actual figures against the targets with at least three actions. If data written during the outage was lost on failback, the reverse replication ran the wrong way -- the failback failure the lesson warns is more common than a failed failover.

Notes

This exercise is the one this whole course has been building toward: hardware redundancy, storage, networking, DNS, security, documentation and backups all had to work for it to pass. Repeat it after any significant change to the lab, because a recovery plan is only as current as its last successful test.

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.