Domain 1 capstone: replace a failed mirror member live, with the paperwork

capstone · 90 min · Objective 1.3

Task

Bring the hardware domain together on one realistic job. A server's RAID 1 boot data mirror has lost a disk. Confirm the failure from the logs, raise a change record, hot-remove the failed member, hot-add a replacement, rebuild, prove the data is intact, and update the inventory -- the complete replacement of a part, done without taking the service down.

Steps

  1. Build a RAID 1 array from the two extra disks with mdadm, create a file system, mount it at /srv/data, and write 100 files of random data. Save sha256sum /srv/data/* to lab/capstone1/before.txt.
  2. Record both member disks with their serial numbers or virtual disk names in lab/capstone1/inventory.csv with header slot,device,serial,status.
  3. Simulate the failure with mdadm --fail on one member. Save mdadm --detail to lab/capstone1/failed.txt and the matching kernel log lines to lab/capstone1/log.txt.
  4. Write the change record in lab/capstone1/change.txt with the lines what:, why:, risk:, rollback: and window:.
  5. Remove the failed member from the array, release it from the kernel, and detach it in the hypervisor while the file system stays mounted. Attach a new disk, add it to the array, and save /proc/mdstat during the rebuild to lab/capstone1/rebuilding.txt.
  6. When the rebuild finishes, save mdadm --detail to lab/capstone1/rebuilt.txt, the checksums to lab/capstone1/after.txt, and update the inventory with the new disk and a replaced row for the old one.

Verify

These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.

grep -Eic 'degraded|faulty' lab/capstone1/failed.txt
grep -Eic 'recovery|resync' lab/capstone1/rebuilding.txt
grep -Eic 'State : (clean|active) *$' lab/capstone1/rebuilt.txt    # mdadm ends the line with a space
diff lab/capstone1/before.txt lab/capstone1/after.txt && echo all 100 checksums unchanged
wc -l < lab/capstone1/after.txt
grep -Ec '^(what|why|risk|rollback|window):' lab/capstone1/change.txt
grep -c 'replaced' lab/capstone1/inventory.csv

The array was degraded, then recovering, then clean; all 100 checksums are identical before and after; the change record has all five fields; and the inventory records the replacement. The file system stayed mounted throughout, which is the whole point of redundancy plus hot-swap: the part changed and the service did not notice.

Notes

On a real server the order is the same, with two additions: check the replacement's compatibility and firmware before inserting it, and take ESD precautions. If the rebuild fails, the change record's rollback line is what you reach for -- which is why it is written before the work starts.

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.