Domain 1 capstone: replace a failed mirror member live, with the paperwork
Task
Bring the hardware domain together on one realistic job. A server's RAID 1 boot data mirror has lost a disk. Confirm the failure from the logs, raise a change record, hot-remove the failed member, hot-add a replacement, rebuild, prove the data is intact, and update the inventory -- the complete replacement of a part, done without taking the service down.
Steps
- Build a RAID 1 array from the two extra disks with
mdadm, create a file system, mount it at/srv/data, and write 100 files of random data. Savesha256sum /srv/data/*tolab/capstone1/before.txt. - Record both member disks with their serial numbers or virtual disk names in
lab/capstone1/inventory.csvwith headerslot,device,serial,status. - Simulate the failure with
mdadm --failon one member. Savemdadm --detailtolab/capstone1/failed.txtand the matching kernel log lines tolab/capstone1/log.txt. - Write the change record in
lab/capstone1/change.txtwith the lineswhat:,why:,risk:,rollback:andwindow:. - Remove the failed member from the array, release it from the kernel, and detach it in the hypervisor while the file system stays mounted. Attach a new disk, add it to the array, and save
/proc/mdstatduring the rebuild tolab/capstone1/rebuilding.txt. - When the rebuild finishes, save
mdadm --detailtolab/capstone1/rebuilt.txt, the checksums tolab/capstone1/after.txt, and update the inventory with the new disk and areplacedrow for the old one.
Verify
These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.
grep -Eic 'degraded|faulty' lab/capstone1/failed.txt
grep -Eic 'recovery|resync' lab/capstone1/rebuilding.txt
grep -Eic 'State : (clean|active) *$' lab/capstone1/rebuilt.txt # mdadm ends the line with a space
diff lab/capstone1/before.txt lab/capstone1/after.txt && echo all 100 checksums unchanged
wc -l < lab/capstone1/after.txt
grep -Ec '^(what|why|risk|rollback|window):' lab/capstone1/change.txt
grep -c 'replaced' lab/capstone1/inventory.csv
The array was degraded, then recovering, then clean; all 100 checksums are identical before and after; the change record has all five fields; and the inventory records the replacement. The file system stayed mounted throughout, which is the whole point of redundancy plus hot-swap: the part changed and the service did not notice.
Notes
On a real server the order is the same, with two additions: check the replacement's compatibility and firmware before inserting it, and take ESD precautions. If the rebuild fails, the change record's rollback line is what you reach for -- which is why it is written before the work starts.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.