Hot-swappable components, and replacing parts safely
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
This lesson closes the hardware domain, and it is where the earlier lessons meet in one routine job: a component has failed, and it has to be replaced without taking the server, or anything else, down with it. The redundancy from the power lessons, the RAID arrays from the storage lessons and the management controller from the out-of-band lesson all exist so that this job can be done during working hours with the service still running.
The job is routine, but the failure modes are not small. Pull the wrong drive from a degraded array and the array is gone. Touch a board carelessly and it may fail weeks later. Most of what follows is about doing the simple thing without the expensive mistake.
The lesson
What can be replaced with the server running -- drives, power supplies, fans -- and what cannot
A hot-swappable component can be removed and replaced while the server is powered on and running, without interrupting its work. On a typical server, that means:
- drives in hot-plug bays, when they are members of a redundant array;
- power supplies, when the server has redundant ones;
- fans, on servers built with redundant, modular fan assemblies.
Other components are cold-swap: the server must be shut down and unplugged first. This normally includes processors, memory, the motherboard and most expansion cards. A few high-end systems support adding memory or cards while running, but it is the exception, and the exam assumes the general rule.
The important qualification is redundancy. Hot-swapping a power supply is safe only because the other supply carries the load while it is out. Remove the only working supply, or a drive from an array with no remaining redundancy, and the "hot-swap" is an outage. Before pulling anything, confirm that whatever remains can carry the server on its own.
Finding the failed part from LEDs and the management controller
Hot-swappable components have status lights, and a failed one normally shows an amber or fault LED on its drive carrier, power supply or fan. The server also has an overall health indicator on its front panel.
In a rack of identical servers, the first problem is finding the right server. The UID (unit identification) LED solves it: a blue light, on both the front and back of the server, that an administrator can switch on remotely through the management controller so whoever is in the data centre can see exactly which machine to work on.
The second problem is finding the right part. The management controller's interface and the RAID controller's software both report which bay, slot or power supply has failed, and most RAID controllers can blink a particular drive's LED on request. Use that locate function, and cross-check the bay number and the drive's serial number, before touching anything. In a degraded array, removing a healthy drive instead of the failed one removes the array's last copy of some data. It is one of the most damaging mistakes in server administration, and entirely avoidable.
ESD precautions and handling components
Electrostatic discharge (ESD) is the static shock that jumps from a person to a metal object. Components can be damaged by discharges far too small to feel, and the damage is not always immediate: a component may work after handling and fail weeks later, which makes the cause hard to trace.
The precautions are straightforward:
- Wear an anti-static wrist strap connected to the server chassis or an earthed point, so you and the server are at the same potential.
- Work on an anti-static mat where possible.
- Keep replacement parts in their anti-static bags until the moment they are fitted, and put removed parts straight into one.
- Handle boards and cards by their edges, never by the gold contacts or components.
Hot-swap drives and power supplies in their carriers are less exposed than bare memory or boards, but the habit costs nothing and the same wrist strap protects the more sensitive parts when you do open a case. Cold-swap work also needs the server shut down and unplugged first, both for the components and for your own safety.
Replacing a drive in an array and watching the rebuild finish
Replacing a failed drive in a redundant array follows a steady sequence:
- Confirm the state. Check that the array is degraded, still running on its remaining redundancy, rather than already failed, and identify the failed drive by bay and serial number.
- Check for a hot spare. If a hot spare has already taken over and rebuilt, the array may be healthy again, and the replacement drive will typically become the new spare.
- Locate the drive with the controller's blink function and confirm it matches.
- Remove the failed drive. Release the carrier and, for a hard disk, give it a moment to spin down before pulling it fully out.
- Insert the replacement. It must be compatible: the same interface and at least the same capacity, and ideally the model and firmware the vendor supports for that server.
- Watch the rebuild start. Most controllers begin automatically; some require the new drive to be assigned. Follow the rebuild's progress in the controller's software.
- Leave the array alone until it finishes. The array stays degraded until the rebuild completes, and removing or disturbing anything else during that window is how a single failure becomes a lost array.
- Verify that the array reports as optimal, and check the event log for any further warnings.
Spares, warranty and returning a failed part
Replacement is only quick if the part is to hand. Keeping spares of the components most likely to fail, particularly drives and power supplies, on site means a failure can be fixed in minutes rather than when a delivery arrives.
For everything else, the warranty or support contract decides the timing. Contracts range from next-business-day parts to replacement within a few hours, and the difference matters when choosing what to keep on the shelf.
Failed parts under warranty usually go back to the vendor through a return merchandise authorisation (RMA). The vendor will want the server's serial number and often a diagnostic log collected from the management controller.
A failed drive is also a data problem. Even a dead drive can hold readable data, so organisations with sensitive data often pay for a keep-your-drive option, retaining failed drives and destroying them securely instead of returning them. The decommissioning lessons later in the course cover how.
Finally, record the replacement: which part, which serial numbers, and when. That history shows up patterns, such as several drives from one batch failing in quick succession.
Practise what you just read
1. In a degraded RAID 5 array, a technician must replace the failed drive. What must be done first?
Select one
Show answer
D. In a degraded array, removing a healthy drive by mistake takes the array offline. Identifying the exact slot using the controller and locate LED before touching anything prevents that.
2. Which components are commonly hot-swappable in a rack server?
Select one
Show answer
C. Servers are designed so drives in hot-swap bays, redundant power supplies and fans can be replaced while running. Processors, memory and the motherboard normally require the server to be powered off.
3. Why should a technician wear an antistatic wrist strap when handling server components?
Select one
Show answer
A. Static charge on a person can discharge into sensitive electronics and damage them without any visible sign. A strap connected to the chassis keeps the technician and the equipment at the same potential.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.