Storage failures and degraded arrays

Listen to this lesson

Episode 40 · 47:59

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 4.1 · Troubleshooting · 28% of the exam

Why this matters

Disks are the component most likely to fail in a server. They are also the one whose failure matters most, because they hold the data. RAID exists so that a single disk failure is an inconvenience, not a disaster, but only if someone notices the failure and acts correctly. An array that loses one disk and runs unnoticed for months is one more failure away from losing everything, and a wrong move during a replacement can finish the job.

This lesson covers the warning signs before a disk fails, the difference between a degraded array and a failed one, replacing a disk and surviving the rebuild, the controller problems that look like disk problems, and dealing with corruption in the file system itself.

The lesson

Reading SMART data and the controller's log

Disks often warn before they fail.

SMART (Self-Monitoring, Analysis and Reporting Technology) is built into hard drives and SSDs, which track their own health. The attributes that matter most:

  • for hard drives, reallocated sectors, bad sectors the drive has replaced with spares, and pending sectors, sectors waiting to be reallocated. A count that keeps rising means the drive is deteriorating;
  • for SSDs, wear indicators such as percentage used or media wearout, since flash has a limited number of write cycles, and the count of spare blocks left;
  • for both, uncorrectable errors, and the drive's overall health status.

On Linux, smartctl, from the smartmontools package, reads SMART data; on Windows, management tools and vendor utilities do. Behind a hardware RAID controller, SMART data is usually read through the controller's own tools, since the operating system sees the array rather than the individual disks.

The RAID controller's log records disk errors, timeouts, disks marked as predictive failures, rebuilds and battery problems. Along with the management controller's hardware log, it is where a failing disk usually shows up first. Many controllers mark a disk as predictive failure from its SMART data, so it can be replaced before it actually fails. Both logs should feed the monitoring system, so warnings arrive as alerts rather than being discovered by chance.

A degraded array against a failed one

A RAID array has three broad states.

  • Optimal or healthy: every disk is working, and full redundancy is in place.
  • Degraded: one or more disks have failed, but the array still works, because the remaining disks hold enough data, or mirror copies, or parity, to reconstruct everything. Data is still available, and users may notice nothing, though performance can drop, especially for parity arrays, which must calculate missing data on the fly.
  • Failed or offline: more disks have failed than the array can tolerate, and the data on it is unavailable.

How many failures an array can survive depends on its level, as the RAID lesson described: RAID 1 and RAID 5 survive one, RAID 6 two, and RAID 10 at least one, more if the failures are in different mirrors. RAID 0 survives none.

The key point is that a degraded array has no margin left, or less margin than it was designed with. It must be treated as urgent: replace the disk promptly, and make sure backups are current, because the next failure may be the last.

A failed array is a recovery situation. Do not start swapping or reinitialising disks in the hope of bringing it back, since the wrong action can destroy data that a specialist or the vendor could still have recovered. Check what the controller reports, contact vendor support, and prepare to restore from backup.

Replacing a drive and watching the rebuild

Replacing a failed disk is routine, and it is also where mistakes happen.

  • Identify the right disk. Use the controller's tools to find the failed disk's slot, and turn on its locate LED, since removing a healthy disk from a degraded array can take the array offline. Check the slot number against the server's labels, not assumptions about the order.
  • Use a compatible replacement, of the same type, and at least the same capacity, ideally the same model or one the vendor supports. A disk even a few sectors smaller cannot replace the failed one.
  • Hot-swap it if the server supports it, as the hot-swap lesson described, without powering down.
  • Rebuild. The controller usually starts rebuilding automatically when it sees the new disk, or uses a hot spare that was already installed. The rebuild reconstructs the missing data onto the new disk.

The rebuild itself is the most dangerous period. It reads every sector of every remaining disk, which puts them under heavy load, and on large disks it can take many hours or days. If another disk fails, or an unreadable sector turns up on a remaining disk during a RAID 5 rebuild, the rebuild can fail and the array with it. This is the main reason RAID 6 is preferred for large disks.

During a rebuild, monitor progress, keep backups current, reduce unnecessary load, and do not touch other disks.

Controller and cache battery failures

Not every storage problem is a disk. The RAID controller itself can fail, and it has a component that fails more predictably: its cache battery or flash-backed cache module.

The controller's write cache stores data in memory and confirms writes before they reach the disks, which makes writing much faster. The battery, or a capacitor and flash module in newer designs, protects that cache: if power fails, the cached data is preserved until it can be written to disk.

When the battery fails, is charging, or reaches the end of its life, the controller protects data by switching from write-back to write-through mode, writing directly to disk. The symptom is a sudden, sharp drop in write performance, often with no other obvious sign, which is easily mistaken for a disk or application problem. The controller's log and management tools report the battery state. Batteries have a limited life, so replacement is routine maintenance.

A failed controller shows as every disk behind it disappearing at once. Replace it with the same model and firmware where possible. Most controllers store the array configuration on the disks themselves, so a replacement can import the existing arrays. Do not create a new array over the disks, which would destroy the data on them.

Corruption, and checking a file system

Sometimes the hardware is healthy but the data structures on it are damaged. A file system can become corrupted by an unexpected power loss, a crash, a failing disk, or a failing controller cache. Symptoms include files that cannot be opened, missing directories, errors in the system log, or a volume that will not mount.

Each platform has a tool to check and repair the file system:

  • On Windows, chkdsk checks a volume; chkdsk /f fixes errors, and /r also locates bad sectors and recovers readable data from them. The system volume can only be fully checked at the next restart. For NTFS, Windows can also scan online and schedule only the repair itself.
  • On Linux, fsck checks a file system, calling the right tool for its type, such as e2fsck for ext4. It must normally be run on an unmounted file system, since checking a mounted one can cause more damage. XFS uses its own tool, xfs_repair.

Before repairing, back up what you can, or better, take an image of the damaged volume, since a repair can discard data it cannot make sense of. Check the underlying hardware too: a file system that keeps becoming corrupted is usually a symptom of a failing disk, controller or memory, and repairing it repeatedly treats the symptom, not the cause.

Try it

An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.

Practise what you just read

1. A RAID 5 array reports as degraded, and users notice nothing. How urgent is it?

Select one

  1. Only urgent if performance drops
  2. Not urgent: it resolves itself
  3. Urgent: it has no redundancy left
  4. Not urgent: the users are unaffected
Show answer

C. A degraded RAID 5 array has lost its only margin, and the next disk failure loses the data. Replacing the disk promptly and confirming backups are current is the response.

2. During a RAID 5 rebuild on large disks, what is the main risk?

Select one

  1. The controller forgetting the array's RAID level partway through
  2. The rebuild overwriting data on the remaining disks with parity
  3. The array growing larger than planned by the end of the rebuild
  4. A second failure or unreadable sector during the long rebuild
Show answer

D. Rebuilding reads every sector of every remaining disk for hours or days. Another failure or read error in that window can fail the rebuild and the array, which is why RAID 6 is preferred for large disks.

3. Write performance on a hardware RAID array suddenly drops sharply, with no failed disks. What should be checked?

Select one

  1. The controller's cache battery or flash-backed cache status
  2. The operating system's page file, the main cause of slow writes
  3. The network adapter's negotiated link speed and driver version
  4. The server's system clock and its time synchronisation source
Show answer

A. When the cache battery fails or recharges, controllers switch from write-back to write-through to protect data, and writes become much slower. The controller's logs report the battery state.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.