RAID levels, and what each one actually survives

Listen to this lesson

Episode 5 · 59:00

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 1.2 · Server hardware installation and management · 18% of the exam

Why this matters

Disks fail. They are among the few moving parts left in a server, and in a room with hundreds of them, one failing is a routine event rather than an emergency. RAID, a redundant array of independent disks, is how a server keeps running through that event.

The exam rarely asks you to define a RAID level. It asks which level fits a workload, how much usable space an array gives, how many failures it survives, and what goes wrong during a rebuild. This lesson is built around those questions.

The lesson

RAID 0, 1, 5, 6 and 10: usable capacity, speed and failures tolerated

RAID combines several physical disks into one logical volume using two techniques: striping, which spreads data across disks so they work in parallel, and redundancy, either by mirroring data or by storing parity, extra information from which a missing disk's data can be recalculated.

Level Method Minimum disks Usable capacity Failures survived
RAID 0 striping 2 all disks none
RAID 1 mirroring 2 one disk one (per mirror)
RAID 5 striping with single parity 3 all but one disk one
RAID 6 striping with double parity 4 all but two disks two
RAID 10 mirrored pairs, striped 4 half the disks one per mirrored pair

A worked example makes the capacity column concrete. With four 4 TB disks: RAID 0 gives 16 TB, RAID 5 gives 12 TB, RAID 6 gives 8 TB, and RAID 10 gives 8 TB. RAID 1 is normally a pair, so two 4 TB disks give 4 TB.

RAID 0 is fast and has no redundancy at all: lose any one disk and the whole array is lost. It is only for data that can be recreated, such as scratch space.

RAID 1 writes identical copies to both disks. It survives one failure, reads well, and wastes half the raw capacity.

RAID 5 stripes data with parity spread across all the disks. It survives one failure and is capacity-efficient, but every write requires the parity to be recalculated and updated. A small random write costs four disk operations, which is the write penalty that makes RAID 5 slow for write-heavy work.

RAID 6 adds a second, independent parity block, so it survives any two failures. The price is one more disk of capacity and a higher write penalty of six operations per small write.

RAID 10 (also written 1+0) mirrors disks in pairs and stripes across the pairs. It survives one failure in each pair, and can therefore survive several failures if they land in different pairs, but it cannot survive losing both disks of the same pair. It has no parity calculation, so writes are fast, and rebuilds are quick because a replacement disk is simply copied from its mirror.

Hardware against software RAID, and the controller's cache and battery

Hardware RAID uses a dedicated controller card, or one built into the server, with its own processor. The operating system sees one logical disk and is not involved in the RAID work. Controllers are managed through their own firmware utility and vendor software.

Software RAID is managed by the operating system: mdadm on Linux, Storage Spaces on Windows, or file systems such as ZFS that do it themselves. It costs nothing extra, is not tied to a particular controller model, and on modern processors the performance cost is small. Its main complications are booting from the array and that the operating system must be running to manage it.

Hardware controllers usually have a cache, a block of memory that holds data on its way to the disks. In write-back mode, the controller tells the operating system a write is complete as soon as it is in the cache, which is much faster. In write-through mode, it waits until the data is on disk.

Write-back caching is only safe if the cache survives a power failure, because otherwise data the operating system believes is written would be lost. That is the job of a battery-backed or flash-backed cache: a battery or capacitor keeps the cache alive, or copies it to flash, until power returns. When the battery fails or is still charging, a controller will typically fall back to write-through to protect data, and the server suddenly gets much slower. A controller cache battery failure is therefore both a hardware alert and a performance problem, and the troubleshooting lessons return to it.

Hot spares, rebuild times, and the second failure during a rebuild

When a member disk fails, the array keeps running in a degraded state. It has lost its redundancy, or some of it, until the failed disk is replaced and its contents are rebuilt onto the new one.

A hot spare is an installed, idle disk that the controller uses automatically: the moment a member fails, the rebuild starts onto the spare without waiting for anyone to notice. A dedicated hot spare belongs to one array; a global hot spare can replace a failure in any array on the controller.

Rebuilds take time, and the time grows with disk size. A rebuild must read the entire contents of the surviving disks, in a parity array all of them, and on large modern drives this can take many hours or even days. During that window:

  • the array has reduced or no redundancy;
  • every surviving disk is working hard, which is when marginal disks tend to fail; and
  • the disks often came from the same batch and have the same age and wear.

For RAID 5 this is the dangerous period. A second failure, or even an unrecoverable read error on one of the surviving disks, can make the rebuild fail and the array unrecoverable. This is the main reason RAID 6 has replaced RAID 5 for arrays of large drives: it can still survive a second problem during the rebuild.

Why RAID is availability and not a backup

RAID protects against one thing: a disk failing. It keeps the server running while a disk is replaced. It does nothing about the other ways data is lost, and several of them are copied to every member of the array instantly:

  • a file deleted by mistake is deleted on every disk;
  • data corrupted by an application or a bug is corrupted everywhere;
  • ransomware encrypts the array as easily as a single disk;
  • a failed controller, a fire, a flood or a theft takes the whole array.

RAID is about availability. A backup is a separate copy, kept somewhere the original's problems cannot reach, from which data can be restored. A server needs both. The backup lessons at the end of the course cover how to do that properly.

Choosing a level for a database, a file share and a boot volume

Three common cases show how the trade-offs decide the choice.

A database does many small, random writes. The parity write penalty makes RAID 5 and 6 slow for this, so databases usually get RAID 10, accepting the loss of half the raw capacity in exchange for write performance and fast rebuilds.

A file share holds a lot of data that is read far more often than it is written. Capacity matters most and write performance matters less, so a parity level fits, and with large drives that should be RAID 6 for the rebuild protection described above.

A boot volume holds the operating system. It needs to survive a disk failure but is small, so two modest disks, often SSDs, in RAID 1 are the usual choice. Keeping the operating system on its own mirrored pair also means data arrays can be rebuilt or replaced without touching it.

The pattern generalises: identify whether the workload is dominated by writes, reads or capacity, and how much a second failure during a rebuild would cost, and the level usually chooses itself.

Try it

An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.

Practise what you just read

1. An array of four 2 TB disks must survive any two disk failures. Which RAID level fits, and what capacity does it give?

Select one

  1. RAID 1, with 2 TB usable
  2. RAID 6, with 4 TB usable
  3. RAID 5, with 6 TB usable
  4. RAID 0, with all 8 TB usable
Show answer

B. RAID 6 uses two disks' worth of parity and survives any two failures; four 2 TB disks give 4 TB. RAID 5 survives only one failure, and RAID 0 survives none.

2. Which statement about RAID 0 is correct?

Select one

  1. It mirrors writes to two disks and survives losing either
  2. It needs at least three disks and uses one for parity
  3. It survives one failure if the disks are identical
  4. It improves performance but survives no disk failures
Show answer

D. RAID 0 stripes data across disks with no redundancy. It is fast and uses all the capacity, but losing any one disk loses the whole array.

3. Why is RAID not a backup?

Select one

  1. RAID needs a hardware controller, unlike backup software
  2. RAID 5 loses the array's data as soon as one member disk fails
  3. It copies deletions and corruption instantly to every member
  4. RAID stores parity rather than data, so files are unreadable
Show answer

C. RAID protects against disk failure, not against data loss in general. A deleted file, ransomware or a corrupted file system is written to every member at once, and only a separate backup can bring it back.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.