Clustering and failover

Listen to this lesson

Episode 18 · 67:18

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 2.4 · Server administration · 30% of the exam

Why this matters

Redundant power supplies, RAID and teamed network adapters protect a server from losing one of its parts. None of them helps when the server itself fails: a motherboard dies, the operating system crashes, or it has to be restarted for patches. To keep a service running through that, it needs more than one server, and a way to move the service between them.

That is what a cluster does. This lesson covers how clusters are arranged, how they decide which servers are in charge, the failure that clustering is most careful to prevent, and why a cluster that has never failed over should not be trusted to do so.

The lesson

Active-active against active-passive clusters

A failover cluster is a group of servers, called nodes, that work together to keep services available. When a node fails, the services it was running are started on another node. That move is failover.

Clusters are arranged in one of two ways.

In an active-passive cluster, one node runs the service while another stands by, doing nothing until it is needed. It is simple, and after a failover the standby node offers exactly the capacity the active one did. The cost is a server that is idle most of the time.

In an active-active cluster, every node runs workloads at the same time. No hardware sits idle, and the load is shared. The trap is sizing: when a node fails, its workload moves onto the survivors, which must have enough spare capacity to absorb it. Two nodes each running at 70 per cent cannot absorb each other; after a failover the survivor would need 140 per cent. Active-active clusters are therefore sized so that the remaining nodes can carry the full load if one fails, a rule often written as N+1.

Quorum, and why a cluster needs a witness

A cluster must always know which nodes are allowed to run its services. It decides by quorum: each node has a vote, and the cluster keeps running only while a majority of the votes can communicate with each other. Nodes that find themselves in a minority stop running clustered services.

With an even number of nodes, a clean split leaves no majority. The answer is a witness, an extra vote that is not a node:

  • a disk witness, a small shared disk that the cluster uses as a tiebreaker;
  • a file share witness, a file share on a separate server;
  • a cloud witness, stored in a cloud storage service, useful when nodes are in different sites.

A four-node cluster with a witness has five votes, so three are a majority, and it survives a split into two and two. A two-node cluster is the case that matters most: without a witness, losing either node or the link between them leaves one vote out of two, which is not a majority. Two-node clusters therefore need a witness to survive the loss of either node.

Heartbeat networks and split-brain

Nodes check on each other constantly by exchanging heartbeat messages over the network. If a node stops answering for long enough, the others declare it failed and fail its services over. Clusters commonly use a dedicated network for heartbeats, or at least more than one path, so a single switch or cable fault does not look like a node failure.

The failure clustering is designed to prevent is split-brain. If the network between nodes fails but the nodes themselves keep running, each side can conclude that the other has died. Without protection, both sides would take over the same services and write to the same shared storage at the same time, corrupting the data.

Two mechanisms prevent it:

  • Quorum ensures that only the side holding a majority keeps running. The minority side knows it has lost and stops.
  • Fencing makes sure the losing node really has stopped, by cutting it off from shared storage or forcibly powering it off through its management controller or a switched PDU. This is sometimes called STONITH, "shoot the other node in the head", and it exists because a node that has merely been asked to stop cannot be trusted to have done so.

Planned and unplanned failover, and failback

An unplanned failover happens automatically when a node fails. The cluster detects the failure, and the service restarts on another node. For most services this means a short interruption: connections drop, clients reconnect, and any work in progress at the moment of failure may have to be repeated. The cluster reduces downtime from hours to seconds or minutes; it does not make failure invisible.

A planned failover is started by an administrator, usually to free a node for maintenance. Its services are moved gracefully before the node is taken down, and virtual machines can often be moved by live migration with no interruption at all. Planned failovers are how clustered servers are patched one node at a time without taking the service down, the rolling update from the firmware lesson.

Failback is returning services to their original node once it is repaired. It can be automatic, but automatic failback causes a second disruption, possibly at a busy time, and a node that keeps failing can cause services to bounce back and forth. Most organisations prefer to fail back manually, at a time they choose, once they are confident the node is healthy.

Testing a failover before you need one

A cluster that has never failed over is a cluster nobody knows will work. The problems that stop a failover are ordinary configuration mistakes: a node that cannot reach the shared storage, an application component installed on one node but not another, differing patch levels, or a licence tied to one machine. Each is invisible until the moment the cluster needs it.

So clusters are tested:

  • Before production, using the platform's validation tools, such as the cluster validation tests in Windows Server, which check storage, networking and configuration across every node.
  • On a schedule, in maintenance windows, by moving services deliberately and, on test clusters, by simulating real failures such as disconnecting a node's network or power.
  • Completely: confirm that clients reconnect, that the surviving node copes with the load, and that monitoring raised the alerts it should have.

Record the results and fix what the test finds. A failover that fails in a test at a planned time is an inconvenience; the same failure during a real outage is a disaster.

Practise what you just read

1. A two-node cluster has no witness. The link between the nodes fails, but both keep running. What does quorum do?

Select one

  1. The node with the lower host name keeps running
  2. Neither node has a majority, so the cluster stops
  3. Both nodes continue, each counting the other's vote
  4. The cluster adds a third voting node automatically
Show answer

B. Each node has one of two votes, and neither holds a majority, so the cluster cannot safely continue. A witness adds a third vote, letting one side keep quorum.

2. What is split-brain in a cluster?

Select one

  1. Both sides believe they are in charge and use shared resources
  2. Two separate clusters that share one witness disk or file share
  3. A cluster with more nodes than its licence allows, so half stop
  4. A single node fitted with two network adapters on separate subnets
Show answer

A. If nodes cannot communicate but keep running, each may take over services and write to the same storage, corrupting data. Quorum and fencing exist to prevent it.

3. In an active-active cluster of two nodes each running at 70 per cent, what happens when one fails?

Select one

  1. Nothing, as each workload's demand is halved
  2. The failed node's workload is discarded
  3. The survivor cannot carry the combined load
  4. Both nodes stop and the cluster goes down
Show answer

C. The surviving node would need 140 per cent of its capacity. Active-active clusters are sized so the remaining nodes can carry the full load after a failure, often described as N+1.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.