High availability, site resilience and power

This course teaches SY0-801, the Security+ exam that launches on or around 17 November 2026. If you are booked on SY0-701, which can be taken until 11 June 2027, use our SY0-701 course instead.

Objective 3.4 · Security Architecture · 19% of the exam

Objective 3.4 asks you to explain why resilience and recovery belong in security architecture. This lesson takes the half that keeps services running through a failure: redundancy, alternate sites, platform diversity, power and capacity. Backups, recovery testing, continuity plans and the recovery metrics are the next lesson.

Why this matters

Availability is a third of the CIA triad and the part security people neglect, because it feels like an operations concern. It is not: ransomware, denial of service, wipers and destructive attacks are all attacks on availability, and the controls in this lesson are what answer them.

The exam tests this as fit: a requirement and a budget, and the redundancy that matches them. Over-engineering is a wrong answer as surely as under-engineering is.

The lesson

Load balancing, clustering and autoscaling, and what each survives

High availability means designing so that a service survives the failure of any single component. Three techniques do most of the work.

Load balancing spreads incoming requests across several independent servers, any of which can answer. Health checks remove a failed node from rotation, which is what makes it a resilience control rather than only a performance one. Session persistence (sticky sessions) sends a user back to the same node, which suits applications holding session state locally and undermines even distribution. The balancer itself must be redundant, or it becomes the single point of failure it was meant to remove.

Clustering joins systems so they act as one and share state. Active-active clusters serve traffic on every node; active-passive clusters keep a standby ready to take over, with a short failover delay. Load balancing suits stateless workloads such as web tiers; clustering suits stateful ones such as databases, where a node failure must not lose the data in flight.

Autoscaling adds or removes instances automatically as demand changes. It survives load -- a traffic spike, a busy season -- rather than component failure. Its security consequences:

  • every new instance must come from a hardened, patched image, because autoscaling copies whatever it starts from;
  • instances that scale in are destroyed along with their local logs, so logs must be shipped centrally;
  • an attack can scale you up as well as down: without maximum limits, a flood of requests becomes a large bill rather than an outage.

All three survive a single node failure. None survives losing the site.

Hot, warm and cold sites, and the environmental risks of where a site stands

Site resilience is what survives losing a whole location.

  • Hot site -- fully equipped, data current, ready in minutes to hours. Most expensive.
  • Warm site -- hardware and connectivity in place, data restored from backup when needed. Hours to days.
  • Cold site -- space, power and connectivity only. Equipment must be brought in and data restored. Days to weeks. Cheapest.

The selection driver is the recovery time objective: four hours rules out a cold site; two weeks makes a hot site an expensive luxury.

Environmental factors decide whether a site survives in the first place. Ask, for each primary and alternate site:

  • is it on a flood plain or a coast exposed to storm surge;
  • is it in a zone of earthquakes, hurricanes, wildfire or extreme heat that will strain cooling;
  • does it share a power grid, water supply or telecoms route with the other site, so one regional failure takes both;
  • is it close to a hazard such as a chemical plant, an airport flight path or a rail line carrying fuel?

That is why the alternate site must be far enough away not to share the threats in your risk assessment. Distance has a cost: synchronous replication only works over limited distances, so very distant sites usually replicate asynchronously and accept some data loss.

Platform diversity across vendors, hardware and hypervisors, and multicloud as resilience

If every server runs the same software at the same patch level, one vulnerability or one bad update affects all of them at once -- a common mode failure. Faulty updates from a single vendor have taken down enormous numbers of machines simultaneously. Platform diversity limits that blast radius:

  • Vendor platform -- more than one operating system or security vendor for the most critical layers.
  • Hardware -- more than one hardware supplier or model, so a firmware flaw or supply shortage does not hit everything.
  • Virtualisation -- more than one hypervisor, so a hypervisor vulnerability or licensing change does not stop every workload.

Multicloud systems apply the same idea to providers: a critical workload that can run in two clouds survives one provider's regional outage.

The cost is real, and the exam expects you to name it: diversity multiplies hardening standards, patch cycles, skills and monitoring, and gives you the union of both platforms' vulnerabilities. Poorly run diversity is less secure than a well-run single platform. It is most defensible where failure would be total: identity, connectivity, and backups that do not share technology or credentials with production.

UPS, redundant power supplies, generators and surge protection

Power is the dependency beneath every other control.

  • A surge protector absorbs voltage spikes from lightning or switching on the grid. It protects equipment from damage; it does nothing during an outage.
  • A redundant power supply (RPS) gives a device a second power supply, or connects it to an external backup unit. Fed from separate circuits, it means a failed supply, tripped breaker or one dead feed does not take the device down.
  • An uninterruptible power supply (UPS) carries the load from the instant mains power fails, from batteries, for minutes. Its jobs are to ride out brief interruptions and to bridge the gap until a generator starts or systems shut down cleanly. Many UPS units also condition power and suppress surges.
  • A generator supplies power for hours to days, limited by fuel. It takes time to start and stabilise -- the gap the UPS covers.

They work in layers, not as alternatives: surge protection for spikes, redundant supplies for component failure, UPS for the transition, generator for the duration.

Capacity planning, and the runtime figure that makes a power plan useful

Capacity planning asks whether there is enough -- for normal load, for peak, and for the degraded state during a failure, when what remains must carry everything. A failover design whose secondary can carry only 60% of production fails exactly when it is used. Capacity covers people (enough trained staff that recovery does not depend on one person), technology (compute, storage, bandwidth, licences) and infrastructure (power, cooling, space).

Power is where capacity planning most often hides an assumption, and the figure that makes a power plan real is UPS runtime at actual load:

  • rated runtime assumes a particular load, and doubling the load cuts runtime to well under half. Racks grow; the eight minutes measured at installation may now be two;
  • batteries age, so they must be load-tested and replaced on a schedule;
  • the UPS runtime must exceed the time the generator needs to start and take the load, with margin;
  • generator fuel must last as long as the outage you plan for, with a refuelling arrangement that still works during a regional event;
  • generators must be tested under load: one that starts but cannot carry the load is the classic finding.

Capacity is a security concern because exhaustion is an availability failure, and attackers cause it deliberately: denial of service exhausts bandwidth or compute, and full log storage silently drops the events your detection needs.

What to take into the exam

  • Load balancing suits stateless work, clustering stateful; autoscaling survives load, needs hardened images and maximum limits. None survives the site.
  • Hot, warm or cold is chosen by RTO; the site must not share environmental threats with the primary.
  • Platform diversity across vendor, hardware and hypervisor prevents common mode failure and multiplies the work.
  • Surge protector for spikes, redundant supply for component failure, UPS for minutes, generator for hours to days.
  • A power plan is only real with a measured runtime at actual load and tested generators.

Practise what you just read

1. Which high-availability technique suits a stateful workload such as a database?

Select one

  1. Round-robin load balancing across web nodes
  2. Autoscaling that adds instances under load
  3. Clustering with shared or replicated state
  4. Sticky sessions on one load-balancer node only
Show answer

C. Load balancing suits stateless work where any server can answer any request. A database must survive a node failure without losing data in flight, which needs a cluster that shares or replicates state, either active-active or active-passive with a short failover delay.

2. A business states a recovery time objective of four hours. Which alternate site type certainly cannot meet it?

Select one

  1. A warm site
  2. A hot site
  3. A hot DR site
  4. A cold site
Show answer

D. A cold site provides space, power and connectivity only, so equipment must be brought in and data restored, which takes days to weeks. Site type is chosen from the RTO, and a four-hour objective needs a site with equipment and data already in place.

3. What does a UPS provide that a standby generator does not?

Select one

  1. Power for hours or days, limited only by its fuel
  2. Instant takeover, bridging minutes until backup starts
  3. Protection from surges, though never during an outage
  4. A second power supply fitted inside each server's chassis
Show answer

B. A UPS carries the load from the instant mains power fails, from batteries, for minutes. A generator runs for hours to days but takes time to start and stabilise, which is exactly the gap the UPS covers. They are layers, not alternatives.

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.