Backups, recovery testing, continuity and recovery metrics

This course teaches SY0-801, the Security+ exam that launches on or around 17 November 2026. If you are booked on SY0-701, which can be taken until 11 June 2027, use our SY0-701 course instead.

Objective 3.4 · Security Architecture · 19% of the exam

Objective 3.4 asks you to explain why resilience and recovery belong in security architecture. The previous lesson kept services running through a failure. This one takes what happens when they stop anyway: backups that survive an attacker, the tests that prove recovery works, the plans that keep the business going, and the four metrics that set the targets. It is also the Domain 3 capstone, because a real restore pulls in classification, segmentation, encryption and site design at once.

Why this matters

Backups decide whether a ransomware incident is an expensive week or the end of the business. They are also the control most likely to be assumed rather than verified: organisations discover their backups do not work at the exact moment they need them.

The exam reflects this. It asks what makes a backup survivable, how recovery is proven, how disaster recovery differs from business continuity, and how to turn a sentence from the business into an RTO or RPO.

The lesson

Retention, immutability and scope: the backup ransomware cannot touch

Ransomware operators look for backups before they encrypt anything, because deleting them is what makes the victim pay. Three properties decide whether yours survive.

Retention is how long backups are kept and how many restore points exist. It must be long enough to reach back before the attacker arrived -- an intruder present for months has contaminated every recent restore point -- and long enough to satisfy legal and regulatory requirements. Rotation schemes such as grandfather-father-son keep daily, weekly and monthly copies so that older restore points exist without keeping every day forever.

Immutability means a backup cannot be changed or deleted, by anyone, until its retention period ends -- not by an administrator, and not by an attacker holding an administrator's credentials. It is provided by write-once storage, object-lock features in cloud storage, or media that is physically offline. Immutable copies should also sit behind separate credentials, so the account that runs production cannot reach them.

Scope is what the backups include, and gaps here are common:

  • the identity system -- without the directory, nothing else authenticates;
  • configuration of network devices, firewalls and cloud services, and the infrastructure-as-code definitions;
  • encryption keys, stored independently of the systems they protect;
  • SaaS data -- under shared responsibility, the provider keeps the service running, and protecting your data in it is often your job;
  • endpoints holding data that exists nowhere else.

The 3-2-1 guideline still frames it: three copies, on two different media, one offsite -- with at least one copy immutable or offline. Snapshots and replication help with fast rollback and hardware failure, and they are not backups on their own: replication faithfully copies deletion and encryption, and snapshots usually live on the storage they protect.

The restore that was never tested, which is the finding this lesson exists for

An unverified backup is not a control. It is a belief.

The failure modes are specific, and each has ended a real recovery:

  • the job reported success for months while silently skipping a locked database;
  • the media was fine and the decryption key was stored only in the system being recovered;
  • restoring the full dataset took eleven days against an RTO of one, because nobody had measured the restore rate;
  • every restore point contained the attacker's access, because retention was shorter than their dwell time;
  • the restore worked and nobody knew the order to bring services back, so the application came up before the directory it authenticates against.

Restoration testing checks that data restores, that it is complete and correct rather than merely present, how long the restore takes end to end, that keys are available independently, and that the recovered service actually works. The exam's phrasing is usually "what should have been done?", and the answer is a regular, documented restore test -- not a backup success report, which is the system marking its own homework.

Failover, simulation and parallel processing as ways to prove recovery

Three ways to prove that recovery works, in rising order of realism and risk:

  • Simulation -- a realistic exercise in which a scenario is played out, some systems are exercised and events are introduced as it runs. It tests people and process under pressure without moving production.
  • Parallel processing -- the recovery environment is brought up and runs alongside production, processing the same work, and the results are compared. It proves the recovery environment can do the job, at real cost, without risking production.
  • Failover -- production is actually moved to the recovery site or secondary system. It is the only test that proves the whole thing works, and the only one that can cause an outage of its own.

A mature programme uses all three at different intervals: simulations often, failover at least annually for the systems that matter most. The reason organisations avoid failover tests is the reason they need them. Discussion-based tabletop exercises belong to incident response and are covered in lesson 39.

Disaster recovery versus business continuity, and what each plan covers

Two terms the exam treats as distinct.

  • Disaster recovery (DR) is about technology: restoring systems, data and infrastructure after a disruptive event, in the right order, to the agreed targets. A DR plan names the systems, their priority, the recovery procedures, the alternate site and who carries out each step.
  • Business continuity (BC) is about the business continuing to operate, by any means, during the disruption. It covers people, premises, suppliers, communications and manual workarounds. DR is one component of it.

The distinction to carry: DR restores the systems; continuity keeps the business running while they are down. Paper processes, staff relocation, delegated authority or a substitute supplier in a scenario are continuity, not disaster recovery.

A continuity plan contains what DR plans usually lack: essential functions ranked by how long the business survives without each; succession of authority for when key people are unreachable; a communication plan that works when email and phones are down; and dependency mapping, including suppliers, because your continuity is limited by theirs.

RTO, RPO, MTTR and MTBF, and deriving them from a business statement

  • RPO (recovery point objective) -- the maximum acceptable data loss, measured in time. It drives backup frequency and replication. "We can afford to lose at most one hour of orders" is an RPO of one hour.
  • RTO (recovery time objective) -- the maximum acceptable downtime. It drives site choice and restore capability. "We must be trading within four hours" is an RTO of four hours, which rules out a cold site.
  • MTTR (mean time to repair) -- how long repair actually takes, on average. It is a measurement; RTO is a target. If MTTR exceeds RTO, the objective is not being met.
  • MTBF (mean time between failures) -- the average interval between failures of a component, used for reliability and replacement planning.

RPO looks backwards from the incident -- how much data is gone. RTO looks forwards -- how long until you are working. Both are set by the business through a business impact analysis, covered in lesson 43, and then costed: near-zero RPO and minutes of RTO are achievable and expensive, and stating them lets the business decide what to buy.

MTBF and MTTR together give availability: MTBF divided by (MTBF plus MTTR). A component that fails every 1,000 hours and takes 10 hours to repair is available about 99% of the time; halving the repair time to 5 hours raises that to about 99.5% without making the component any more reliable.

Two consistency checks the exam likes: backups must run at least as often as the RPO, and restore capability must be faster than the RTO. A fifteen-minute RPO with nightly backups is an unmet objective, whatever the plan says.

What to take into the exam

  • Retention must reach back before the attacker arrived; immutable copies sit behind separate credentials; scope includes identity, configuration, keys and SaaS data.
  • Snapshots and replication are not backups on their own.
  • The answer to "what went wrong?" is often an untested restore.
  • Simulation tests people, parallel processing proves the recovery environment, failover proves everything and risks an outage.
  • DR restores systems; business continuity keeps the business running meanwhile.
  • RPO is data loss and drives backup frequency; RTO is downtime and drives site choice; MTTR is measured repair time; MTBF is reliability.

Practise what you just read

1. Backup jobs have reported success every night for a year, yet the restore fails. What should have been done?

Select one

  1. A longer retention period for logs
  2. A second backup job every night
  3. A more detailed backup success report
  4. A regular, documented restore test
Show answer

D. An unverified backup is a belief, not a control, and a success report is the system marking its own homework. Restoration testing checks that data is complete and correct, measures the restore time, confirms keys are available independently and proves the recovered service works.

2. Why is a storage snapshot not a backup on its own?

Select one

  1. It cannot be scheduled to run automatically
  2. It usually lives on the storage it protects
  3. It records only the blocks that have changed
  4. It needs the source offline for every restore
Show answer

B. Snapshots give fast rollback but share the fate of their storage, so losing or encrypting that storage takes them too. Replication has the same weakness in another form: it faithfully copies deletion and encryption, so it protects against hardware failure rather than ransomware.

3. What does the 3-2-1 backup guideline say, and what does the lesson add to it?

Select one

  1. 3 sites, 2 providers, 1 region; plus a second region too
  2. 3 tests a year, 2 full backups, 1 differential; plus audits
  3. 3 copies, 2 media, 1 offsite; plus 1 immutable or offline
  4. 3 retention tiers, 2 keys, 1 escrow; plus a key custodian
Show answer

C. Ransomware operators hunt for backups before encrypting, because deleting them is what makes the victim pay. A copy that production credentials can delete is one the attacker can delete, so at least one copy must be immutable or offline, behind separate credentials.

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.