Recovery testing, RTO and RPO
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
Nobody actually wants backups. They want the ability to restore. A backup that has never been restored is a hope, not a capability: the job may have been silently skipping files, the media may be unreadable, the encryption key may be lost, or the restore may take four days when the business can survive one. Organisations regularly discover these problems at the worst possible moment, in the middle of a real disaster.
This lesson covers the two targets that define what recovery must achieve, testing restores routinely, the kinds of disaster recovery exercise, the written procedure that makes recovery repeatable, and measuring what recovery actually achieves against what was promised.
The lesson
Recovery time and recovery point objectives, and who sets them
Two measures define the recovery a system needs.
The recovery time objective (RTO) is the maximum acceptable time a system can be unavailable after a disruption: how long until it must be working again. An RTO of four hours means the service must be restored within four hours of failing.
The recovery point objective (RPO) is the maximum acceptable data loss, measured in time: how far back the restored data may be. An RPO of one hour means that no more than an hour of data may be lost, so the organisation must have a recoverable copy of the data at least every hour.
The two drive technical choices directly:
- RPO determines backup frequency. A nightly backup gives an RPO of up to 24 hours. An RPO of minutes needs frequent log backups, continuous data protection, or replication.
- RTO determines the recovery method. Restoring terabytes from off-site tape may take days; an RTO of minutes needs a standby system, clustering, or failover to a replicated copy.
Shorter objectives cost more, often much more, so they are set system by system. Related measures include the maximum tolerable downtime (MTD), the absolute limit beyond which the business suffers serious harm, which the RTO must be well inside, and, for hardware, MTBF (mean time between failures) and MTTR (mean time to repair).
Who sets them? Not IT alone. RTO and RPO are business decisions, made by the owners of each system and service, usually through a business impact analysis (BIA) that asks what an outage or data loss would cost the organisation over time. IT's role is to explain what each objective costs to achieve, and then to build and test systems that meet the objectives agreed.
Testing restores on a schedule
The only proof that a backup works is a successful restore. Testing restores is part of routine operations, not something done once.
A sensible testing programme:
- Restore individual files and folders regularly, such as monthly, from different servers and different ages of backup, including older backups from long-term retention and off-site copies.
- Restore whole servers and applications periodically, into an isolated test environment, and check that they actually work: the operating system starts, services run, the database opens and its data is consistent, and the application can be used. A backup that restores files but not a working application has not been tested.
- Test every backup type and medium: disk, tape, cloud and immutable copies, since each has its own ways of failing.
- Include the encryption keys and credentials that a real restore would need, obtained the way they would be in a disaster.
- Check backup job reports daily for failures, warnings and skipped files, which are early signs of problems a restore test would find later.
Many backup products can automate restore verification, for example by booting a backed-up virtual machine in an isolated sandbox and confirming it starts. Automation helps, but does not replace occasionally performing the full manual procedure.
Tabletop exercises against full disaster recovery tests
Testing a single restore proves a backup works. Testing disaster recovery proves that the organisation can recover, with its people, procedures, and dependencies. There is a range of exercises, from cheap and safe to expensive and realistic:
- A walkthrough or checklist review: participants read through the recovery plan to check it is complete and current.
- A tabletop exercise: the key people meet and talk through a realistic scenario, such as "the main data centre has flooded", step by step, discussing what each would do. It costs little and disrupts nothing, and is very good at finding gaps in the plan: missing contacts, unclear decisions, dependencies nobody wrote down. It does not prove that anything technical actually works.
- A simulation or parallel test: systems are actually recovered at the recovery site or in a test environment, while production keeps running. It proves the technical recovery without risking the live service.
- A full interruption test, or full failover: production is actually shut down or moved to the recovery site. It is the only test that proves everything, and the riskiest, since a failure means a real outage. It is done rarely, carefully planned, and only once the lesser tests pass.
Most organisations use a combination: regular tabletop exercises, periodic parallel tests, and occasional full tests for the most critical systems, as the next lesson on failover sites discusses.
Documenting the restore procedure
In a real disaster, the person who set up the backups may not be available, and the people who are available will be under great pressure. Recovery must not depend on memory.
A restore procedure, or recovery runbook, as described in the documentation lesson, sets out step by step how to recover each important system:
- where the backups are, and how to get the media or access the cloud copies;
- where the encryption keys and necessary credentials are, and how to obtain them;
- the order of recovery, since systems depend on each other: directory services and DNS before the applications that use them, databases before application servers;
- the exact steps for each system, including any configuration needed after the restore;
- how to verify that the restored system works;
- contacts: the system owner, vendors and support contracts.
Keep the procedure where it can be reached during an outage, as the documentation lesson warned: not only on the servers it describes. And update it after every test, since tests reliably reveal steps that are missing, wrong or out of date.
Measuring the actual recovery against the target
Every restore test and disaster recovery exercise should be timed and measured, because objectives only mean something if the organisation knows whether it can meet them.
Measure:
- the actual recovery time: from the start of the incident, or the test, to the service working and verified, compared with the RTO. Include every stage, such as finding people, retrieving media and getting keys, not just the restore itself;
- the actual recovery point: how recent the restored data was, compared with the RPO;
- whether the restored system worked fully, and what had to be fixed.
If actual recovery falls short of the targets, the gap must be dealt with: by improving the recovery method, such as faster backup storage or a standby system, by improving the procedure, or by going back to the business owners to agree new objectives, or the budget to meet the existing ones. What must not happen is for the organisation to go on believing it can recover in four hours when every test says it takes twelve.
Record the results, share them with system owners, and track them over time. Recovery capability quietly degrades as data grows and systems change, and measurement is the only way to notice.
Practise what you just read
1. Backups run every night at 01:00. The server fails at 16:00. What is the most data that could be lost?
Select one
Show answer
C. Everything written since the 01:00 backup, about 15 hours, is lost. The worst case is just under 24 hours, just before the next backup, which is why RPO sets backup frequency.
2. What does the recovery time objective (RTO) define?
Select one
Show answer
B. RTO is how quickly a system must be restored after a disruption. RPO is how much data, measured in time, may be lost, and together they determine recovery methods and backup frequency.
3. Who should set a system's RTO and RPO?
Select one
Show answer
A. RTO and RPO reflect what downtime and data loss would cost the business. IT explains what each target costs to achieve, and then builds and tests systems that meet the agreed objectives.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.