Failover sites, and validating disaster recovery
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
Clusters survive a failed server; RAID survives a failed disk; redundant power survives a failed supply. None of them survives the loss of the building. Fire, flood, a regional power cut or a cyber attack that takes down an entire data centre needs a different answer: somewhere else to run.
This is the capstone of the course, because disaster recovery draws on almost everything before it. Backups and replication supply the data, networking and DNS redirect users, documentation tells people what to do, and testing proves any of it works. This lesson covers the kinds of recovery site, replicating data and failing over to them, the harder job of failing back, proving that failover works, and the business continuity plan that disaster recovery sits inside.
The lesson
Hot, warm and cold sites
A recovery site, or disaster recovery site, is a second location where an organisation can run its systems if the primary site is lost. They are classed by how ready they are.
- A hot site is a fully equipped, running duplicate: servers, storage and networking in place, with data replicated continuously or near-continuously. It can take over in minutes to hours. It is by far the most expensive, since it effectively doubles the infrastructure.
- A warm site has the facilities, power, network connections and some or all of the hardware, but systems are not running current data. Recovery means restoring recent backups or bringing replicated systems up to date and starting them, taking hours to days. It is a middle ground in cost and speed.
- A cold site provides space, power, cooling and connectivity, but little or no equipment. Hardware must be obtained, installed and configured, and data restored from backup, taking days to weeks. It is cheapest, and suitable only for systems whose RTO allows that long.
The choice follows the RTO and RPO from the previous lesson. A hot site is justified for systems whose loss would stop the business within hours; a cold or warm site may be enough for others. Organisations often mix them, recovering critical systems at a hot site and the rest more slowly.
Options also include using a cloud provider as the recovery site, where infrastructure is only paid for in full when it is used, sometimes called disaster recovery as a service (DRaaS), and reciprocal agreements with another organisation to host each other in a disaster. Whatever the site, it must be far enough from the primary that one regional event cannot affect both.
Replication, and failing over to a recovery site
For a warm or hot site, data must reach it continuously, not only through nightly backups. That is replication: copying changes from the primary site to the recovery site as they happen.
Replication comes in two forms, with a trade-off between data loss and performance:
- Synchronous replication writes each change to both sites before confirming it to the application. No confirmed data is lost if the primary fails, giving an RPO of zero. But every write waits for the round trip to the recovery site, so it needs low latency and is practical only over short distances.
- Asynchronous replication confirms writes locally and sends them to the recovery site shortly afterwards. It works over any distance with no performance cost, but the most recent changes, typically seconds to minutes, can be lost if the primary fails suddenly.
Replication can be done by storage arrays, by hypervisors replicating whole virtual machines, or by applications themselves, such as database replication.
Replication is not a backup. It copies everything, including mistakes: a deleted table or files encrypted by ransomware are faithfully replicated to the recovery site. Backups, with their history of restore points, are still needed alongside.
Failover to the recovery site follows the disaster recovery plan: formally declare a disaster, since failover is itself disruptive and must be decided, not drifted into; bring up systems at the recovery site in dependency order, as the restore procedure sets out; redirect users, usually by updating DNS records, where the low TTLs planned in the DNS lesson pay off, or through global load balancing; then verify that services work, and tell users and stakeholders.
Failing back
Once the primary site is repaired, operations must move back, and failback is often harder than failover.
During the time at the recovery site, the recovery systems have been running production and accumulating new data. Failing back means:
- repairing or rebuilding the primary site and confirming it is healthy;
- replicating changes from the recovery site back to the primary, reversing the normal direction, until it is up to date;
- a planned cutover during a maintenance window: stopping changes at the recovery site, completing a final synchronisation, switching services and DNS back, and verifying;
- restoring normal replication from primary to recovery site, so protection is in place again.
Unlike failover, failback can be scheduled, and it should be: rushed failbacks cause second outages and data loss, as the clustering lesson noted for automatic failback within a cluster. Some organisations, when the recovery site is fully capable, simply make it the new primary and treat the repaired site as the recovery site.
The failback procedure deserves the same documentation and testing as failover. It is often neglected, and organisations have found themselves running at a recovery site for months because nobody knew how to get back.
Proving that failover works, rather than assuming it
A disaster recovery plan that has never been exercised is an assumption. The clustering lesson made the same point about a single cluster; at site level, the stakes and the number of things that can go wrong are far larger.
Common discoveries during disaster recovery tests:
- a system that was never added to replication, or whose replication stopped months ago without alerting anyone;
- a dependency nobody documented, such as an application needing a licence server or authentication service that exists only at the primary site;
- network paths, firewall rules or DNS changes that do not work as expected;
- recovery site capacity too small for the production load, because the primary grew and the recovery site did not;
- keys, credentials or documentation that are only available at the primary site;
- recovery taking far longer than the RTO.
So failover is tested using the range of exercises from the previous lesson, tabletop, parallel and, for the most critical systems, full failover, on a regular schedule and after significant changes. Each test is measured against the RTO and RPO, its findings are fixed, and the plan is updated. Monitoring replication continuously, with alerts when it lags or stops, catches the most common silent failure between tests.
The business continuity plan around it
Disaster recovery is concerned with restoring IT systems. It sits within a larger business continuity plan (BCP), which is concerned with keeping the business operating through a disruption, whatever its cause.
A business continuity plan covers much that is not technical:
- which business functions are critical, identified through the business impact analysis, and in what order they must be restored;
- how people will work: alternative offices, working from home, or manual procedures, such as paper forms, to keep essential operations going while systems are down;
- people and roles: who leads, who decides, who is contacted, and deputies for each, since key people may be unavailable in a disaster;
- communication with staff, customers, suppliers, regulators and the media;
- suppliers and dependencies outside the organisation, and what happens if one of them fails.
The disaster recovery plan is the IT component: the recovery sites, replication, backups, restore procedures and failover processes covered across this course. It supports the business continuity plan by restoring the systems that critical business functions depend on, within the objectives the business has set.
Both plans must be maintained: reviewed regularly, updated after changes and tests, stored where they can be reached in a disaster, and practised, so that when a real disaster comes, the people carrying them out have done it before.
Practise what you just read
1. A business needs its critical systems running again within an hour of losing its main data centre. Which recovery site fits?
Select one
Show answer
D. A hot site has equipment running with replicated data and can take over in minutes to hours. Warm sites take hours to days, and cold sites days to weeks.
2. What is the trade-off of synchronous replication to a recovery site?
Select one
Show answer
B. Synchronous replication confirms writes only when both sites have them, giving an RPO of zero, but adds round-trip latency to every write, limiting the distance between the sites.
3. Why does replication to a recovery site not replace backups?
Select one
Show answer
C. Replication faithfully copies every change, including mistakes and malicious encryption, within seconds or minutes. Backups keep earlier restore points to go back to when something goes wrong.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.