Backup types: full, incremental, differential and snapshots
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
Almost every lesson in this course has ended at the same safety net: if it goes wrong, restore from backup. RAID does not protect against deletion, snapshots depend on the disk they sit on, ransomware encrypts whatever it can reach, and a failed patch sometimes leaves nothing else. Backups are the last line of defence, and CompTIA places them, with disaster recovery, in the troubleshooting domain, because recovery is where troubleshooting ends when nothing else works.
A backup strategy is a set of trade-offs between how long backups take, how much storage they use, and how long and how complicated a restore will be. This lesson covers the backup types and what each restore needs, synthetic fulls, snapshots compared with backups, fitting backups into the time available, and how backup software knows what has changed.
The lesson
Full, incremental and differential backups, and what each restore needs
There are three basic backup types.
A full backup copies all the selected data, every time. It is the simplest to restore, since everything is in one backup set, but it takes the longest and uses the most storage.
An incremental backup copies only what has changed since the last backup of any kind, full or incremental. Each incremental is small and quick. But a restore needs the last full backup plus every incremental since, applied in order. If the full backup ran on Sunday and incrementals ran on each day after it, restoring on Friday needs Sunday's full and Monday's, Tuesday's, Wednesday's and Thursday's incrementals. If any one of them is missing or damaged, everything after it is lost.
A differential backup copies everything that has changed since the last full backup. Each differential grows through the week, because it includes all changes since the full, so it takes longer and uses more storage than an incremental. But a restore needs only two sets: the last full and the latest differential.
The trade-off in summary:
| Type | Backup time and size | Restore needs |
|---|---|---|
| Full | Largest, slowest | The full backup only |
| Incremental | Smallest, fastest | Last full plus every incremental since |
| Differential | Grows until the next full | Last full plus the latest differential |
Typical schemes combine them: a weekly full with daily incrementals saves time and space; a weekly full with daily differentials makes restores simpler. Exam questions often give a schedule and ask which backup sets a restore on a particular day will need.
Synthetic full backups
Full backups are the easiest to restore from, but taking them regularly means reading every file on the server and sending it all across the network, which takes time and load that production systems may not be able to spare.
A synthetic full backup avoids that. Rather than reading from the server again, the backup system builds a new full backup from backups it already holds: it combines the last full with the incrementals taken since, on the backup server or storage, producing a new full backup that is identical to one taken directly.
This gives the advantages of both approaches:
- the production server only ever sends incremental changes, so its backup window stays short;
- restores still start from a recent full, so they need few backup sets;
- the long chain of incrementals is kept short, reducing the risk that one damaged incremental breaks a restore.
The work moves to the backup infrastructure, which needs the storage performance to do it. Many modern backup systems go further with an incremental forever approach: after the first full, only incrementals are taken from the server, and synthetic fulls or equivalent processing keep restores simple.
Snapshots against backups
The virtualisation lesson warned that a snapshot is not a backup, and it is worth making the difference exact.
A snapshot captures the state of a volume, a virtual machine or a storage array at a moment. It is almost instant and is ideal for rolling back a change quickly. But snapshots are usually stored on the same storage as the original, and many depend on the original data, since they record only what has changed since. If the storage fails, is corrupted or is encrypted by ransomware, the snapshots go with it. Kept too long, they also degrade performance.
A backup is an independent copy stored separately: on a different system, different media, ideally in a different place. It survives the loss of the original storage entirely.
The two work together. Backup systems commonly use snapshots: they take a snapshot so they can copy a consistent, frozen image of a volume or virtual machine while it keeps running, and then copy that image to separate backup storage. On Windows, the Volume Shadow Copy Service (VSS) does this, and coordinates with applications such as databases so that the snapshot is application-consistent, with no transactions half-written. A backup that is only crash-consistent, like pulling the power at that moment, may need repair before a database can use it.
Storage array snapshots, replicated to a second array at another site, can serve as backups, because they are then independent of the original storage. The test is always the same: would this copy survive the loss of the original system?
Backup windows
A backup window is the period during which backups are allowed to run, usually overnight or at weekends, when the load they place on servers, storage and the network causes least disruption.
The difficulty is that data keeps growing while the window stays the same length, and eventually backups no longer fit. Symptoms are backups still running when the working day starts, slowing production systems, or jobs failing because they ran out of time.
Ways to fit backups into the window:
- incremental and synthetic full backups instead of frequent full backups;
- deduplication and compression, which reduce how much data is sent and stored, sometimes at the source before it crosses the network;
- snapshot-based backups, which freeze the data briefly and copy from the snapshot, so the production system is only affected for moments;
- staggering jobs, so not every server backs up at once;
- a dedicated backup network, so backup traffic does not compete with users;
- faster backup targets, such as disk rather than tape for the first copy, which can be copied to tape later, outside the window, a scheme called disk-to-disk-to-tape.
Some systems, such as busy databases, run around the clock and have no quiet time at all. For those, continuous protection methods, such as frequent log backups or replication, protect data throughout the day.
The archive bit and change tracking
Incremental and differential backups need to know which files have changed.
On Windows, the traditional mechanism is the archive bit, a file attribute that is set whenever a file is created or modified. Backup software uses it this way:
- a full backup copies every file and clears the archive bit on each;
- an incremental backup copies files whose archive bit is set, and clears it, so the next incremental starts afresh;
- a differential backup copies files whose archive bit is set, but does not clear it, which is why each differential includes everything changed since the last full.
That difference, clearing or not clearing the bit, is the whole mechanical difference between incremental and differential backups, and is a favourite exam question. A copy backup, by contrast, copies everything and leaves the archive bit alone, so it does not disturb the regular schedule.
Modern backup systems mostly use more efficient methods. Many compare modification timestamps or keep their own catalogue of what they backed up. Hypervisors provide changed block tracking (CBT), which records exactly which disk blocks have changed since the last backup, so backup software reads only those blocks, not whole files. Windows file systems provide a change journal recording file changes. These methods are faster, and they avoid a weakness of the archive bit: any program, or a second backup product, can change the bit and silently cause files to be skipped.
Try it
An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.
Practise what you just read
1. A full backup runs on Sunday and differentials run Monday to Saturday. The server fails on Thursday afternoon. Which sets does the restore need?
Select one
Show answer
A. Each differential contains everything changed since the last full backup, so only the latest differential is needed with the full. Applying all of them would give the same result more slowly.
2. A full backup runs on Sunday and incrementals run Monday to Saturday. The server fails on Thursday afternoon. Which sets does the restore need?
Select one
Show answer
B. Each incremental contains only the changes since the previous backup of any kind, so the restore needs the full and every incremental since, applied in order. Losing one breaks the chain.
3. Which statement about the archive bit on Windows is correct?
Select one
Show answer
C. Incrementals clear the bit so the next one copies only newer changes; differentials leave it set, so each captures everything changed since the last full, which clears it.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.