Patching failures, and rolling back safely
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
The patching lesson in the security domain built a process to make updates safe: test, stage, deploy and verify. Even so, updates fail. Sometimes they refuse to install. Sometimes they install and leave a server waiting for a reboot nobody performed. And sometimes they install perfectly and break the application the server exists to run.
A failed patch creates pressure from two directions: the service is broken now, and the vulnerability the patch fixed is still there. This lesson covers diagnosing updates that fail, recognising an update that broke an application, the three ways back, making patch testing catch these problems earlier, and keeping people informed while the fix is in progress.
The lesson
Failed updates and pending reboots
An update can fail to install for many reasons, and the error code is the starting point. Common causes:
- Insufficient disk space, especially on the system volume, where updates are downloaded and unpacked.
- A corrupted update cache or download. On Windows, resetting the update components, which includes clearing the SoftwareDistribution folder, forces a fresh download.
- Corrupted system files, which the System File Checker, sfc /scannow, and the DISM tool, with DISM /Online /Cleanup-Image /RestoreHealth, can detect and repair.
- A missing prerequisite, such as a servicing stack update that must be installed first.
- Connectivity to the update server, whether WSUS, a repository mirror or the vendor, or an update not yet approved for that server's group.
- On Linux, package dependency conflicts, a locked package database because another package manager process is running, or a repository that is unavailable or has an expired signing key.
Logs explain most failures: on Windows, the update history, the Setup and System event logs, and the Windows Update log; on Linux, the package manager's own log, such as /var/log/dnf.log or /var/log/apt/history.log.
Pending reboots are a quieter problem. An update that has installed but not yet taken effect leaves the server unprotected and sometimes unstable, with some components updated and others not, and further updates may refuse to install until the reboot happens. Management tools, scripts and compliance reports can detect pending reboots, as the patching lesson described. The fix is simply to schedule the reboot, in a window, and check that the server returns healthy.
Updates that break an application
The harder failure is an update that installs without error and breaks something.
Typical signs appear shortly after the maintenance window: an application that will not start, errors in its logs, a feature that stops working, poor performance, or a stop error on boot. The troubleshooting method's question "what changed?" leads straight to it, provided the change was recorded, with the list of updates installed and when.
Confirm the connection before acting:
- Compare the timing: did the problem start with the first restart after patching?
- Check update history for exactly what was installed on this server.
- Check whether other servers that received the same update show the same problem, and whether servers that did not receive it are fine.
- Search the vendor's release notes and known issues for the update. Vendors often publish known problems and workarounds within days, and application vendors often state which operating system updates they support.
Sometimes the update is not really at fault: it restarted the server, and the restart exposed an older problem, such as a service with the wrong startup type or an expired certificate. Check that the fault truly follows the update before removing it.
Rolling back: uninstall, snapshot or restore
When an update has broken a service, there are three ways back, from the most targeted to the broadest.
Uninstall the update. On Windows, installed updates can be removed through Settings, Control Panel, or wusa /uninstall with the update's KB number; on Linux, package managers can revert, for example with dnf history undo, or by installing the previous package version. Choosing a previous kernel from the GRUB menu reverses a kernel update. Uninstalling is precise, removing only the update, but some updates cannot be removed, and cumulative updates may take other fixes with them.
Revert to a snapshot. If a virtual machine snapshot was taken before patching, as the patching lesson recommended, reverting returns the whole server to that moment in minutes. It also discards everything that changed since, including application data written after the snapshot, which matters for databases and file servers. Delete the snapshot once the patch is confirmed good, as the virtualisation lesson warned.
Restore from backup. The broadest option, when neither of the others is possible, such as a physical server or a patch that damaged the system too badly to boot. It takes longest, and loses changes made since the backup.
Whichever route is taken:
- follow the change process, recorded as an emergency change if necessary;
- block the update from reinstalling automatically, by declining it in WSUS, hiding it, or excluding the package, until a fix is available;
- remember the vulnerability is open again, and apply mitigations, such as restricting access, until a fixed update can be installed.
Testing patches before production
Every broken production server is a question for the testing process: why did testing not catch this?
The patching lesson's process depends on the test environment resembling production. Common gaps:
- the test servers do not run the same applications, versions or configurations as production;
- testing checks that servers boot, but not that applications work;
- the pilot group contains only unimportant servers, so problems affecting the critical application are only found in production;
- testing is done, but too briefly to catch problems that appear after a day or a week, such as memory leaks.
Improvements include keeping a test environment that mirrors production, built from the same baselines; writing a short test checklist for each application, covering the functions that matter; including at least one server running each important application in the pilot; and allowing time between rings for problems to show. After a failure, the lessons learned step of the troubleshooting method feeds back into the testing process, so the same class of problem is caught next time.
Communicating an outage while it is being fixed
While a patching failure is being fixed, users and managers need to know what is happening. Silence causes duplicate support calls, guesswork and loss of trust, and pulls the people fixing the problem away from it to answer questions.
Good outage communication:
- Acknowledge quickly: tell users there is a known problem, which services are affected, and that it is being worked on, even before the cause is known.
- Use agreed channels: a status page, email, messaging, or a recorded message on the help desk line, as set out in the incident process. Remember that if email is the service that is down, it cannot carry the message.
- Give updates at set intervals, such as every thirty minutes, even when there is no news, and give an estimated time of restoration only when it is realistic.
- Say what users should do: a workaround, or simply not to report it again.
- Announce resolution, and confirm with users that the service is working.
Afterwards, a short post-incident report explains to stakeholders what happened, how long it lasted, what was done, and what will prevent it happening again. Clear, honest communication during an outage does much to preserve confidence in the team, even when the outage itself was their doing.
Practise what you just read
1. An update breaks an application on a physical server with no snapshot. What is the most targeted rollback?
Select one
Show answer
D. Uninstalling the specific update removes only the change that caused the problem, preserving everything else. A full restore loses every change made since the backup was taken.
2. A virtual machine is reverted to its pre-patch snapshot. What is lost?
Select one
Show answer
B. Reverting returns the entire machine to the snapshot moment, discarding application data written since. For databases and file servers this matters as much as the rollback itself.
3. After rolling back a security update that broke an application, what must happen next?
Select one
Show answer
A. The vulnerability the update fixed is open again, and automatic updating may reinstall the faulty update. Blocking it and applying mitigations holds the line until a fixed update arrives.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.