Patching and OS updates for servers
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
The malware lesson made the case that patching is among the most effective defences available. Patching servers is harder than patching workstations, though. A patch can break the application a server runs, most patches need a reboot that interrupts service, and a busy environment has hundreds of servers on different operating systems, each with its own schedule and owner.
So organisations patch through a process, not by clicking "install updates" whenever someone remembers. This lesson covers that process, the infrastructure that delivers updates, the windows that schedule them, what happens when a patch cannot wait, and how to know which servers are actually up to date.
The lesson
A patch process: test, stage, deploy and verify
A reliable patch process has four stages.
Test. When updates are released, such as Microsoft's monthly Patch Tuesday or a Linux distribution's security advisories, the patch team reviews what each fixes and how severe the vulnerabilities are. Updates are installed first on test servers, ideally in an environment that mirrors production, and checked: do the applications still work, do services start, does anything behave differently?
Stage. Updates are then rolled out in rings or waves, starting with a small group of less critical production servers, a pilot group, and widening only if no problems appear. Clustered and load-balanced servers are patched one node at a time, moving work off each first, so the service stays up throughout.
Deploy. The remaining servers are patched during their scheduled maintenance windows, using central tools rather than by hand, so nothing is missed and every server gets the same updates.
Verify. After deployment, confirm that updates actually installed, servers rebooted where needed, services are running and applications work. Monitoring alerts and application health checks help here.
Throughout, keep a way back. Take a snapshot of virtual machines or ensure a recent backup exists before patching, and know how to uninstall an update if it causes problems. Recovering from a failed patch is covered in the troubleshooting domain later in the course.
Update servers and repositories
In a large environment, servers do not download updates individually from the internet. Central update infrastructure gives administrators control over which updates are installed, when, and where.
On Windows, WSUS (Windows Server Update Services) downloads updates once from Microsoft and serves them to internal servers. Administrators approve updates for particular groups of computers, so an update can be approved for the test group first and for production later. Microsoft Configuration Manager extends this with scheduling and reporting, and cloud services such as Azure Update Manager offer the same for cloud and hybrid servers. Group Policy tells each server where to get updates and when to install them.
On Linux, updates come from package repositories, managed with the distribution's package manager: apt on Debian and Ubuntu, dnf or yum on Red Hat and its relatives. Organisations run local mirrors or repository management tools, such as Red Hat Satellite, which can hold snapshots of repositories. A tested snapshot is promoted from test to production, so every production server installs exactly the versions that were tested, not whatever happened to be newest on the day it was patched.
Firmware and applications need their own channels: the management controller and vendor tools for firmware, as in the firmware lesson, and each application's own update mechanism.
Maintenance windows and reboots
Many updates only take effect after a reboot, because files in use by the running system cannot be replaced while it runs. Kernel updates on Linux and cumulative updates on Windows nearly always require one. A server that has installed updates but not rebooted is not yet protected, and may be in an inconsistent state.
Reboots interrupt service, so they happen in maintenance windows: agreed, recurring periods, typically at night or at weekends, when service interruptions are acceptable and users have been told to expect them. Windows are agreed with the owners of each system, recorded in change management, and matched to the business, since a retailer's quiet time is different from a hospital's.
Good practice around reboots:
- Schedule reboots rather than allowing servers to restart themselves whenever updates finish.
- Reboot servers in dependency order: databases before the applications that use them, and never every domain controller at once.
- Check each server after rebooting, since a server that fails to come back is far better discovered during the window than by users the next morning.
- Watch for pending reboots, servers with installed updates awaiting a restart, which management tools and scripts can report.
Some platforms reduce the need for reboots, such as live kernel patching on Linux, which applies certain kernel fixes to a running system.
Emergency patches outside the window
Sometimes a vulnerability cannot wait for the next maintenance window. A critical vulnerability that is being actively exploited, especially in a system exposed to the internet, may justify patching within hours or days. These are often called out-of-band patches when a vendor releases them outside its normal schedule.
Emergency patching uses the organisation's emergency change process: the same steps as a normal change, compressed, with approval from a smaller group, and recorded afterwards. Testing is shortened but not abandoned: even a brief test on one server can catch a patch that breaks the application.
The decision weighs two risks: the risk of an outage from applying a patch with little testing, against the risk of compromise from leaving it unapplied. Where the patch cannot be applied immediately, mitigations buy time, such as disabling the affected feature, restricting access to the vulnerable service, or adding a firewall rule, as vendors often describe in their advisories.
Tracking which servers are compliant
Deploying patches is not the same as having servers patched. Installations fail, servers are offline during the window, reboots are postponed, and new servers appear that no one added to the patch groups.
So organisations track patch compliance: which servers have which updates, and which are missing required ones. The tools that deploy updates, such as WSUS, Configuration Manager and Linux management platforms, report each server's status. Vulnerability scanners check independently, by examining servers for known vulnerabilities, which also catches software the patch tools do not manage.
Compliance reporting should answer:
- which servers are missing critical or security updates, and for how long;
- which servers have a pending reboot;
- which servers have not reported at all, which may mean a broken agent, or a server nobody knew existed;
- which servers run unsupported operating systems that no longer receive patches, and so can never be compliant.
Organisations set targets, such as critical patches applied within a number of days, and measure against them. Servers that cannot be patched, perhaps because a vendor application does not support the update, are recorded as exceptions, with the risk accepted by their owner and extra protection such as network isolation.
Practise what you just read
1. Patches have installed on a server but it has not been rebooted. What is its security state?
Select one
Show answer
D. Many updates, including kernel and cumulative updates, only take effect after a reboot because files in use cannot be replaced. A pending reboot leaves the vulnerability open.
2. What is the purpose of a pilot group in a patch rollout?
Select one
Show answer
C. Updates are deployed in rings, starting with a small pilot group. Problems show up there and can be fixed or rolled back before the update reaches every server.
3. An actively exploited critical vulnerability affects an internet-facing server, and the next maintenance window is in three weeks. What should happen?
Select one
Show answer
A. Actively exploited critical vulnerabilities on exposed systems justify an emergency change with compressed testing. If the patch cannot be applied at once, mitigations reduce the risk meanwhile.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.