Firmware, UEFI and driver updates without bricking the server
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
Firmware is the software built into hardware, and a server is full of it. Updating it fixes bugs, closes security holes and adds support for newer components. It is also one of the few routine tasks that can leave a server unable to start at all, because a firmware update interrupted halfway, or applied in the wrong order, can leave a component with no working code.
That risk is manageable. It is managed with the same discipline every time: know what you are updating, read what the vendor says about it, test it, schedule it, and have a way back.
The lesson
Where firmware lives: UEFI, the BMC, the RAID controller, NICs and drives
A single server can carry a dozen separate firmware images:
- System firmware, the UEFI or BIOS, which initialises the hardware and starts the boot process.
- BMC firmware, for the out-of-band management controller covered in the previous lesson.
- Storage controller firmware, for the RAID controller or host bus adapter.
- Network adapter firmware, for each NIC.
- Drive firmware, on every hard disk and SSD.
- Smaller components such as power supplies and backplanes.
Firmware has a partner on the software side. Drivers are the operating system's code for talking to a device, and the two must be compatible: a new controller firmware may require a newer driver, and an old driver may misbehave with new firmware.
This is why server vendors publish tested bundles of firmware and drivers together, such as HPE's Service Pack for ProLiant and Dell's update packages applied through the Lifecycle Controller. A bundle is a combination the vendor has tested as a set, which is safer than picking the newest version of each component individually.
Update order, dependencies, and reading the release notes first
Every firmware update comes with release notes, and reading them before applying the update is the single most useful habit in this objective. They state:
- what the update fixes and whether it is a security update;
- prerequisites and minimum versions: an update may require another component to be at a certain version first, or may need an intermediate version before the latest can be applied;
- known issues introduced by the update;
- whether the update can be rolled back, because some cannot.
Order matters when components depend on each other. A common pattern is to update the BMC first, since it is often the component that applies the other updates, then the system firmware, then controllers and devices, and then the matching drivers in the operating system. Follow the vendor's stated order where one is given.
Do not update for its own sake. Good reasons are a security fix, a fix for a problem you actually have, or keeping to the vendor's supported baseline. A working server updated without a reason is taking a risk for nothing.
UEFI against legacy boot, and what Secure Boot checks
Modern servers use UEFI (Unified Extensible Firmware Interface), which replaced the older BIOS. UEFI can boot from disks partitioned with GPT, which allows boot volumes larger than 2 TB; it starts faster; and it supports Secure Boot. Most UEFI firmware can also emulate the old behaviour through a legacy or compatibility (CSM) mode, which boots from MBR disks.
The boot mode must match how the operating system was installed. An operating system installed in UEFI mode will not boot if the firmware is switched to legacy mode, and the reverse is also true. Changing the boot mode on a working server is therefore a common way to make it fail to boot, and a common exam scenario.
Secure Boot protects the start-up process itself. The firmware holds databases of trusted and revoked signing keys and certificates, and before it runs a boot loader, an operating system kernel or a device's option ROM, it checks that the code is signed by a trusted key and has not been revoked. Unsigned or tampered code does not run. This defeats bootkits, malware that hides in the boot process before the operating system and its security tools start. The cost is that unsigned drivers or custom kernels may be blocked until they are signed or their key is enrolled.
Staging an update with a rollback plan
Firmware updates should never reach production servers untested.
- Test first on a non-production server of the same model, or the least critical member of a group, and confirm it boots and runs its workload normally.
- Update redundant systems one at a time. In a cluster, move the workload off one node, update it, confirm it is healthy, bring it back, then move to the next. This is a rolling update, and it keeps the service running.
- Protect the power. An update interrupted by a power loss is the classic way to leave a component with corrupt firmware, so updates should run on equipment protected by a UPS.
- Know the way back. A virtual machine snapshot does not cover firmware. Rolling firmware back means reinstalling the previous version, so keep it available, and check whether the vendor supports downgrading. Many servers keep a backup or recovery firmware image, and BMCs often offer a rollback to the previous version.
- Verify afterwards: confirm the new versions are reported, check the hardware health and the event log, and check that the workload is normal.
Maintenance windows and the change record that goes with them
A maintenance window is an agreed, scheduled time when a system may be disrupted, usually outside business hours, and announced to the people who depend on it. Firmware updates that need a reboot belong in one.
Each update also belongs in a change record, raised and approved through the organisation's change management process, often by a change advisory board. The record states what will change, why, the risk, the tested rollback plan, who will do it and when. It exists so that the change is deliberate, reviewed and reversible, and so that when something unexpected happens later, anyone investigating can see what changed and when.
After the work, update the record with what was actually done, and update the server's documentation with the new firmware versions. The next person to plan an update will start from those versions.
Practise what you just read
1. A server was installed in UEFI mode. After a firmware reset it will not boot. What is a likely cause?
Select one
Show answer
C. An operating system installed in UEFI mode needs UEFI boot. A reset can restore defaults, including legacy mode or a different boot order, leaving the firmware unable to find the boot loader.
2. What does Secure Boot do on a server?
Select one
Show answer
B. Secure Boot checks digital signatures on the boot loader and early drivers against keys in the firmware, blocking unsigned or tampered code such as bootkits. Disk encryption is a separate feature.
3. Before updating a server's firmware, which step matters most?
Select one
Show answer
D. Firmware updates can have dependencies on drivers or other firmware, and some cannot be reversed. Checking the vendor's compatibility information and planning a rollback turns a failed update into a recoverable one.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.