OS errors and boot failures

Listen to this lesson

Episode 42 · 51:52

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 4.2 · Troubleshooting · 28% of the exam

Why this matters

The power and POST lesson separated a server that will not power on, one that fails its hardware checks, and one that passes them but cannot load an operating system. This lesson picks up at that third stage, and at the operating system failures that happen after a successful start: the server that crashes to a blue screen, the Linux host that stops with a kernel panic, and the server that boots after a hardware change and cannot find its own disks.

Operating system faults are often recoverable without reinstalling, provided you know where to look and which tools the platform provides for getting a broken system far enough to fix it. This lesson covers boot problems, crashes and the files they leave behind, drivers after hardware changes, recovery environments, and reading the system log.

The lesson

Boot loader and boot configuration problems

After POST, the firmware hands control to a boot loader on the boot disk, which loads the operating system. Several things can break this chain.

Boot order. The firmware may be trying to boot from the wrong device: a USB stick left in a port, a network boot attempt, or a new disk placed ahead of the boot disk. Symptoms include "no bootable device" or "operating system not found". Check the boot order in the firmware setup.

Firmware mode. A system installed in UEFI mode will not boot if the firmware is switched to legacy BIOS mode, or the reverse. Settings can change after a firmware update or reset, as the firmware lesson warned. Secure Boot can also block a boot loader or driver that is not signed.

Boot configuration. The boot loader's configuration may be damaged or point to the wrong place:

  • On Windows, boot settings are held in the Boot Configuration Data (BCD) store. The bcdedit command views and edits it, and in the recovery environment, bootrec and bcdboot can rebuild it and restore the boot files.
  • On Linux, the boot loader is usually GRUB. A missing or damaged GRUB, or a configuration pointing to a disk that has changed, stops the boot. The configuration is regenerated with grub2-mkconfig or update-grub, and GRUB can be reinstalled from a rescue environment. A wrong entry in /etc/fstab, the list of file systems to mount, is another common cause of a Linux server stopping partway through boot, typically dropping into an emergency shell.

Also check that the boot disk or array is visible at all, since a failed array or disconnected disk shows the same symptoms, and the storage lesson's checks come first.

Stop errors, kernel panics and dump files

When an operating system meets an error it cannot safely recover from, it stops deliberately, to avoid corrupting data.

On Windows, this is a stop error, commonly called a blue screen. It shows a stop code, such as a name like IRQL_NOT_LESS_OR_EQUAL, and often the name of the driver involved. On Linux, the equivalent is a kernel panic, which prints diagnostic information to the console.

The most common causes are faulty drivers, failing hardware, particularly memory, corrupt system files, and overheating. A crash that started after a driver or update was installed points to that change; crashes at random times, with different stop codes, suggest hardware, and memory in particular.

Crashes leave evidence:

  • Windows writes a memory dump file, by default MEMORY.DMP in the Windows folder, plus smaller minidumps, recording the state of the system at the moment of the crash. Debugging tools such as WinDbg analyse them to identify the responsible driver or module.
  • Linux can capture crash dumps with kdump, which, once configured, saves the kernel's memory to disk for analysis.

Make sure dump settings are configured before they are needed, with enough disk space for the dump, so that a crash leaves something to investigate. On servers that restart automatically after a crash, the dump file and the system log may be the only record that anything happened.

Missing drivers after a hardware change

A server that worked yesterday may fail to boot after a hardware change, such as a new storage controller, a replaced motherboard, or a migration to different hardware or a virtual machine, as the migration lesson described for P2V.

The classic symptom is the operating system starting to load and then failing because it cannot find its boot disk: on Windows, a stop error such as INACCESSIBLE_BOOT_DEVICE; on Linux, a failure to find the root file system, dropping into an emergency shell. The cause is that the driver for the new storage controller is not in the installed system, or not loaded early in boot.

Remedies:

  • install the new controller's driver before making the hardware change, where possible;
  • on Windows, inject the driver from the recovery environment, or revert the hardware and add the driver first;
  • on Linux, the driver must be included in the initramfs, the initial file system loaded at boot; rebuild it with dracut or update-initramfs after installing the driver;
  • check the controller's mode in the firmware: changing a controller between modes, such as AHCI and RAID, has the same effect as a new controller.

Other devices without drivers, such as a new network adapter, do not stop the boot but simply do not work. Device Manager shows them as unknown devices, and lspci on Linux lists hardware whether or not a driver is loaded.

Recovery environments and safe mode

When a system will not start normally, each platform offers a way to start it minimally, far enough to fix the problem.

On Windows:

  • Safe Mode starts Windows with only essential drivers and services. If a server starts in Safe Mode but not normally, the cause is likely a driver or service that Safe Mode leaves out, which can then be disabled or removed.
  • The Windows Recovery Environment (WinRE) provides a command prompt, startup repair, system restore where configured, the option to uninstall recent updates, and tools such as bootrec, chkdsk and sfc, which checks and repairs protected system files.

On Linux:

  • The GRUB menu offers older kernels, often the quickest fix after a kernel update causes problems.
  • Rescue or single-user mode, and the emergency target, start a minimal system with a root shell, from which files such as fstab can be corrected.
  • Booting from installation or rescue media allows the installed system to be mounted and repaired from outside.

On physical servers, the management controller's remote console and virtual media make all of this possible without standing in front of the machine, as the out-of-band management lesson described.

Reading the system log

The system log is where the operating system records what it has been doing, and it is usually the most direct route to the cause of an OS problem.

On Windows, Event Viewer's System log records drivers, services, crashes and hardware events. Useful events include unexpected shutdowns, recorded by the Kernel-Power source with event ID 41 when the system restarted without shutting down cleanly, and BugCheck events recording stop errors. Filter the log by level, time and source, and look at what happened just before the failure.

On Linux, journalctl shows the journal. journalctl -b shows messages from the current boot, and journalctl -b -1 from the previous one, which is essential after a crash, since the problem happened in the boot before this one. journalctl -k shows kernel messages, as does dmesg, which is where hardware and driver errors appear. Traditional logs in /var/log hold the same information on many systems.

Read logs with the troubleshooting method in mind: establish when the problem started, look for the first error rather than the many that follow from it, and compare with what the change records say happened around that time.

Try it

An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.

Practise what you just read

1. After a new storage controller is installed, Windows starts to load and stops with INACCESSIBLE_BOOT_DEVICE. What is the likely cause?

Select one

  1. The Windows licence has expired on the server
  2. The driver for the new controller is missing
  3. The boot disk failed at exactly the same moment
  4. The network adapter is unplugged from the switch
Show answer

B. Windows cannot reach its boot disk without a driver for the controller it sits behind. Installing the driver before the change, or injecting it from recovery, resolves it.

2. A Linux server drops into an emergency shell at boot after a new data disk entry was added to fstab. What is the likely cause?

Select one

  1. A wrong or missing device in /etc/fstab
  2. The kernel mounts one file system at boot
  3. A failed network interface on the server
  4. A full /tmp directory on the root disk
Show answer

A. systemd waits for every fstab entry, and a missing or mistyped device fails the mount and stops the boot. Correcting the entry, or marking non-essential mounts nofail, fixes it.

3. Which Linux command shows the error messages from the previous boot, after a crash?

Select one

  1. last -b, to list errors logged at the last boot
  2. dmesg -1, to print the last boot's kernel buffer
  3. journalctl -b -1 -p err, to filter for errors
  4. systemctl status, to show units that failed
Show answer

C. journalctl -b -1 selects the previous boot, and -p err limits it to errors, which is essential after a crash because the evidence belongs to the boot that failed.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.