Power faults, and a server that will not POST
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
A server that will not start is one of the most alarming faults, and one of the most often misdiagnosed. "It's dead" can mean it receives no power at all, it powers on but fails its hardware checks, or it passes them and then cannot load an operating system. Each of those points to a different set of causes, and replacing the wrong part wastes time and money while the server stays down.
This lesson applies the troubleshooting method to hardware that will not start: telling the three stages apart, finding power faults, reading what the server reports about itself, tracking down memory and processor faults, and recognising a server that is shutting itself down to avoid overheating.
The lesson
No power, no POST and no boot, and telling them apart
Starting a server happens in stages, and the first job is to identify which stage fails.
- No power. Nothing happens: no fans, no lights, no response to the power button. The problem is between the wall and the motherboard: the supply, the cables, the PSUs, or the motherboard's power circuitry.
- No POST. The server powers on, and fans spin and lights come on, but the power-on self-test, the firmware's check of processors, memory and essential hardware, does not complete. There may be no video, error lights, beep codes or a POST code displayed. The problem is in core hardware or firmware.
- No boot. POST completes, the firmware screen appears, but no operating system loads: a message about no boot device, or a boot loader error. The problem is in storage, boot order, or the operating system, covered in the OS errors lesson.
Asking "how far does it get?" before doing anything else narrows the field enormously. The management controller helps here: because it runs on standby power, it can often be reached remotely even when the server will not start, reporting the power state, hardware health and event log.
Failed power supplies, tripped breakers and PDU faults
For a server with no power, work outward from the server along the power path, as the power lesson laid it out, checking the simple things first:
- Is the server actually receiving power? Check that the cords are firmly seated at both ends. Servers with retention clips on their power inlets exist because cords get knocked loose.
- The PDU. Check that the outlet is live: PDU outlets can be switched off remotely or individually, and PDUs have their own breakers, which trip when their load is exceeded. A metered PDU shows its load and alerts.
- The circuit. A tripped breaker in the building's distribution board cuts every PDU on that circuit. Tripped breakers are often caused by overload: a new server added to a circuit that was already near its limit, or all servers on one circuit after the other failed.
- The UPS. A UPS in fault or overload state, or one with exhausted batteries, may have stopped supplying power.
- The power supplies. Most server PSUs have an LED showing their state, and the management controller reports each PSU's health. A failed PSU in a redundant pair should have raised an alert already, and the server keeps running on the other. If both fail at once, suspect the supply to them rather than the PSUs.
A server with redundant PSUs connected to the same PDU or circuit is not protected against that PDU or circuit failing, which is why the power lesson insisted on separate feeds. Swap-testing a PSU with a known-good spare, or moving a cord to a known-good outlet, isolates the fault quickly.
POST codes, beep codes and the management controller's log
When a server powers on but does not complete POST, it usually tells you why, in one of several ways.
- POST codes are numeric or hexadecimal codes, shown on a small display on the motherboard or front panel, or in the management controller, which mark each stage of POST. The last code shown indicates where the process stopped.
- Beep codes are patterns of beeps that signal particular failures, such as memory not detected. Their meanings vary between firmware vendors and models.
- Diagnostic LEDs on the front panel or beside components light up to show which component has failed, such as a particular memory slot or processor.
- The management controller's log, the system event log (SEL) in IPMI terms, records hardware events with timestamps: failed components, voltage and temperature problems, and POST errors.
Codes differ between manufacturers and models, so look them up in the server's own documentation, not a general list. The management controller's log is usually the most informative source of all, since it is readable remotely and often names the failed component outright. It also records what happened before the failure, which can show a pattern, such as a memory module reporting correctable errors for weeks before it failed completely.
Memory and CPU faults
Memory is among the most common causes of POST failure and instability.
- A server with no usable memory, or memory in the wrong slots, will not POST. Server memory must be installed following the manufacturer's population rules, which specify slot order and matching modules across channels and processors.
- Modules that are incompatible, such as a mix of registered and unregistered memory, or unsupported speeds, can stop POST or reduce performance.
- ECC memory, standard in servers, corrects single-bit errors and logs them. A rising count of correctable errors on one module is a warning to replace it before it produces uncorrectable errors, which crash the server.
- Faulty modules are isolated by checking the log and diagnostic LEDs, then by removing or swapping modules one at a time. Many servers can also disable a failing module automatically and continue with less memory.
Processor faults are rarer. Signs include a server that fails POST with a processor error, a processor missing from the system inventory, or repeated machine check errors in the log. Causes include bent socket pins after a processor was fitted, a processor not supported by the current firmware, often fixed with the firmware update from the firmware lesson, and inadequate cooling. Always check the socket carefully when a processor has been installed or reseated recently.
Reseating a component, removing it and fitting it again, fixes a surprising number of faults caused by poor contact, especially after a server has been moved. Take the ESD precautions from the hot-swap lesson whenever you do.
Overheating and thermal shutdown
A server that shuts down unexpectedly, restarts repeatedly, or slows under load may be protecting itself from heat. Servers monitor temperatures constantly, and when components get too hot, they first throttle, slowing processors down, and then perform a thermal shutdown to prevent damage.
Signs of a thermal problem include:
- fans running at full speed;
- temperature warnings or thermal events in the management controller's log;
- shutdowns under heavy load, or at the hottest time of day;
- performance falling for no other apparent reason, as processors throttle.
Common causes, many from the cooling lesson in the hardware domain:
- a failed fan, reported by the management controller;
- blocked airflow: missing blanking panels, cables blocking vents, or equipment installed facing the wrong way, drawing hot air from the hot aisle;
- dust clogging heatsinks and filters;
- a heatsink poorly seated, or thermal paste dried out or not applied after a processor was replaced;
- room cooling failing or overloaded, which the environmental monitoring in the environmental controls lesson should have reported, and which affects many servers at once.
A thermal fault on one server points to that server; thermal alerts across a row or a room point to the room.
Practise what you just read
1. A server powers on, its fans spin and lights come on, but there is no video and it beeps a pattern. Which stage is failing?
Select one
Show answer
B. Fans and lights show power is arriving. Beeps and no video mean the power-on self-test failed, pointing at core hardware such as memory, processor or firmware rather than power or the OS.
2. A server with redundant power supplies loses both at once. What should be suspected first?
Select one
Show answer
C. Two supplies failing together is far more likely to be a common feed, a tripped breaker or a PDU fault, than two independent failures. That is why supplies belong on separate circuits.
3. Which source usually identifies a failed component most directly on a server that will not POST?
Select one
Show answer
D. The BMC runs on standby power and logs hardware events, often naming the failed component, and can be read remotely. The operating system has not started, so its logs contain nothing.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.