The troubleshooting method for servers

Listen to this lesson

Episode 38 · 44:01

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 4.1 · Troubleshooting · 28% of the exam

Why this matters

When a server fails, the pressure is to do something immediately: restart it, replace a part, roll back the last change. Sometimes that works. Often it destroys the evidence of what went wrong, fixes a symptom while the cause remains, or makes things worse, such as restarting a server whose disks are failing and which then does not come back.

A method keeps troubleshooting systematic under pressure. The troubleshooting domain is the largest on the exam, and scenario questions often test whether you know which step comes next, not just the technical fix. This lesson covers the method, the information to gather first, the steps from theory to verification, documenting as you go, and knowing when a problem belongs to someone else.

The lesson

CompTIA's troubleshooting steps, and why the order matters

The troubleshooting method used throughout CompTIA's exams runs in this order:

  1. Identify the problem. Gather information, question users, identify recent changes, determine the scope, and reproduce the problem if possible.
  2. Establish a theory of probable cause. Start with the obvious and the simple, and consider several possibilities.
  3. Test the theory to determine the cause. If the theory is confirmed, move on; if not, form a new theory, or escalate.
  4. Establish a plan of action to resolve the problem, and identify its potential effects.
  5. Implement the solution, or escalate if necessary.
  6. Verify full system functionality and, where applicable, implement preventive measures.
  7. Document findings, actions and outcomes.

The order matters. Each step depends on the one before. A theory formed without information is a guess. A fix applied before its theory has been tested may change nothing, or change something else. A problem declared solved without verification may return that evening. Exam questions often describe a technician doing things out of order, such as replacing parts before asking what changed, and ask what they should have done first.

For servers, one step is often added within the first: back up before making changes, where possible, because a troubleshooting action that loses data turns one problem into two.

Gathering information: logs, users and recent changes

Identifying the problem is where most troubleshooting succeeds or fails.

Ask the people involved. Users and application owners describe symptoms: what exactly does not work, the precise error message, when it started, and whether it affects everyone or only some. Ask open questions and write down the answers. "The server is slow" might mean one report, one user, or every application.

Determine the scope. Is it one server, several, or everything in a site? One application, or all applications on the server? One user, or all users? Scope separates a failed disk in one server from a failed switch affecting many, and tells you how urgent the problem is.

Find out what changed. Problems very often follow changes: patches, new software, configuration edits, hardware work, even changes elsewhere, such as a firewall or DNS update. The change records described in the documentation lesson exist to answer this question.

Collect evidence. Read the system and application logs, the management controller's hardware log, monitoring history, and alerts. Compare current performance with the baseline from the monitoring lesson. Collect this before restarting anything, since a restart may clear the very state that explains the problem.

Reproduce it if you can, safely, because a problem that can be made to happen on demand can be tested.

Theory, test, plan, implement and verify

Theory. From the information gathered, list the likely causes, starting with the simplest: is it plugged in, is the service running, is the disk full, has a password expired? These account for a large share of problems. Look for a common element when several symptoms appear at once: several failing applications on one server may all share one full disk.

Test. Check the theory without changing anything important where possible: examine the disk space, check the service status, test with a known-good cable. If the test disproves the theory, move to the next one. If every theory fails, gather more information or escalate.

Plan. Before fixing, plan the fix. What exactly will be done? What could it affect, such as a restart interrupting other services on the server? Does it need a change approval or a maintenance window? What is the rollback if it goes wrong?

Implement. Carry out the plan, changing one thing at a time, so that if it works you know why, and if it does not, you have not introduced several new variables.

Verify. Confirm that the whole system works, not just the part that was broken. Ask the users who reported it to confirm. Check that monitoring is green, that dependent services work, and that nothing else was affected. Then consider preventive measures: what would stop this recurring? Where the cause is not obvious, a root cause analysis asks why it happened, not just what failed: the disk filled, but why was nobody alerted, and why was log rotation not configured?

Documenting what was found and what was done

Documentation is the last numbered step, but it happens throughout. Note observations, times, and each action as you take it, rather than trying to reconstruct them afterwards.

The final record, usually in the ticketing or incident system, includes:

  • the symptoms and how the problem was reported;
  • the cause that was found;
  • the actions taken, in order, including those that did not work;
  • the outcome, and how it was verified;
  • any follow-up: preventive measures, changes to monitoring, or documentation to update.

This serves several purposes. The next person who meets the same symptoms can search for it and find the answer in minutes. Recurring problems become visible when records are compared. Actions taken during an incident are recorded for anyone reviewing it later. And a knowledge base built from these records is how a team's experience survives staff changes, as the documentation lesson argued.

Knowing when to escalate

Escalation means passing a problem to someone with more expertise, more authority or more access. It is a normal part of the method, not a failure.

Escalate when:

  • the problem is beyond your knowledge or your permissions, such as a storage array problem needing a specialist, or a change needing administrative rights you do not hold;
  • time matters more than learning: if a critical service is down and you have not made progress in a reasonable time, the business needs the problem solved, not solved by you;
  • the problem is someone else's system: the network team's switch, the application vendor's software, or hardware still under a support contract, where the vendor must do the repair;
  • you suspect a security incident, which follows the incident response plan and goes to the security team, as the SIEM lesson described;
  • the fix needs a decision beyond your authority, such as an outage during business hours.

A good escalation passes on everything you have learned: symptoms, scope, evidence, theories tested and their results, and anything already changed. Nobody should have to start again from the beginning. Organisations define escalation paths and contacts in advance, as part of the documentation covered earlier in the course.

Practise what you just read

1. A technician hears that a file server is slow and immediately replaces its network card. Which step of the troubleshooting method was skipped?

Select one

  1. Identifying the problem and establishing a theory
  2. Documenting findings before any hardware is replaced
  3. Verifying full system functionality afterwards
  4. Escalating the fault to the vendor's support
Show answer

A. Replacing parts before gathering information and forming a tested theory is guessing. The method starts by identifying the problem, including what changed, and only acts once a theory has been tested.

2. What should a technician ask early when identifying a server problem?

Select one

  1. How old the server hardware is
  2. Which vendor made the server
  3. What has changed recently
  4. Who installed the operating system
Show answer

C. Problems very often follow changes such as patches, configuration edits or hardware work. Change records answer the question quickly and usually point to the most likely cause.

3. A theory of probable cause is tested and proves wrong. What is the next step?

Select one

  1. Document the problem as resolved
  2. Establish a new theory, or escalate
  3. Reboot the server and test again
  4. Apply the original theory's fix anyway
Show answer

B. When a test disproves a theory, the method returns to forming another one from the information gathered, gathering more if needed, or escalates if no theory can be confirmed.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.