Documentation that survives the person who wrote it

Listen to this lesson

Episode 24 · 62:12

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 2.7 · Server administration · 30% of the exam

Why this matters

Every server environment is documented somewhere. The question is whether that documentation is written down where others can find it, or held in the head of the one administrator who built everything. The second kind works until that person is on holiday, ill, asleep during an outage, or leaves the organisation.

Good documentation is what lets a colleague fix a server they have never seen, at three in the morning, without guessing. This lesson covers what to record, how assets are tracked, why changes need records, what a server looked like when it was built, and the easily missed problem of documentation that goes offline with the systems it describes.

The lesson

What to document: configurations, diagrams, runbooks and contacts

Four kinds of documentation cover most of what an administrator needs.

Configuration documentation records how each server is set up: its name, role, addresses, operating system and version, installed roles and applications, storage layout and RAID configuration, and anything unusual about it. This is what the addressing lesson meant by recording addresses so the next change does not collide.

Diagrams show how things connect. A physical diagram shows racks, cabling and which port connects to which; a logical diagram shows networks, subnets, VLANs and how traffic flows between servers and services. A rack diagram or rack elevation shows what sits in each unit of each rack.

Runbooks, also called procedures or playbooks, are step-by-step instructions for specific tasks: restarting a service in the right order, failing over a cluster, restoring from backup. A good runbook is written so someone unfamiliar with the system can follow it under pressure, and tested by having someone do exactly that.

Contacts are the part most often missing when needed: who owns each application, how to reach vendor support, what the support contract numbers are, who is on call, and who must be told about an outage.

Asset tags, inventory and the configuration management database

Organisations need to know what equipment they have, where it is and who is responsible for it.

Each device is given an asset tag, a label with a unique number, often as a barcode or QR code, fixed to the hardware. The tag links the physical object to its record in an inventory, which holds its make, model, serial number, location, owner, purchase date and warranty. When a server is moved, repaired or retired, the record is updated with it.

A configuration management database (CMDB) goes further. It stores configuration items, which include hardware, software, services and documents, and, crucially, the relationships between them: this application runs on these two servers, which use this storage array and sit behind this load balancer. Those relationships answer questions that matter during changes and outages: if this server goes down, what else is affected?

An inventory or CMDB is only useful if it is accurate, which means it must be updated as part of every change, not once a year. Automated discovery tools that scan the network help find equipment that was installed and never recorded.

Change records, and why they matter at three in the morning

When a service breaks, the first question in troubleshooting is what changed? A change record is how that question gets answered.

Change management is the process of proposing, approving, carrying out and recording changes. A change record typically includes what is being changed and why, who is making it, when, the expected impact, the rollback plan if it goes wrong, and afterwards the outcome. Significant changes are reviewed and approved before they happen, often by a change advisory board, and scheduled in agreed maintenance windows.

The value of change records shows at three in the morning, when an application has stopped working and the person on call finds that a firewall rule was altered at 22:00 that evening. Without the record, they may spend hours looking in the wrong place. With it, they have their most likely cause and, from the rollback plan, the way to undo it.

The same records keep documentation current: a change is not complete until the configuration documentation, diagrams and CMDB reflect it.

Baselines and as-built documentation

As-built documentation records a server exactly as it was actually built and handed over, which is not always the same as the design it was built from. Plans change during installation: a different disk layout, an extra network adapter, a changed address. The as-built record captures reality, including hardware, firmware versions, operating system build, installed software and configuration.

A configuration baseline is the agreed standard configuration a server should match. Baselines serve two purposes. New servers can be built to them, so every server of a type is the same. And existing servers can be compared against them to find configuration drift, the gradual, undocumented changes that make one server behave differently from its supposedly identical neighbours.

This complements the performance baseline from the monitoring lesson. The performance baseline records how a server normally behaves; the configuration baseline records how it is normally set up. Both let you answer "what is different?" when something goes wrong.

Keeping documentation reachable during an outage

There is a trap in storing documentation on the systems it documents. If the runbook for restoring the file server is kept on the file server, it is unavailable at the exact moment it is needed. The same is true of a wiki on a virtual machine in the cluster that has failed, or a password vault that needs the directory service that is down.

Documentation that matters during an outage should be reachable without the infrastructure it describes:

  • keep a copy off site or in a separate service, such as a cloud-hosted documentation system independent of the main environment;
  • keep printed or offline copies of the most critical runbooks, recovery procedures and contact lists, stored securely, since documentation describes exactly how to get into the systems;
  • make sure the people who will need it know where the copy is, and that it is updated when the original is.

Disaster recovery plans, covered in the troubleshooting domain later in the course, depend on this completely: a recovery plan that can only be read once the systems are recovered is not a plan.

Practise what you just read

1. An application stops working at 03:00. What record most quickly answers 'what changed?'

Select one

  1. The change records
  2. The asset register
  3. The network diagram
  4. The visitor log
Show answer

A. Change records show what was altered, when and by whom, with a rollback plan. A change made that evening is usually the first suspect, and the record shows how to undo it.

2. Where should the runbook for restoring the file server be kept?

Select one

  1. In the file server's own event log, with its errors
  2. Somewhere reachable when the file server is down
  3. On the file server, next to the data it describes
  4. In the head of the administrator who first set it up
Show answer

B. Documentation stored on the system it describes is unavailable exactly when it is needed. Off-site, separate-service or printed copies keep runbooks reachable during outages for the people who need them.

3. What does a CMDB record that a simple inventory does not?

Select one

  1. Only each item's hardware serial number
  2. The purchase price of each item that it lists
  3. Relationships between configuration items
  4. User passwords for each managed server
Show answer

C. A configuration management database stores items and how they depend on each other, such as which applications run on which servers, answering what else is affected when something fails.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.