Server roles, monitoring, and performance baselines
Listen to this lesson
Every episode of this course is also a podcast: listen on Spotify.
This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.
Why this matters
A server exists to do a job, and the job decides everything about how it is built and watched: which resources it leans on, what counts as healthy, and which failures matter. A database server short of memory and a file server short of disk space are different emergencies with different warning signs.
This lesson covers the common server roles, then how to tell whether a server is performing normally, which is impossible without first knowing what normal looks like, and how to be told about problems without being buried in noise.
The lesson
Common roles: file, print, web, database, directory and application
The exam expects you to recognise the main roles a server plays and what each depends on.
- A file server stores shared files and serves them over SMB or NFS. It leans on storage capacity and throughput, and uses permissions and quotas to control who stores what.
- A print server manages shared printers and their queues and distributes printer drivers to clients.
- A web server, such as IIS, Apache or nginx, serves websites and web applications over HTTP and HTTPS, on ports 80 and 443.
- A database server, running SQL Server, MySQL, PostgreSQL or Oracle, stores and queries structured data. It is typically the most demanding role, heavy on memory and on fast random storage.
- A directory server holds user accounts, groups and policies and handles authentication. On Windows networks this is an Active Directory domain controller, using LDAP and Kerberos. If it is down, people cannot log in.
- An application server runs business applications and the middleware they depend on.
Many other services, such as DNS, DHCP, mail and time, are also roles in the same sense. On Windows Server, roles and features are installed deliberately through Server Manager or PowerShell, and the principle is the same on any platform: install only the roles a server needs. Every extra role is extra attack surface and extra load. Virtualisation makes this easy to follow, because giving each major role its own virtual machine costs little and keeps a problem in one role from affecting another.
Baselines, and why a number means nothing without one
Is 70 per cent processor use a problem? On a lightly used print server it is alarming; on a busy database server at month end it may be perfectly normal. A performance figure means nothing on its own. It only means something compared with what is normal for that server.
A baseline is that comparison point: a record of a server's normal performance, measured over a period long enough to include its ordinary cycles, such as the working day, the week, and busy periods like month-end processing or backups.
Take a baseline when a server enters service, and again after any significant change, such as an upgrade, a new application or a hardware change. Then, when users report that something is slow, you can compare current figures with the baseline and see immediately which resource has moved away from normal, instead of guessing.
CPU, memory, disk and network metrics, and the thresholds that matter
Four resources account for most performance problems, and each has a few measurements that matter.
- Processor. Utilisation percentage, and the processor queue length, the number of threads waiting for a processor. Short spikes are normal; utilisation that stays high with a growing queue means the server needs more processing capacity or less work.
- Memory. Available memory and, more tellingly, paging or swapping: when the operating system runs short of memory it moves data to disk, which is far slower. Heavy, sustained paging is the clearest sign of memory pressure.
- Disk. Free space, and latency, the time each read or write takes. High or rising latency, and long disk queues, mean storage is the bottleneck, and users feel it as slowness everywhere.
- Network. Throughput compared with the link's capacity, and errors and dropped packets, which point to a failing cable, adapter or switch port.
The tools are built in. Windows has Performance Monitor for detailed counters and logging over time, Resource Monitor for a live view, and Task Manager. Linux has top or htop, vmstat for memory and processes, iostat for disk activity, free for memory and df for disk space.
The key distinction is sustained against momentary. A spike to 100 per cent for a few seconds is a server doing its job. The same figure held for an hour is a problem.
Alerts that fire on a real problem and stay quiet otherwise
Monitoring is only useful if the right person is told about a real problem in time. Alerting design decides that.
- Base thresholds on the baseline, not on round numbers, so an alert means "this is abnormal for this server".
- Require a duration. Alert when a threshold is exceeded for several minutes, not for a single sample, so brief spikes do not trigger alerts.
- Use severity levels, so a warning that disk space is getting low is distinguished from a critical alert that a service is down.
- Make every alert actionable. If nobody would do anything about it, it should not be an alert.
- Monitor the service, not just the resources. A check that actually requests a web page or runs a test query catches failures that resource graphs miss.
- Test that alerts fire. An alert that has never been triggered on purpose may not work at all.
The failure to avoid is alert fatigue. When a system sends so many alerts that most are ignorable, people learn to ignore them, and the important one is lost among the noise. A quiet monitoring system in which every alert matters is far more valuable than a noisy one that reports everything.
Event logs and journals, and collecting them in one place
Performance figures say that something is wrong; logs usually say what.
On Windows, Event Viewer holds the main logs: System for the operating system and drivers, Application for software, and Security for logons and auditing. Each event has a level (critical, error, warning or information), a source and an event ID that can be looked up.
On Linux, journald collects logs, read with journalctl: for example journalctl -u followed by a service name shows one service's log, and -p err shows only errors. Traditional log files are kept in /var/log, including the main system log and the authentication log.
Logs spread across many servers are hard to use, and a server that fails or is compromised may lose its own. So production environments forward logs to a central collector, using syslog forwarding on Linux, event forwarding on Windows, or a SIEM, which the security domain covers later in the course. Central collection also makes it possible to line up events across servers, which only works if every server's clock is synchronised, one more reason for the NTP configuration in the installation lesson.
Try it
An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.
Practise what you just read
1. Users report a server is slow. Its processor is at 70 per cent. Why is that figure alone not enough to act on?
Select one
Show answer
C. Seventy per cent might be normal for a busy database at month end and alarming for a print server. Only a baseline of that server's normal behaviour shows whether the figure has moved.
2. A server's available memory is low and it is paging heavily to disk. What does this indicate?
Select one
Show answer
D. Sustained heavy paging means the operating system is short of physical memory and moving pages to disk, which is far slower. The fix is more memory or less demand, not faster disks.
3. Which alert design best avoids alert fatigue?
Select one
Show answer
A. Thresholds based on each server's normal behaviour, which must be exceeded for several minutes, ignore brief spikes and fire on real deviations. Alerts that fire constantly teach people to ignore them.
7 more questions on this objective are part of the full course.
Hands-on labs
Part of the free CompTIA Server+ SK0-005 course — 51 lessons and 72 hands-on labs.
This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.