Take a performance baseline, then catch a deviation from it

short · 50 min · Objective 2.3

Task

Record a performance baseline for lin-srv -- processor, memory, disk and network -- over a quiet period, then generate load and record the same counters again. Compare the two to show which resource moved and by how much. A number means nothing without a baseline; this lab produces both halves.

Steps

  1. With lin-srv idle, record ten one-second readings of each resource: vmstat 1 10, iostat -dx 1 10, sar -n DEV 1 10. Save them to lab/baseline/idle-vmstat.txt, idle-iostat.txt and idle-net.txt.
  2. Summarise the idle values in lab/baseline/summary.csv with header state,cpu_idle_pct,free_mem_mb,disk_util_pct,swap_in as a row with state idle, using the averages.
  3. Start stress-ng --vm 2 --vm-bytes 75% --io 2 --timeout 120s and, while it runs, record the same three captures as load-vmstat.txt, load-iostat.txt and load-net.txt.
  4. Add a load row to the summary.
  5. Record in lab/baseline/alert.txt one alert rule you would set from this baseline, with a threshold and a duration, such as free memory below a value for five minutes.

Verify

These checks run in a POSIX shell: Terminal on macOS or Linux, and on Windows Git Bash (it comes with Git for Windows) or WSL. A stock Windows PowerShell or Command Prompt has no awk or grep, so there the first line fails.

ls lab/baseline/idle-*.txt lab/baseline/load-*.txt | wc -l
awk -F, 'NR>1 {c[$1]=$2; m[$1]=$3} END {print "cpu idle "c["idle"]" -> "c["load"]", free mem "m["idle"]" -> "m["load"]; exit !(c["load"]<c["idle"])}' lab/baseline/summary.csv
grep -Eic 'minute|second|for [0-9]' lab/baseline/alert.txt

Six captures exist, and processor idle time under load is lower than the baseline, so the second command exits zero. Free memory should have dropped too. Your alert must include a duration: a threshold that fires on a single reading alerts on every brief spike, which the monitoring lesson calls alert fatigue.

Notes

On Windows the same work uses Performance Monitor's data collector sets, and the counters to capture are % Processor Time, Available MBytes, Avg. Disk sec/Transfer and Bytes Total/sec.

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.