Write one detection per indicator, and prove none of them cries wolf
Task
Turn the lesson's indicator table into working detections over logs you generate, then grade each detection twice: does it fire on the event it is for, and does it stay silent on ordinary traffic? A rule that fires on everything is ignored within a week, so false positives are half the mark.
Steps
- Write
/tmp/genlogs.pythat produces three files:/tmp/web.log(client, time, method, path, status, bytes),/tmp/dns.log(time, client, query type, name) and/tmp/auth.log(time, source, user, event). Give them at least 500 ordinary lines between them -- varied users, paths, names and sources -- so that normal has some texture. - Have the same script plant a known number of indicator lines and record each one in
/tmp/truth.csv, with the headerfile,line_number,attack. Plant at least three each of: directory traversal (parent-directory sequences in a path, plain and percent-encoded), ssrf (a request whose URL parameter points at an internal or link-local metadata address), dns tunnelling (long, random-looking labels in text-record queries to one domain), spraying (one source, one failure each across many users) and privilege escalation (a standard user added to an admin group). - Write
/tmp/detect.pywith one function per attack. Together they write/tmp/alerts.csvwith the same header and one row per line they flag, naming the attack. - Spraying cannot be judged one line at a time: detect it in aggregate -- distinct users per source per window -- and flag every line belonging to a source over your threshold. Record the threshold and why you chose it in
/tmp/detections.md. - Run the detector and compare
/tmp/alerts.csvwith/tmp/truth.csv. Tune each rule until it catches every planted line and flags no ordinary one. - Then add ten tricky but innocent lines to the generator -- a path that happens to contain
etc, a long but legitimate CDN hostname, one user mistyping their own password twice -- rerun everything, and record in/tmp/detections.mdwhich rules needed tightening.
Verify
wc -l /tmp/web.log /tmp/dns.log /tmp/auth.log
python3 /tmp/detect.py
python3 - <<'PY'
import csv, collections
key = lambda r: (r['file'].strip(), r['line_number'].strip())
truth = {key(r): r['attack'].strip().lower() for r in csv.DictReader(open('/tmp/truth.csv'))}
alerts = {key(r): r['attack'].strip().lower() for r in csv.DictReader(open('/tmp/alerts.csv'))}
per = collections.Counter(truth.values())
print('planted:', dict(per))
assert len(per) >= 5 and min(per.values()) >= 3, 'plant three lines for each of five attacks'
missed = [k for k in truth if k not in alerts]
noise = [k for k in alerts if k not in truth]
misnamed = [k for k in truth if k in alerts and alerts[k] != truth[k]]
print('missed', len(missed), '| false positives', len(noise), '| misnamed', len(misnamed))
assert not missed and not noise and not misnamed, 'tune until all three counts are zero'
PY
The assertion grades both directions at once. Missed lines mean a rule is too narrow; false positives mean it is too broad; misnamed lines mean two rules overlap. All three must be zero, and they must stay zero after the tricky lines in the last step, which is where most first drafts fail.
Notes
Every rule you wrote is a statement about what normal looks like, which is the lesson's closing advice: read for the anomaly, then name it. The spraying rule is the instructive one -- it is the only detection here that cannot exist at the level of a single line, and it is the one most monitoring is configured to miss.
This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.