DLP and SIEM: stopping data leaving, and seeing what happened

Listen to this lesson

Episode 33 · 53:21

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 3.4 · Security and disaster recovery · 24% of the exam

Why this matters

Prevention eventually fails somewhere. A user emails a spreadsheet of customer records to their personal account; an attacker with a stolen password quietly copies files for weeks. Two kinds of system deal with these situations. Data loss prevention watches for sensitive data leaving where it should not. Security information and event management brings together the logs from every server and device, so that attacks can be seen and investigated.

Both depend on work server administrators do: sending logs, keeping clocks right, and knowing what to do when an alert arrives. This lesson covers what each system does, the constant work of tuning alerts, how long logs are kept and why their timestamps must agree, and how to respond when an alert fires.

The lesson

Data loss prevention: what it inspects, and where

Data loss prevention (DLP) systems detect sensitive data and stop it being sent, copied or stored where policy forbids.

DLP recognises sensitive data by inspecting content:

  • Pattern matching finds data with a recognisable format, such as payment card numbers, national identity numbers or bank account details, often checked further, for example with the checksum that valid card numbers pass.
  • Keywords and dictionaries find terms such as project code names.
  • Fingerprinting recognises specific documents, or parts of them, that have been registered as sensitive.
  • Classification labels, as in the data retention lesson, mark files so DLP can act on them directly.

DLP works in three places, matching the states of data:

  • Data in motion, at the network edge and mail gateway, inspecting email, web uploads and file transfers as they leave.
  • Data at rest, on servers and storage, scanning file shares and databases to find sensitive data stored where it should not be, such as card numbers in a share open to everyone.
  • Data in use, on endpoints, controlling copying to USB drives, printing and uploads to personal cloud storage.

When DLP finds a match, it can log it, warn the user, block the action, or encrypt the data automatically. New DLP policies usually start by logging only, so their accuracy can be judged before they begin blocking legitimate work. Encrypted traffic is a limit: DLP cannot inspect what it cannot decrypt.

SIEM: collecting logs, correlating them and alerting

A security information and event management (SIEM) system collects logs and events from across the organisation, including servers, firewalls, directory services, applications and endpoint protection, into one place.

It does three things with them:

  • Aggregation and normalisation. Logs arrive in many formats, which the SIEM converts to a common structure so events from different sources can be searched and compared together.
  • Correlation. The SIEM applies rules that link events across sources. A single failed logon means nothing. Hundreds of failed logons against many accounts from one address, followed by a successful one, then a new administrator account being created, is an attack in progress, and only visible when events from several systems are put together.
  • Alerting and reporting. When a rule matches, the SIEM raises an alert for analysts, and it produces reports and dashboards, including those needed for compliance.

Servers feed a SIEM by forwarding logs, using syslog on Linux and network devices, Windows Event Forwarding, or agents, as the monitoring lesson described for central log collection. A server whose logs are not collected is invisible to the SIEM, which is why checking that every server is forwarding is part of building one.

Tuning alerts against false positives

A SIEM or DLP system, straight out of the box, produces a great many alerts, most of them false positives: alerts about activity that is actually legitimate, such as a backup job reading every file, or an administrator's scheduled script logging on to many servers. A false negative, by contrast, is a real attack that raises no alert at all.

Too many false positives cause alert fatigue, the problem the monitoring lesson warned about: analysts learn to dismiss alerts, and the real attack is lost among them. Tuning is the ongoing work of reducing them:

  • adjusting thresholds, so a rule triggers on patterns that are genuinely unusual for this environment;
  • adding exceptions for known, legitimate activity, kept narrow and documented, like antivirus exclusions;
  • using context, such as asset importance and user role, so alerts on critical servers rank higher;
  • retiring or rewriting rules that never produce a useful alert.

Tuning must not go too far. Every exception is a gap an attacker could use, and a quiet SIEM is only good if it is quiet because nothing is happening.

Log retention, and time synchronisation with NTP

Logs are evidence, and they are often needed long after the events they record. Attacks are frequently discovered weeks or months after they began, and an investigation needs logs from the start. Regulations and standards also set minimum log retention periods. So logs are kept according to a retention policy, as in the data retention lesson, typically with recent logs in fast, searchable storage and older ones archived more cheaply.

Logs must also be protected. Attackers try to delete or alter logs to hide their tracks, which is one reason to send logs off each server to a central collector as they are written, where access is restricted and records cannot be changed.

Correlating events across systems only works if their timestamps agree. If a firewall's clock is four minutes ahead of the web server's, events that happened together appear minutes apart, and an investigation reconstructs the wrong sequence. Every server and device should synchronise its clock with NTP (Network Time Protocol) from the same reliable sources, as the installation lesson set up. In Active Directory, domain members synchronise from domain controllers automatically; the controller holding the PDC emulator role should synchronise from a reliable external source. Recording times in UTC, or with time zones, avoids confusion across sites.

Responding to an alert

An alert is the start of a process, not the end. Organisations define that process in an incident response plan, and the usual stages are:

  1. Triage. Confirm whether the alert is real or a false positive, and judge its severity. What system, what account, what data is involved?
  2. Containment. Stop it spreading: isolate the affected server from the network, disable a compromised account, block an address. Where possible, isolate rather than switch off, since memory contents can be valuable evidence.
  3. Investigation. Use the SIEM, the server's logs and EDR records to establish what happened, how it started, and what else was affected. Preserve evidence carefully if the incident may lead to legal action.
  4. Eradication and recovery. Remove the cause, restore affected systems from known-good backups or rebuild them, and reset credentials that may have been exposed.
  5. Lessons learned. Review what happened and improve defences, detection rules and the response plan itself.

Server administrators are often the first to see an alert or an unusual symptom. The key rules are to follow the plan, escalate to the security team rather than investigate alone, and record every action taken and when.

Practise what you just read

1. A single failed logon is harmless, but hundreds against many accounts followed by a success is an attack. Which SIEM capability detects this?

Select one

  1. Log compression of events held in storage
  2. Normalisation of every event into one common format
  3. Correlation of events across sources and time
  4. Data loss prevention rules applied to each server
Show answer

C. Correlation rules link related events, from many systems and over time, into a pattern that indicates an attack, which no single event reveals. Normalisation makes that comparison possible.

2. Why must every server's clock be synchronised with NTP for a SIEM to work well?

Select one

  1. NTP encrypts the log traffic that each server sends
  2. Events can only be ordered correctly if timestamps agree
  3. A SIEM uses NTP to find the servers it should collect from
  4. Accurate clocks make each log entry smaller to store
Show answer

B. If clocks differ, events that happened together appear minutes apart, or in the wrong order, and correlation and investigation reconstruct the wrong sequence of events.

3. A DLP rule for card numbers flags many order references. What tuning reduces the false positives while keeping real card numbers?

Select one

  1. Disable the rule, because it does not work
  2. Block every sixteen-digit number outright
  3. Lower the number of digits in the pattern
  4. Validate matches with the Luhn checksum
Show answer

D. Real card numbers pass the Luhn checksum, while most random digit strings do not. Adding the check removes false positives without missing genuine card numbers.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.