Data types, states, classification and the methods that protect them

This course teaches SY0-801, the Security+ exam that launches on or around 17 November 2026. If you are booked on SY0-701, which can be taken until 11 June 2027, use our SY0-701 course instead.

Objective 3.3 · Security Architecture · 19% of the exam

Objective 3.3 asks you to summarize how data is protected. It is a broad objective, so this course gives it two lessons. This one is about the data itself: what shape it takes, how sensitive it is, what state it is in, and the techniques that protect it. The next lesson takes the people, places, life cycle and legal categories around it.

Why this matters

Every control in this course costs something. Classification is how you decide what to spend it on: you cannot protect everything at the highest level, and an organisation that tries either runs out of money or quietly stops doing it.

The exam tests this objective largely as matching. Given a state -- at rest, in transit, in use -- which control applies? Given a requirement, is the answer masking, tokenisation, hashing or encryption? Given a dataset shared for analysis, what does de-identification leave behind? Those are separate question shapes, and each has a clean answer once the distinctions are sharp.

The lesson

Structured versus unstructured data, and why the second is harder to protect

Structured data lives in a defined format: database tables, spreadsheets with consistent columns, records with named fields. Because you know where every value is, you can protect it precisely -- encrypt one column, mask one field in a report, grant a role access to some columns and not others, and audit every query against it.

Unstructured data has no fixed format: documents, email, chat messages, presentations, images, scanned forms, audio and video recordings. Most of an organisation's data is unstructured, and it is harder to protect for reasons the exam likes:

  • You do not know where the sensitive parts are. A card number can sit in the middle of an email or a photo of a form.
  • It copies itself. An attachment forwarded five times exists in five mailboxes, two laptops and a sync folder.
  • Access control is coarse -- usually per file or folder, not per value.
  • Content inspection is needed to find it, which is what discovery tools and data loss prevention do, imperfectly.

Data in between, such as logs, JSON or XML, is sometimes called semi-structured: it has labelled fields but no fixed schema. The practical lesson: the most reliable way to protect unstructured data is to classify and label it when it is created, because finding it afterwards is slow and never complete.

Classification levels from public to top secret, and the protection each earns

Classification is the label you attach, and it decides handling. Two families of scheme appear on the exam.

Government schemes classify by the damage unauthorised disclosure would do. The United States, for example, uses three levels: confidential, secret and top secret, in ascending order of harm, each with its own clearance, storage and transmission rules.

Commercial schemes use their own names. A common set, from least to most protected:

  • Public -- approved for release; integrity still matters, secrecy does not.
  • Sensitive or internal -- not for outsiders, limited harm if leaked.
  • Confidential -- significant harm if disclosed; access on a need-to-know basis.
  • Restricted -- the highest commercial tier, for the most damaging data, with named access and the strongest controls.

Critical is a different axis: data the organisation cannot operate without. It is about availability and integrity as much as secrecy, so it earns backups, replication and recovery testing rather than only access control.

Exact names vary between organisations, and the exam does not require one taxonomy. It requires the mechanics: every classified item has an owner; the level determines who may access it, how it is stored, transmitted, retained and destroyed; and labels make the level visible so that people and tools can act on it.

Two failures the exam tests. Over-classification -- labelling everything confidential -- makes the label meaningless and trains people to ignore it. And aggregation: a name is public and a job title is public, but a complete staff list with titles, locations and email formats is a phishing target.

Data at rest, in transit and in use, and the control for each state

The cleanest mapping in the objective, examined directly.

  • At rest -- stored on disk, in a database, in object storage, on backup media. Controls: encryption (full-disk, file, database, field), access permissions, physical security.
  • In transit -- crossing a network. Controls: TLS, IPSec, SSH and other transport encryption, plus certificate validation so you are encrypting to the right party.
  • In use -- loaded in memory and being processed. The hard one, because the application must see plaintext to work on it. Partial answers: process isolation, secure enclaves and confidential computing, tokenisation so the sensitive value is never processed at all, and tight control over who can reach the running system.

The exam point: if a scenario asks how to protect data while it is being processed, encryption at rest and encryption in transit are both wrong answers.

Masking, tokenisation, hashing, encryption and obfuscation, compared

Method Reversible? Use it when
Encryption Yes, with the key You need the original back, and only key holders should get it
Hashing No You need to verify integrity or store passwords (salted, stretched)
Masking Usually no People need to see part of a value, or test data must look real
Tokenisation Only through the token vault The real value should never enter most systems
Obfuscation Often, by anyone who works it out Making casual reading harder, never alone
  • Encryption protects data at rest and in transit; its weak point is key management.
  • Hashing is one-way. It is not for data you need to read again.
  • Masking hides part of a value for display -- the last four digits of a card number on a screen -- and is the usual way to make realistic non-production data.
  • Tokenisation swaps the value for a meaningless token and keeps the mapping in a separate, tightly protected vault. There is no key and no mathematical relationship, so a stolen token is worthless. This is the card-payment pattern, and it shrinks the number of systems in compliance scope.
  • Obfuscation makes data or code harder to interpret. It slows a reader down and does not stop a determined one.

Underneath all of them sits data minimisation: data you never collected cannot be breached and costs nothing to protect.

Filtering, transposition and de-identification, and what each leaves behind

Three more methods, each defined by what survives it.

  • Filtering removes or withholds data that the destination does not need: exporting three columns instead of thirty, stripping fields from an API response, or blocking outbound content that matches a sensitive pattern. What it leaves behind is everything that passed the filter, so the filter is only as good as its rules -- a new field added upstream flows straight through a filter written as "remove these".
  • Data transposition rearranges data rather than replacing it: swapping values between records, or reordering characters within a value, so that the content is present but the arrangement that gave it meaning is broken. Shuffling a column of salaries across employees keeps a realistic dataset for testing without attaching any salary to the right person. What it leaves behind is the real values themselves, and their patterns and statistics, so on its own it is weak protection.
  • De-identification removes or alters the details that link data to a person. Direct identifiers such as names and account numbers are removed or replaced; indirect identifiers such as birth date, postcode and job title are generalised. If a separately held key can reverse it, that is pseudonymisation, and under privacy laws such as the GDPR pseudonymised data is still personal data. If it cannot be reversed by anyone, that is anonymisation. What de-identification leaves behind is re-identification risk: well-known research has shown that a few indirect details, such as postcode, birth date and sex, are enough to single out most people, especially when the dataset is combined with another one.

What to take into the exam

  • Structured data can be protected field by field; unstructured data must be found first, so classify it at creation.
  • Government schemes run confidential, secret, top secret; commercial schemes run roughly public to restricted; critical is about availability.
  • At rest: encryption and permissions. In transit: TLS or IPSec. In use: enclaves, tokenisation, access control.
  • Tokenisation has no key; masking is for display and test data; hashing is one-way; obfuscation is never enough alone.
  • Filtering leaves what passed the rules, transposition leaves the real values, and de-identification leaves re-identification risk.

Practise what you just read

1. Which controls protect data while an application is actively processing it?

Select one

  1. Full-disk encryption plus TLS on every connection made
  2. Database encryption together with encrypted backup media
  3. Secure enclaves, tokenisation and tight access control
  4. Certificate validation on every client that connects to it
Show answer

C. Data in use is the hard state because the application must see plaintext to work on it. Encryption at rest and in transit are both wrong answers here; enclaves and confidential computing, tokenisation so the real value is never processed, and control over who can reach the running system are the partial answers.

2. Which protection method has no key and no mathematical link to the original value?

Select one

  1. Salted hashing
  2. Encryption
  3. Field encryption
  4. Tokenisation
Show answer

D. A token is a meaningless substitute, and the mapping lives in a separate, tightly protected vault. A stolen token database has nothing to decrypt, which is why tokenisation is the card-payment pattern that shrinks compliance scope, and why the vault is the asset to defend.

3. Names in a dataset are replaced with codes, and a separately held key can reverse them. What is this, and is it still personal data?

Select one

  1. Anonymisation, so it is no longer personal data
  2. Tokenisation, so it falls outside privacy law
  3. Pseudonymisation, and it is still personal data
  4. Transposition, which keeps values but hides links
Show answer

C. If a separately held key can reverse the de-identification, it is pseudonymisation, and under laws such as the GDPR pseudonymised data remains personal data. Only de-identification that nobody can reverse is anonymisation, and even then indirect identifiers can leave re-identification risk.

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.