De-identify a dataset and measure what each method leaves behind

applied · 75 min · Objective 3.3

Task

Apply masking, tokenisation, transposition and de-identification to one synthetic dataset, then measure what survives each: whether any full card number is left, whether every token resolves only through the vault, whether the real salaries are still there, and how many people can still be singled out by postcode, birth date and sex. The lesson defines each method by what it leaves behind; this lab counts it.

Steps

  1. Generate /tmp/deid/people.csv with Python: 500 rows and the header id,name,postcode,birth_date,sex,card,salary. Draw postcodes from about 40 five-digit values, birth dates between 1950 and 2005 as YYYY-MM-DD, and card numbers as 16 digits starting 4471 (synthetic, never real ones).
  2. Masking: write /tmp/deid/masked.csv, identical except that card shows only its last four digits, for example ************1234. This is the display and test-data method.
  3. Tokenisation: write /tmp/deid/tokenised.csv, identical except that card holds a random token, tok_ followed by 16 hex characters from secrets.token_hex. Keep the mapping in /tmp/deid/vault/vault.csv as token,card lines with no header, and chmod 600 it. The vault is now the only place the real numbers live.
  4. Transposition: write /tmp/deid/transposed.csv with the salary column shuffled across the rows and everything else unchanged.
  5. De-identification: write /tmp/deid/deid.csv with the direct identifiers (name, card) removed and the indirect ones generalised: birth_date cut to its year and postcode to its first three digits. Keep the columns postcode,birth_date,sex,salary under those names.
  6. Write /tmp/deid/findings.md: for each method, what it leaves behind -- part of a value, a vault to protect, the real values, or re-identification risk -- and which output you would hand to a developer, a support agent and an outside researcher.

Verify

python3 - <<'PY'
import csv, os, stat, collections
rd = lambda p: list(csv.DictReader(open(p)))
orig = rd('/tmp/deid/people.csv')
cards = {r['card'] for r in orig}
masked = open('/tmp/deid/masked.csv').read()
tok = open('/tmp/deid/tokenised.csv').read()
print('full card numbers left in masked   :', sum(c in masked for c in cards))
print('full card numbers left in tokenised:', sum(c in tok for c in cards))
assert not any(c in masked for c in cards), 'masking left a full card number'
assert not any(c in tok for c in cards), 'tokenisation left a full card number'
vault = dict(csv.reader(open('/tmp/deid/vault/vault.csv')))
assert all(r['card'] in vault for r in rd('/tmp/deid/tokenised.csv')), 'a token the vault cannot resolve'
mode = stat.S_IMODE(os.stat('/tmp/deid/vault/vault.csv').st_mode)
print('vault permissions: %o' % mode)
assert mode & 0o077 == 0, 'the vault is readable by other accounts'
tr = rd('/tmp/deid/transposed.csv')
assert sorted(r['salary'] for r in tr) == sorted(r['salary'] for r in orig), 'transposition changed the values'
same = sum(a['salary'] == b['salary'] for a, b in zip(orig, tr)) / len(orig)
print('rows still holding their own salary: %.0f%%' % (same * 100))
assert same < 0.2, 'the salary column was not really shuffled'
def unique(rows):
    key = lambda r: (r['postcode'], r['birth_date'], r['sex'])
    c = collections.Counter(key(r) for r in rows)
    return sum(1 for r in rows if c[key(r)] == 1)
before, after = unique(orig), unique(rd('/tmp/deid/deid.csv'))
print('people unique on postcode + birth date + sex: before', before, 'after', after)
assert after < before, 'generalising did not reduce re-identification risk'
PY

Every assertion must pass. Read the printed numbers as well as the verdict: transposition passes precisely because every real salary is still in the file, and the uniqueness count after de-identification is rarely zero. Those two lines are the residue the lesson says each method leaves behind.

Notes

With a full postcode, birth date and sex almost every synthetic person is unique, which is the well-known research result the lesson cites. Generalising cuts that count, but joining the output to another dataset can undo it, which is why pseudonymised data stays personal data and why anonymisation is a claim to test rather than assume.

This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.