De-identify a dataset and measure what each method leaves behind
Task
Apply masking, tokenisation, transposition and de-identification to one synthetic dataset, then measure what survives each: whether any full card number is left, whether every token resolves only through the vault, whether the real salaries are still there, and how many people can still be singled out by postcode, birth date and sex. The lesson defines each method by what it leaves behind; this lab counts it.
Steps
- Generate
/tmp/deid/people.csvwith Python: 500 rows and the headerid,name,postcode,birth_date,sex,card,salary. Draw postcodes from about 40 five-digit values, birth dates between 1950 and 2005 asYYYY-MM-DD, and card numbers as 16 digits starting4471(synthetic, never real ones). - Masking: write
/tmp/deid/masked.csv, identical except thatcardshows only its last four digits, for example************1234. This is the display and test-data method. - Tokenisation: write
/tmp/deid/tokenised.csv, identical except thatcardholds a random token,tok_followed by 16 hex characters fromsecrets.token_hex. Keep the mapping in/tmp/deid/vault/vault.csvastoken,cardlines with no header, andchmod 600it. The vault is now the only place the real numbers live. - Transposition: write
/tmp/deid/transposed.csvwith thesalarycolumn shuffled across the rows and everything else unchanged. - De-identification: write
/tmp/deid/deid.csvwith the direct identifiers (name,card) removed and the indirect ones generalised:birth_datecut to its year andpostcodeto its first three digits. Keep the columnspostcode,birth_date,sex,salaryunder those names. - Write
/tmp/deid/findings.md: for each method, what it leaves behind -- part of a value, a vault to protect, the real values, or re-identification risk -- and which output you would hand to a developer, a support agent and an outside researcher.
Verify
python3 - <<'PY'
import csv, os, stat, collections
rd = lambda p: list(csv.DictReader(open(p)))
orig = rd('/tmp/deid/people.csv')
cards = {r['card'] for r in orig}
masked = open('/tmp/deid/masked.csv').read()
tok = open('/tmp/deid/tokenised.csv').read()
print('full card numbers left in masked :', sum(c in masked for c in cards))
print('full card numbers left in tokenised:', sum(c in tok for c in cards))
assert not any(c in masked for c in cards), 'masking left a full card number'
assert not any(c in tok for c in cards), 'tokenisation left a full card number'
vault = dict(csv.reader(open('/tmp/deid/vault/vault.csv')))
assert all(r['card'] in vault for r in rd('/tmp/deid/tokenised.csv')), 'a token the vault cannot resolve'
mode = stat.S_IMODE(os.stat('/tmp/deid/vault/vault.csv').st_mode)
print('vault permissions: %o' % mode)
assert mode & 0o077 == 0, 'the vault is readable by other accounts'
tr = rd('/tmp/deid/transposed.csv')
assert sorted(r['salary'] for r in tr) == sorted(r['salary'] for r in orig), 'transposition changed the values'
same = sum(a['salary'] == b['salary'] for a, b in zip(orig, tr)) / len(orig)
print('rows still holding their own salary: %.0f%%' % (same * 100))
assert same < 0.2, 'the salary column was not really shuffled'
def unique(rows):
key = lambda r: (r['postcode'], r['birth_date'], r['sex'])
c = collections.Counter(key(r) for r in rows)
return sum(1 for r in rows if c[key(r)] == 1)
before, after = unique(orig), unique(rd('/tmp/deid/deid.csv'))
print('people unique on postcode + birth date + sex: before', before, 'after', after)
assert after < before, 'generalising did not reduce re-identification risk'
PY
Every assertion must pass. Read the printed numbers as well as the verdict: transposition passes precisely because every real salary is still in the file, and the uniqueness count after de-identification is rarely zero. Those two lines are the residue the lesson says each method leaves behind.
Notes
With a full postcode, birth date and sex almost every synthetic person is unique, which is the well-known research result the lesson cites. Generalising cuts that count, but joining the output to another dataset can undo it, which is why pseudonymised data stays personal data and why anonymisation is a claim to test rather than assume.
This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.