Data retention and the data lifecycle

Listen to this lesson

Episode 27 · 36:57

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 3.1 · Security and disaster recovery · 24% of the exam

Why this matters

It is tempting to keep everything for ever. Storage is cheap, and deleting something that turns out to be needed is embarrassing. But keeping data has costs and risks of its own: every record held is a record that can be stolen, that must be searched when a legal request arrives, and that some laws say must not be kept longer than necessary. Deleting too early has its own consequences, from failed audits to penalties for destroying evidence.

Retention is where security, law and server administration meet. This lesson covers why data must be kept for set periods, how data is classified, the stages it passes through, the one situation that overrides every schedule, and how backups fit into all of it.

The lesson

Retention policies, and the legal and regulatory reasons behind them

A retention policy states how long each type of data must be kept, and what happens to it afterwards. Retention periods come from several sources:

  • Legal and regulatory requirements. Laws require many records to be kept for minimum periods: financial and tax records, employment records, and records in regulated industries such as healthcare and finance. The periods vary by country and sector, which is why retention policies are written with legal and compliance staff rather than by IT alone.
  • Privacy law. Data protection laws such as the EU's GDPR work in the other direction: personal data should be kept no longer than necessary for the purpose it was collected for. Keeping it indefinitely can itself be a breach of the law.
  • Contracts and business needs. Agreements with customers or partners may set retention periods, and the organisation may simply need some data for a time to operate.

The policy therefore sets both a minimum, the time data must be kept, and often a maximum, after which it should be destroyed. The server administrator's job is to make systems enforce it: storage, archiving, and deletion when the period ends, rather than leaving the policy as a document nobody implements.

Data classification: public, internal, confidential and restricted

Not all data needs the same protection. Data classification sorts it into levels according to the harm its exposure would cause, so protection can be matched to the risk. A common four-level scheme:

  • Public: intended for anyone, such as published marketing material and product information. Exposure causes no harm.
  • Internal: for staff, but not secret, such as internal procedures and directories. Exposure would be undesirable but not seriously damaging.
  • Confidential: sensitive business or personal data, such as customer records, contracts and employee information. Exposure would cause real harm.
  • Restricted: the most sensitive data, such as payment card data, health records, trade secrets and credentials. Exposure could cause severe harm or legal penalties, so access is limited to named people.

Names vary between organisations, and governments use their own schemes, but the principle is the same. Classification drives everything else: which data is encrypted, who may access it, where it may be stored, how long it is kept, and how it is destroyed. Data should carry its classification, through labels, metadata or its storage location, so that systems and people can treat it correctly.

The lifecycle from creation to destruction

Data passes through a series of stages, and each has its own controls.

  1. Create or collect: data is generated or received, and should be classified at this point.
  2. Store: data is kept on servers or in services, protected according to its classification with access controls and encryption.
  3. Use: data is read, processed and changed by people and applications, with access limited to those who need it.
  4. Share: data is sent to others, inside or outside the organisation, over encrypted channels and only where permitted.
  5. Archive: data no longer in active use but still within its retention period moves to cheaper, slower storage, still protected and still searchable if it must be produced.
  6. Destroy: at the end of the retention period, data is deleted or its media destroyed so it cannot be recovered, and the destruction is recorded.

The last stage is the one most often skipped. Data that should have been destroyed years ago lingers on old shares, forgotten servers and archived backups, where it adds risk and nothing else. The methods for destroying data properly are covered in the media sanitisation lesson later in this domain.

Legal holds, and why they override the schedule

A legal hold, or litigation hold, is an instruction to preserve data because it may be relevant to a lawsuit, investigation or regulatory inquiry. It is usually issued by the legal department when litigation is expected or begins.

A legal hold overrides the retention schedule. Data covered by the hold must not be deleted, altered or destroyed, even if its retention period has ended and even if the normal policy would delete it automatically. Destroying data under a hold, even by an automated cleanup job running as usual, can lead to severe penalties and damage the organisation's case.

For server administrators, this means:

  • automated deletion, cleanup scripts, mailbox purges and backup rotation must be able to exempt data under hold, and should be checked when a hold is issued;
  • many systems support holds directly, such as the litigation hold features in email and document platforms, which preserve items even when users delete them;
  • a hold stays in force until legal staff formally release it, after which normal retention resumes.

A hold is also the clearest reason not to keep excessive data: everything the organisation holds when a hold arrives may have to be searched and produced.

Keeping backups in line with the retention policy

Backups are copies of data, so they are subject to the same retention policy, and they are where policies are most often undermined.

Two problems are common. The first is backups kept too long: if customer data is deleted from live systems after seven years but backups are kept for ten, the organisation still holds it, and a legal request or breach can reach it. The second is backups kept too briefly: if the policy requires records for seven years but backups rotate after a month, any record deleted from the live system is lost for good well before its time.

Backup retention should therefore be set deliberately to match the policy. Rotation schemes, covered in the backup lessons later in the course, keep daily, weekly, monthly and yearly copies for different periods; archives, kept separately for long-term retention, serve the longest requirements. Individual records inside a backup usually cannot be deleted selectively, so the whole backup's expiry date has to respect both the longest period that applies to what it contains and any legal hold that covers it.

Try it

An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.

Practise what you just read

1. A legal hold covers records that have passed their retention period. What must happen to them?

Select one

  1. They should be moved to cloud storage for safekeeping
  2. They should be encrypted, then deleted on the schedule
  3. They should be deleted, as retention overrides holds
  4. They should be preserved until the hold is released
Show answer

D. A legal hold overrides the retention schedule. Deleting held data, even through an automated cleanup, can lead to severe penalties, so deletion jobs must be able to exempt it.

2. Why might keeping personal data longer than necessary break the law?

Select one

  1. Old personal data cannot be encrypted under the current standards
  2. Privacy laws such as GDPR require it to be kept no longer than needed
  3. Storage of personal data is taxed per gigabyte per year in most countries
  4. It does not, because storing data collected lawfully is permitted
Show answer

B. Data protection laws require personal data to be kept only as long as its purpose requires. Retention policies therefore set maximum periods as well as minimum ones.

3. Backups are kept for ten years, but the policy says customer data must be deleted after seven. What is the problem?

Select one

  1. No problem, as backup copies are exempt from retention policy
  2. The old backups will fail to restore after seven years
  3. The organisation still holds data it should have deleted
  4. The backups use too little storage to hold ten years of data
Show answer

C. Backups are copies of data and fall under the same policy. Keeping them longer than allowed means the data still exists and can be reached by a breach or a legal request.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.