Q1.06.2De-identification and retention limitationdesignresearch

Research data need de-identification and bounded retention

Aliases: de-identification · retention schedule · data minimization

What it is

De-identification and retention limitation remove direct identifiers no longer needed, reduce re-identification through indirect combinations, and delete or further aggregate data when its purpose ends. De-identified is not anonymous: codes, precise times, locations, rare occupations, voices, and screen content may still reconnect a record to a person. Retention follows research and audit needs rather than defaulting to permanence.

Why it happens

Re-identification emerges from combinations, not only names. Time, trajectory, and small-group attributes can become unique when linked. Longer retention increases the window for breach, purpose drift, accumulated access, and changes in participant circumstances. Separating identity keys, reducing granularity, limiting roles, and enforcing expiry reduce both external attack and internal misuse. Replacing a name with an ID alone preserves linkability.

Studying it

Maintain a data inventory and flow map with field purpose, source, sensitivity, linkability, access role, and deletion date. Threat modeling should consider direct, inferred, linkage, and membership disclosure across exports, backups, transcription, and collaboration services. Audit whether retained material still has an active purpose and record deletion evidence and backup expiry. Appropriate anonymization depends on data type, attacker capability, and applicable requirements.

Where it stops holding

Long retention increases risk, but the object of deletion matters. Destroying the identity key while retaining de-identified records may later make a participant’s record impossible to locate for targeted removal; that is different from deleting the research data and is not a reason to retain everything longer. Each data class should remain only as long as research necessity, the approved protocol, and applicable legal or ethical duties justify. Small qualitative datasets remain identifiable through quotations and context after names are removed, and public datasets, model parameters, and aggregates are not automatically disclosure-free.

Applying it

  • Ask whether each field is necessary to answer the question before collection.
  • Store contact and consent records separately from research data under a controlled mapping.
  • Assign owners and expiry dates to raw, transcribed, analyzed, and backup material.
  • Rehearse export and deletion with least-privilege accounts, including local copies, version history, trash, and provider backups.

Related

  • Same group: Q1.06.1 Explain purpose and withdrawal · Q1.06.3 Vulnerable populations need additional safeguards
  • Adjacent: Q1.13 Informed consent and research ethics · Q3.11 Telemetry and instrumentation
  • Search terms: de-identification · data retention · re-identification risk

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q1.06.2