What is de-identification, and does it actually protect my customers?
In this article
What gets removedHow it is doneWhat remains, and why it is still valuableThe limitsWhat you approve, and whenWhat gets removed
Direct identifiers: names, email addresses, phone numbers, postal addresses, account numbers, card numbers, government IDs, IP addresses, usernames. Indirect identifiers that combine into a person: a job title plus a small company plus a date. Health information, which is excluded entirely rather than scrubbed. And anything you have placed under a restricted category, which is removed or excluded whether or not it identifies anyone.
How it is done
Structured fields are dropped or replaced with placeholders. Free text is run through detection that finds names, numbers and addresses in context and replaces them with tokens, then sampled by a human reviewer. Attachments are screened the same way or excluded. The buyer does this before anyone reviews the records, and the process is written into the agreement you sign. What every buyer commits to.

See what your company's records could get.
Ten questions, under five minutes, built from the buyers' own published ranges.
Get my estimate →What remains, and why it is still valuable
The shape of the work. "Customer reports the unit short-cycles; technician finds a failed capacitor; replaces; invoice $412; follow-up call three days later, resolved." Who the customer was adds nothing to that record for a buyer training a model on how HVAC work gets done. The process, the timing, the pricing and the outcome are the asset.
The limits
De-identification reduces risk; it does not make re-identification impossible in every case, especially for records about very small groups or very unusual events. That is why buyers also sign use restrictions, why some categories are excluded rather than scrubbed, and why you approve the categories before anything is scoped. If a category feels wrong, it comes out.
What you approve, and when
Before anything is scoped: the list of systems and the categories of record inside them. Before anything is exported: the scrub rules for each category. Before anything is used: nothing further, because the agreement covers use, retention and deletion. The four things that make it legal.
In short
- De-identification removes names, contact details, identifiers and anything that points to a person, before review.
- Health information is excluded entirely. Restricted categories are removed whether or not they identify anyone.
- What remains is the work: process, timing, pricing, outcome. That is the asset.
- It reduces risk rather than eliminating it, which is why use restrictions and your approval sit on top.
Questions owners ask.
Can I see the scrubbed output before it is used?
Yes. Buyers provide a sample after scrubbing and before use, and you can pull a category at that point.
What about my employees' names?
Removed the same way. Internal threads keep the roles ("ops lead", "technician") and lose the names.
Is de-identification the same as anonymization?
Close. Anonymization is the stronger claim that re-identification is practically impossible; de-identification is the process. Buyers describe what they do precisely in the agreement.
Briggs Analytics