Most data classification schemes are written to satisfy a control, adopted with a training slide, and ignored. Four tiers with names like Restricted and Internal, definitions that read like legal text, and no effect on anything anyone does on a Tuesday.
The failure is not that people are careless. It is that the scheme asks them to make a judgement and then does nothing with the answer, so the judgement has no consequence and stops being made.
This post is how to build one that changes behaviour. For whoever has been asked to produce a data classification policy.
Classify systems, not documents
Asking every person to label every file is the version that fails, because the number of decisions is enormous and each one is low stakes to the person making it.
Classify the places data lives instead. This database holds customer records. This bucket holds exports. This wiki space holds internal planning. This spreadsheet holds nothing sensitive. The count is small, the owners are identifiable, and the classification lasts.
Then the rules attach to the place rather than to the item, and the person who puts something somewhere inherits the handling requirements without having to decide anything.
Three tiers, not five
Every tier you add multiplies the boundary cases, and the arguments about whether something is Confidential or Restricted consume more effort than the distinction is worth.
Three works for most organisations:
- Public: safe if it appears on the internet tomorrow.
- Internal: business information, not harmful if it leaked, embarrassing if published.
- Sensitive: customer data, credentials, financials, anything regulated. Handling rules apply.
If you need a fourth, make it a small explicit category for a specific obligation, such as regulated health or payment data with its own legal requirements, rather than another shade of confidential.
Each tier must change what happens
This is the part that makes it real. A tier with no consequence is a label.
For each tier, state what is different: where it may be stored, who may access it by default, whether it may leave the country, whether it may go to a third party, retention, and what happens on deletion. If the answers are the same for two tiers, you have one tier with two names.
The rules should be things the organisation actually does. "Sensitive data may only be stored in the systems on this list" is enforceable and checkable. "Sensitive data must be handled with care" is not a rule.
Make the right thing the default
The scheme succeeds or fails on whether the compliant path is the easy one.
If sensitive data must live in a specific system, that system needs to be the convenient place to put it, with the access requests that people need being fast. If it is slow, people will use a spreadsheet, and the classification will be perfectly documented and entirely fictional.
The corollary: before writing rules that restrict where data may go, find out where it currently goes. A policy that forbids the workflow everybody depends on produces quiet non-compliance rather than change.
Wire it to the decisions it should inform
A classification is useful when other processes consume it:
- Access reviews prioritise systems holding sensitive data.
- Vendor review depth scales with the tier of data the vendor will touch.
- Backup and retention policies differ by tier.
- Incident severity is partly determined by the classification of what was involved.
- Deletion requests know which systems to search, because the inventory of sensitive systems already exists.
If none of your processes reference the classification, it has no consumers, and something with no consumers stops being maintained.
Keep it current with the one thing nobody forgets
Classification drifts as systems are added. The same trick that works for subprocessors and contractors works here: attach the review to something with its own momentum.
New systems get classified as part of being provisioned, ideally as a required field rather than a follow-up. Existing ones get reviewed on the same cycle as your access review, using the same list. Two processes, one list, and the list stays alive because it is used twice.
What an auditor will ask
Expect: do you have a classification scheme, is it approved, does it define handling requirements per tier, do you maintain an inventory of systems and their classification, and can you show that handling matches the policy for a sample.
The last one is where schemes fail, because the policy says sensitive data is encrypted, restricted and retained for two years, and the sample shows a spreadsheet in a shared drive. Being able to answer honestly, including "this system is classified sensitive and here is the exception we recorded and why", is better than a policy that describes a company you are not.
The concession
There is a real argument that for a small company this is overhead with no benefit: everybody knows what the sensitive systems are, and writing it down changes nothing. That is often true at fifteen people.
It stops being true at the point where somebody joins who does not have the shared context, or where a customer asks, or where you need to answer which systems hold their data. The trigger is not size, it is the first time the answer lives only in one person's head. At that point, a one-page list of systems with three tiers and a handling table is a morning's work and it is the artefact everything else references.
The implication
The purpose is not to have a policy. It is that when someone asks where customer data lives, or which vendors touch it, or what happens when it must be deleted, the answer exists and is the same answer every time.
If your scheme cannot answer those three questions, the tiers are decoration regardless of how carefully they are defined.