Data Classification Policy Levels, Labels, and Rules That Hold Up
A data classification policy tells everyone in your organization which data is sensitive, how sensitive it is, and exactly what they are allowed to do with it. This guide covers the standard classification levels, what belongs in the written policy, and how Petronella Technology Group builds programs that survive a real audit.
Key Takeaways
- Data classification is the practice of sorting information into named sensitivity levels so that access, storage, and disposal rules can be applied consistently.
- Most organizations land on four data classification levels: Public, Internal, Confidential, and Restricted. Three is often enough for a small business.
- Classification is based on the harm that disclosure would cause, not on who created the file or which department owns it.
- A policy without a handling matrix is decoration. Each level needs written rules for storage, transmission, sharing, retention, and destruction.
- Classification only works on data you know you have, so it starts with an IT asset inventory and a data discovery pass.
Before You Write a Single Level
The most common failure we see is a beautiful four-level policy written in a conference room by people who have never looked at the file server. Somebody downloads a template, swaps in the company name, and files it for the auditor. Six months later nothing is labeled, nobody can tell you where the customer records live, and the policy is evidence of a control that does not exist. Start with discovery. Find the data first, then name the levels around what you actually found.
What Is Data Classification?
The plain answer, before the detail.
Data classification is the process of sorting an organization's information into defined sensitivity levels so that consistent security rules can be applied to each level. Instead of making a judgment call about every individual file, a person consults the classification label and follows the handling rules already written for that label. A data classification policy is the document that names those levels, defines what belongs in each one, and states the rules.
The practical purpose is decision-making at scale. An organization might hold millions of files across email, network shares, cloud storage, and line-of-business applications. No security team can review each one. Classification converts an unmanageable per-file question into a small, repeatable set of rules: this is Confidential, so it is encrypted at rest, restricted to the finance group, never sent to a personal email address, and destroyed after seven years.
What Data Classification Is Based On
Classification is based on impact: the harm that would result if the information were disclosed, altered, or lost. That harm can be legal, financial, competitive, or reputational. It is not based on file type, on the department that produced the document, or on how recently it was touched. A spreadsheet is not Confidential because finance made it. It is Confidential because it contains information that would damage the business or trigger a regulatory obligation if it leaked.
Three factors drive the impact judgment in practice. First, regulatory obligation: protected health information, controlled unclassified information, cardholder data, and personal data carry mandatory handling requirements regardless of how your business feels about them. Second, contractual obligation: a customer agreement or a defense subcontract can impose handling terms stricter than any regulation. Third, business sensitivity: pricing models, source code, acquisition plans, and unreleased product designs may face no legal requirement at all and still deserve the tightest controls you have.
Data Classification Versus Data Categorization
The two words get used interchangeably, and the distinction matters when an auditor is reading your policy. Classification assigns a sensitivity level: how protected does this need to be. Categorization assigns a content type: what kind of information is this. A record can be categorized as "customer contact data" and classified as "Confidential." Good programs do both, because the category tells you which regulation applies and the classification tells you which controls apply. If your policy uses one word for both ideas, define your terms explicitly in the opening section.
The Four Data Classification Levels
Names vary between organizations. The structure almost never does.
1. Public
Information approved for release outside the organization. Marketing pages, published press releases, job postings, product documentation, and regulatory filings that are already in the public record. The security objective for Public data is integrity rather than confidentiality: nobody is harmed by reading it, but real damage follows if an attacker alters it. Defacement of a public site or tampering with published pricing is a Public-data incident, and it is why Public still belongs in the policy instead of being left unlabeled.
2. Internal
Information intended for employees and contractors but not for the general public. Internal process documents, org charts, meeting notes, most routine email, and internal wikis. Disclosure would be embarrassing or mildly useful to a competitor without triggering a legal obligation. This is the default level in most policies, which is the right design choice: unlabeled data should fall into Internal automatically, not into Public.
3. Confidential
Information that would cause meaningful harm if disclosed. Customer records, employee files, financial statements before release, contracts, pricing models, and vendor agreements. Confidential data typically requires encryption at rest and in transit, access limited to a named business need, logging of access, and an approval step before it leaves the organization. Most regulated data that is not subject to a specific federal framework lands here.
4. Restricted
The smallest and most tightly held tier. Information whose disclosure would cause severe harm, trigger mandatory breach notification, or breach a federal requirement. Protected health information, controlled unclassified information, cardholder data, Social Security numbers, authentication secrets, and encryption keys. Restricted data usually carries additional requirements: multi-factor authentication, dedicated storage locations, prohibition on removable media, formal approval for every external transfer, and defined retention and destruction schedules.
How Many Levels Should You Use?
Four is the common answer, and for a business with fewer than roughly fifty employees, three is often better. Every additional level multiplies the training burden and the number of edge cases people have to reason about. If your staff cannot state the levels from memory, you have too many. We have seen seven-level schemes at organizations that could not correctly classify a single sample document, and three-level schemes at organizations where every employee got it right. Fewer levels applied accurately beats more levels applied inconsistently.
The one situation that justifies a fifth level is a business holding two distinct categories of severely regulated data with genuinely different handling rules, such as a company that handles both controlled unclassified information for a defense contract and protected health information for a clinical line. Even then, most organizations do better with four levels plus category tags than with five levels.
What Each Level Actually Requires
The handling matrix is the part auditors read. Levels without rules prove nothing.
| Control | Public | Internal | Confidential | Restricted |
|---|---|---|---|---|
| Access | Anyone | All employees and contractors | Named business need, reviewed quarterly | Named business need plus documented approval |
| Encryption at rest | Not required | Recommended | Required | Required, with managed keys |
| Encryption in transit | Recommended | Required | Required | Required, approved channels only |
| External sharing | Permitted | Manager approval | Written agreement in place | Formal approval plus transfer log |
| Removable media | Permitted | Permitted, encrypted | Encrypted, registered devices | Prohibited |
| Access logging | Not required | Recommended | Required | Required and monitored |
| Retention | Business need | Business need | Defined schedule | Defined schedule, regulator-driven |
| Destruction | Standard deletion | Standard deletion | Secure wipe, logged | Certified destruction with certificate |
Adapt the specifics to your environment and your regulators, but keep the shape. Every row is a question an assessor will eventually ask, and "we handle that case by case" is not an answer that produces evidence. Pair the matrix with technical enforcement where you can: data loss prevention rules that block Restricted content from leaving approved channels turn a written rule into a control that actually stops something.
"Petronella Cybersecurity provides outstanding service! Their team is extremely knowledgeable, responsive, and truly cares about protecting their clients. They take the time to explain complex issues in simple terms and deliver real solutions, not just promises."
GB Entrainement, verified TrustIndex reviewWhat Belongs in the Written Policy
Nine sections cover what every framework we work with expects to see.
Purpose and scope. State why the policy exists and exactly what it covers: which systems, which business units, which data, and whether contractors and third parties fall inside the boundary. Scope gaps are where assessments fail. A policy that silently excludes the marketing team's cloud storage has an undocumented exception, and undocumented exceptions become findings.
Definitions. Define classification, categorization, data owner, data custodian, and each level by name. Assume the reader is a new hire in accounting, not a security engineer.
The classification levels. Name each level, describe the harm threshold that puts data into it, and give three or four concrete examples drawn from your actual business. Generic examples produce generic compliance. "Customer payment records in the billing system" beats "sensitive financial information."
Roles and responsibilities. Name who classifies data at creation, who reviews classifications, who approves declassification, and who owns the policy. Data owners are business leaders, not IT staff. IT is the custodian that enforces the rules the owner sets, and blurring that line is how classification decisions end up being made by whoever last touched the file server.
The handling matrix. The table above, adapted to your controls and your systems.
Labeling requirements. State how labels are applied: document headers and footers, file metadata, email subject tags, cloud platform sensitivity labels, or physical markings on printed material. State what happens to unlabeled data, which should default to Internal or higher, never Public.
Reclassification and declassification. Sensitivity changes over time. Quarterly earnings are Restricted before release and Public after. Write the process for changing a label and who may approve it, or you will have people quietly downgrading data to make sharing easier.
Exceptions. There will be exceptions. Define who approves them, how they are documented, and when they expire. An exception process is a strength in an assessment; an undocumented workaround is a finding.
Enforcement and review. State the consequence of noncompliance and set a review cadence. Annual review is the norm, with an out-of-cycle review triggered by a merger, a new regulated data type, or a significant incident.
How to Build the Program in Six Steps
The order matters. Discovery before levels, levels before labels, labels before enforcement.
Inventory the Systems
You cannot classify data in systems you have forgotten. Start from a current IT asset inventory covering servers, endpoints, cloud tenants, software as a service applications, backups, and any place a departing employee might have left a copy. Shadow cloud storage is the usual surprise.
Run Data Discovery
Scan those systems for the patterns that matter: Social Security numbers, payment card numbers, health record identifiers, and defense contract markings. Data discovery and classification tooling does the first pass; a human confirms the findings. Expect to find regulated data in places nobody expected, usually old shared drives and personal mailboxes.
Define Levels Around What You Found
Now write the levels, informed by real data rather than a template. If the discovery pass turned up no controlled unclassified information, you do not need a level built for it. If it turned up three distinct regulated types, your handling matrix needs to say something specific about each.
Assign Owners and Label
Every repository gets a named business owner who is accountable for the classification of what it holds. Then label, working highest-risk first. Labeling everything at once fails; labeling the finance share and the clinical records this quarter succeeds.
Enforce Technically
Translate the matrix into configuration: access groups, encryption settings, sensitivity labels in your cloud platform, and data loss prevention rules that block Restricted content from unapproved destinations. Route the resulting alerts somewhere staffed, such as a managed detection and response service.
Train, Measure, and Review
Teach people the levels in the context of their own work, not in the abstract. Measure how often new documents get labeled correctly, and review the policy annually. Pair the rollout with security awareness training so classification is part of how people think rather than a slide they clicked past.
Classifying PII, PHI, CUI, and Cardholder Data
Where a general policy meets specific federal and contractual obligations.
PII Data Classification
Personally identifiable information is any data that can identify a specific person, alone or combined with other data. Names and business email addresses are usually low-sensitivity. Social Security numbers, driver license numbers, financial account numbers, and biometric data belong in Restricted in nearly every policy we write. The trap is combination: a birth date is not identifying by itself, and a birth date next to a ZIP code and a gender field very often is. Classify the record as it actually exists in the database, not field by field in isolation.
PHI and HIPAA
Protected health information carries specific handling requirements under the HIPAA Security Rule, and the eighteen identifiers that make health data protected are broader than most practices assume. Appointment schedules, billing records, and even a voicemail transcript can qualify. As Craig Petronella details in How HIPAA Can Crush Your Medical Practice, the failures that generate penalties are rarely exotic attacks; they are ordinary mishandling of records nobody had classified. Healthcare organizations should read the classification policy alongside their HIPAA compliance program rather than as a separate exercise.
CUI and CMMC
Defense contractors face the strictest version of this problem. Controlled unclassified information must be identified, marked, and protected under NIST SP 800-171, and CMMC assessment turns on whether you can show where it lives. Classification is the practical foundation of scoping: the boundary of your assessment is defined by where controlled unclassified information is processed, stored, or transmitted, so a wrong classification produces a wrong boundary and a wrong score. Craig Petronella, a CMMC Registered Practitioner and author of the CMMC 2.0 Certification Guide, works through this scoping logic in detail, and our CMMC compliance guide covers how the levels map to the 110 controls. The free SPRS score calculator is a reasonable place to check where you stand.
Cardholder Data
Payment card data has the least ambiguous rules of the four. The primary account number is Restricted, full magnetic stripe data and card verification codes may not be stored after authorization at all, and the systems that touch them define your PCI DSS scope. Classification here is mostly a scoping exercise: find every system that stores, processes, or transmits cardholder data, and shrink that list as far as the business allows.
GDPR and Personal Data
European personal data is defined more broadly than United States personally identifiable information, and it includes special categories such as health, biometric, and political data that require heightened protection. If you serve European customers, your levels need to accommodate that broader definition rather than assuming a United States privacy standard covers it. The same reasoning applies to state privacy laws, which increasingly define personal data closer to the European standard than to the older domestic one.
Frameworks such as SOC 2 do not dictate levels, but they do expect you to have defined them and to show the controls following from them. That is the underlying pattern across every framework: none of them tells you what your levels must be, and all of them ask what your levels are.
Why Most Classification Policies Fail
Five patterns we see repeatedly in organizations that already have a policy on file.
Writing the Policy Before the Discovery
A template downloaded and renamed is not a policy, it is a description of somebody else's business. The levels have to be defined around data that actually exists in your environment, which means discovery comes first. This single sequencing error accounts for more dead classification programs than every technical problem combined.
Classifying Everything as Confidential
When people are unsure, they overclassify, and it feels responsible. It is not. If most data is Confidential, the label carries no information, the handling rules become impractical, and staff start ignoring them selectively. That selective ignoring is worse than no policy, because now your genuinely Restricted data is protected by a rule everyone has already learned to bypass.
Making IT the Data Owner
IT staff do not know whether a contract is commercially sensitive or whether a clinical spreadsheet contains identifiers. Business leaders do. Assign ownership where the knowledge lives and let IT act as custodian. Getting this backward means classification decisions get made by whoever is least equipped to make them.
No Technical Enforcement
A label with no control behind it is a suggestion. If Restricted data can still be attached to a personal email or copied to an unencrypted drive, the policy documents an intention rather than a control. Assessors increasingly ask for evidence that the rule is enforced, not just written.
Never Revisiting It
Business changes. A new product line, an acquisition, a first defense subcontract, or a move to a new cloud platform can invalidate a classification scheme that was accurate last year. Annual review with event-driven triggers keeps the policy connected to reality. Where a program lacks the internal bandwidth for that cadence, a virtual CISO engagement is usually the cheapest way to keep it current.
How Petronella Technology Group Builds Classification Programs
Petronella Technology Group has been securing regulated businesses from Raleigh since April 2002, and classification work sits at the front of nearly every compliance engagement we run. We start with discovery rather than documentation, because the policy is only as good as the picture of your data underneath it. That means an asset inventory, a scan for regulated data patterns, and a walkthrough with the business owners who know what the records actually are.
The written policy is generated through ComplianceArmor®, our compliance documentation platform, which produces the classification policy, the handling matrix, and the supporting evidence artifacts in the format the relevant framework expects. That matters more than it sounds: a great deal of assessment friction comes from having the right controls described in the wrong structure. ComplianceArmor® covers CMMC, HIPAA, SOC 2, PCI DSS, and CCPA modules, so a business subject to more than one framework does not maintain four separate versions of the same policy.
From there the work is enforcement and monitoring. Access groups and encryption settings implement the matrix, data loss prevention rules stop Restricted content at the boundary, and our Managed XDR service watches the alerts that result. Where classification surfaces vulnerable systems holding sensitive data, vulnerability management prioritizes remediation by what the system holds rather than by raw severity score, which is a materially better use of a limited patching window.
Craig Petronella, our founder, is a CMMC Registered Practitioner, an NC Licensed Digital Forensics Examiner, and the author of fifteen books on cybersecurity and compliance. The forensics background shapes how we approach this work: after an incident, the first question is always which data was exposed, and organizations that classified their data in advance can answer it in hours instead of weeks. Our digital forensics team has watched that gap decide how bad a breach notification gets. You can find the full library on our books page.
Petronella Technology Group is a CyberAB Registered Provider Organization, RPO #1449, and has held a BBB A+ rating since 2003. We serve Raleigh, Durham, Chapel Hill, Cary, Apex, and the wider Research Triangle, along with clients nationwide. To talk through your data, call 919-348-4912 or use the contact form.
Data Classification Policy Questions
What is a data classification policy?
What are the four data classification levels?
What is data classification based on?
How many classification levels should a small business use?
Is data classification required for compliance?
What is the difference between data classification and data categorization?
How do you classify data you already have?
Who is responsible for classifying data?
Know What You Hold Before Someone Else Finds Out
Petronella Technology Group builds data classification programs that start with discovery, produce audit-ready documentation, and end with controls that actually enforce the rules. Call 919-348-4912 or reach out to scope the work.
Reviewed by Craig Petronella, CMMC Registered Practitioner and NC Licensed Digital Forensics Examiner (DFE #604180)
Last Updated: August 13, 2026