The short answer

Separating PHI from business data for AI takes five steps.

  • Find where PHI really lives, including the places it shouldn’t be.
  • Classify and tag data by sensitivity.
  • Wall the two apart with access controls at the data layer, so a tool reaches only what a user is already cleared to see.
  • De-identify or redact where AI doesn’t need identifiers.
  • And confirm a Business Associate Agreement before any tool touches patient data.

Focus does this data work for healthcare practices as the step before AI goes on.

Why do you have to separate PHI from business data before using AI?

doctor thinking of phi data

Because HIPAA’s Minimum Necessary Standard applies to AI the same way it applies to a person. Under 45 CFR 164.502(b), access to protected health information has to be limited to the minimum needed for a task, and that limit follows the access rather than the actor: an AI tool querying on a staff member’s behalf is still reaching PHI on the covered entity’s behalf, so the same ceiling applies.

An AI tool that can reach every field in a patient record is already outside the standard, whatever the system technically permits. Auditors measure against that ceiling, so it’s the first thing to settle when you scope a pilot.

The exposure isn’t theoretical. IBM’s 2025 Cost of a Data Breach Report found that 97% of breached organizations that experienced an AI-related security incident said they lacked proper AI access controls. Point an AI agent at years of documentation and it can surface protected health information no human would have gone looking for.

And consumer AI tools aren’t a safe outlet: general assistants such as the free and standard ChatGPT tiers aren’t covered by a Business Associate Agreement, so pasting a patient note into one is an impermissible disclosure that OCR can treat as a breach.

Separating the two data sets first is what makes AI usable without turning it into a liability. It’s also the step most healthcare groups skip, because they’ve never had to draw the line inside their own systems before.

What counts as PHI versus business data?

Protected health information is any data that ties a person to their care: names, dates, chart numbers, encounter notes, lab results, images, and the other identifiers HIPAA lists, in any format. Business data is everything that runs the practice without pointing at a specific patient’s health: aggregate billing, payroll, contracts, staffing, and operational reporting.

The line matters for AI because a tool cleared for business data should never be able to reach the clinical side. In most practices patient data and business data currently sit in the same systems, which is where the work starts.

Properly de-identified health information is no longer PHI under HIPAA. That’s why de-identification shows up as a separation tool and not just as a compliance footnote. Strip the identifiers correctly and the same data set becomes safe for uses that don’t need identity, though it only counts if the stripping meets HIPAA’s standard.

Step 1: Find where PHI really lives

lock on top of several documents

PHI rarely stays where it belongs. It ends up in spreadsheets, shared drives, email threads, and channel messages, and you can’t separate what you haven’t located. A readiness scan with data loss prevention (DLP) tooling is how you surface it, and it’s the same first move behind any AI rollout.

Treat this as an inventory: which systems hold PHI, which hold business data, and where the two have bled together. Our guide on what AI readiness takes for a healthcare practice covers the broader assessment this fits inside.

Step 2: Classify and tag data by sensitivity

Data classification and tagging are what turn a pile of files into something an AI tool can respect. Sort data into tiers, then tag every document, email, and record by audience, so a tool can tell what it’s looking at and who it’s for. Without a valid tag it can’t tell a clinical note from a business memo, however capable the model is, and missing context is where wrong answers and over-sharing begin.

At minimum, separate protected health information, business and financial data, and general internal content, and carry a tag on each. The tags are the instructions your access controls and AI tools read in the next step.

How do you separate PHI from business data in practice?

Wall the two apart with access controls at the data layer, so an AI tool can only reach the information a given user was already cleared to see. There are four methods, and most groups use more than one together.

Method What it does Best for Watch-out
Environment segmentation Isolates the systems and storage that hold PHI from those that hold business data Keeping AI processing away from patient records The proposed 2026 Security Rule may require isolating AI environments
Row-level and role-based access control Limits which records and fields a given user or tool can reach Letting AI answer from only what a user is already cleared to see Must be enforced at the data layer, not just in the app
De-identification Strips the identifiers that make data PHI Analytics or model use where identity is not needed Must meet HIPAA’s de-identification standard, or it is still PHI
Redaction and DLP Detects and removes PHI before it reaches a tool Stopping copy-paste of PHI into general AI tools An approved tool can still over-share; needs data-level policy

 

Segmentation and access control keep the two data sets apart in daily use. De-identification and redaction handle the cases where PHI doesn’t belong in the tool at all. Either way, the enforcement has to sit at the data layer, because a rule that only lives in the application is a rule an agent can route around. Which of the four you start with usually depends on where your data already sits.

Step 3: Enforce access at the data layer with row-level security

This mirrors how your staff already work. Accounting has access to accounting data, not patient data, and providers have access to clinical data, not the financials. Set row-level and role-based access so AI sits inside those same lines. Enforced down in the data platform, that’s what satisfies the Minimum Necessary Standard for an AI agent, because the agent inherits the limits of whoever it’s acting for.

An approved-tool list is necessary but not enough. An approved tool can still receive more PHI than a task needs if the data underneath it isn’t scoped. Data-level control is what closes that gap, which is why this work tends to land with whoever runs your data platform.

Step 4: De-identify or redact where AI does not need identifiers

De-identification removes the elements that make data PHI, and properly de-identified data falls outside HIPAA, which makes it safe for analytics and model use. Where a use case doesn’t need to know who the patient is, strip the identity out. Redaction and DLP tooling handle the live cases, catching PHI before it reaches a tool that shouldn’t see it.

De-identification only counts if it meets HIPAA’s de-identification standard; a half-stripped record is still PHI. Done correctly, this is the step that lets a group get real analytic value from clinical data without carrying clinical risk into every AI query. Most groups make this call per use case, since the answer changes with what the analysis actually needs.

Step 5: Put a BAA in place before any tool touches PHI

two people discussing business agreement

A Business Associate Agreement is a legally binding contract, and without one, sending PHI to a vendor is an impermissible disclosure. Confirm one with any AI platform that will create, receive, maintain, or transmit protected health information, and do it before the tool goes live. For groups running in Microsoft Azure, a BAA with Microsoft is typically part of the subscription, one reason Azure-based tooling is a common healthcare starting point.

Sequence matters. Treat the BAA as the gate, and read the terms for how the tool handles PHI, what happens to data after processing, and whether anything is used to improve the vendor’s models.

The PHI separation checklist

  • Inventory where PHI lives today, including spreadsheets, drives, email, and channels where it does not belong.
  • Classify and tag data by sensitivity and audience across every system.
  • Set row-level and role-based access at the data layer so tools reach only what a user is cleared for.
  • Segment AI processing environments away from systems that hold patient records.
  • De-identify data for uses that do not need identity, to HIPAA’s de-identification standard.
  • Add redaction and DLP to catch PHI before it reaches a general AI tool.
  • Confirm a signed BAA with any platform that will touch PHI, before it goes live.

Where healthcare groups get PHI separation wrong

They treat it as an allow-list problem when it is a data problem. Approving a few tools and blocking the rest feels like control, but an approved tool with unscoped data underneath it can still over-share patient data, and staff who need a capability that isn’t approved will find a workaround. The durable fix is separation and access control at the data layer, so the limit travels with the data instead of depending on which app is open.

Get the data clean, classified, and walled off before you turn the lights on with AI. Done in that order, a group can hand its staff AI tools that only ever see the data they’re cleared for, and get the productivity without carrying every patient record into every prompt.

Focus has spent 16+ years working only in healthcare, and only in IT, security, and data. We’ve run more than 2,000 EHR conversions for healthcare practices, which means separating clean, tagged, access-controlled data from the messy reality of a live practice is the core of what we do. Separating PHI from business data before AI goes on is the same discipline, applied one step earlier.

Frequently asked questions

magnifying glass checking a puzzle pattern

Can you use ChatGPT with PHI?
Not the consumer tiers. Free and standard ChatGPT plans aren’t covered by a Business Associate Agreement, so entering protected health information is an impermissible disclosure under HIPAA that OCR can treat as a breach. Only an AI platform under a signed BAA, with access controls in place, should ever handle PHI.

What is the minimum necessary standard for AI?
Under 45 CFR 164.502(b), PHI access must be limited to the minimum needed for a task. That ceiling applies to an AI tool querying on a staff member’s behalf, so the tool must only reach the data its specific job requires, not everything the system can technically access. In practice you enforce that with row-level access control in the data platform.

Do you have to de-identify PHI before using AI?
Not always. If the use case needs patient identity, keep the data as PHI and protect it with access controls and a BAA. If it doesn’t, de-identify to HIPAA’s standard, which takes the data outside HIPAA and makes it safe for analytics or model use. Which route you take comes down to whether the work needs to know who the patient is.

Is business data safe to use with AI?
Only once it’s genuinely separated from PHI. Business data with no patient identifiers carries far less risk, but if clinical and business data sit in the same tables or drives, a tool cleared for business data can reach PHI by accident. Confirm they don’t before you open the business side up.

What happens if PHI and business data are not separated?
An AI tool inherits whatever the underlying data exposes, so it can surface protected health information to users or vendors who should never see it. That’s a HIPAA violation and, in IBM’s data, the condition behind the 97% of AI-related incidents where organizations said they lacked proper access controls. The exposure lives in the data, so that’s where it has to be closed.