All articles

HIPAA, PHI, and Why Clinical Practices Need Private AI

Medical private AI and HIPAA: why cloud LLMs are a compliance risk for clinical practices, and how on-prem AI keeps PHI inside your office.

A psychiatrist in Encinitas pastes a session note into ChatGPT to clean up the language before it goes in the chart. A front-desk coordinator at an OB/GYN practice in Carlsbad drops a patient's intake form into an AI summarizer to speed up triage. An oncology practice manager uses a cloud transcription tool to convert dictated referral letters into finished drafts.

Every one of those workflows is a HIPAA problem. Not "might be" — is. And most of the clinicians running them have no idea, because the tool is fast, the output is good, and nobody at the vendor sent them a scary email.

This is the conversation clinical practices need to be having in 2026. Not "should we use AI" — you already are, whether leadership signed off or not. The real question is where the model lives, who has access to the prompts, and whether the vendor will put their signature on a Business Associate Agreement that actually covers what the model does with your data.

The HIPAA framing nobody wants to say out loud

Under HIPAA, any vendor that creates, receives, maintains, or transmits Protected Health Information on behalf of a covered entity is a Business Associate. Full stop. That's not a SentriCraft interpretation — that's the definition in 45 CFR 160.103.

An LLM that ingests a clinical note, a patient email, or an intake form is doing exactly that. It's receiving PHI. It's transmitting PHI back to you. In many cases, it's maintaining PHI in logs, in caching layers, and — depending on the vendor's terms — in the training pipeline for future model versions.

That makes the LLM vendor a Business Associate. Which means you need a signed BAA before a single byte of PHI touches their infrastructure.

Here's where it gets uncomfortable:

  • Most consumer AI tools don't sign BAAs at all. ChatGPT's default consumer product, Claude.ai's free and Pro tiers, Gemini's consumer product — none of these are covered.
  • Some enterprise tiers will sign a BAA, but the agreement carves out significant categories of processing. Read the fine print on model improvement, telemetry, and abuse-monitoring pipelines. "We won't train on your data" is not the same thing as "your PHI never leaves the isolated tenancy."
  • The vendors that do sign real BAAs (Azure OpenAI Service under a specific configuration, AWS Bedrock with the right controls, a handful of clinical-specific SaaS platforms) require you to configure the environment correctly. A misconfigured Azure deployment is not a compliant Azure deployment.

The staff member pasting a chart note into a browser tab does not know any of this. They just know it works.

The four workflows that keep pulling clinicians toward AI

Before we talk infrastructure, it's worth naming why this problem exists in the first place. Clinicians aren't reaching for AI because they read a Gartner report. They're reaching for it because the paperwork is crushing them.

Clinical note summarization. A 45-minute psychiatry session generates a lot of raw material. Turning it into a structured progress note that satisfies billing, malpractice defense, and the next clinician who reads the chart is a real cognitive tax. LLMs are genuinely good at this.

Intake triage. New-patient forms, prior records, and referral packets arrive as PDFs, faxes, and portal uploads. Summarizing the salient history before the first visit saves 15–20 minutes per patient. In a busy OB/GYN or oncology practice, that's the difference between running on time and running an hour behind by lunch.

Patient communication drafts. Portal messages, appointment follow-ups, medication clarifications. AI-drafted responses that a clinician edits and signs are dramatically faster than starting from a blank text box. The volume of asynchronous patient messaging has grown every year since 2020 and shows no sign of reversing.

Referral letter generation. Structured, formulaic, and time-consuming. A model trained on the practice's existing letter templates produces first drafts that need light editing rather than full authorship.

These use cases aren't hypothetical. They're already happening in your practice. The only question is whether they're happening on infrastructure you control.

On-prem private AI: what it actually means

Private AI, in the sense we deploy it, means the model runs on hardware physically located inside your practice — or in a dedicated colocation cage you control — and no prompt, no completion, and no PHI ever leaves that boundary.

The technical stack looks like this:

A GPU server (typically a single workstation-class or 1U rack unit with an NVIDIA RTX 6000 Ada, an L40S, or similar) sits on an isolated VLAN inside your network. It runs an open-weights model — Llama 3.3, Qwen 2.5, or a medically fine-tuned derivative — served through a local inference engine like vLLM or Ollama. Staff access it through a browser-based chat UI or through direct integration with the EHR, but the traffic never leaves the LAN. The model has no internet egress. It cannot phone home. It has nothing to phone home with.

Compare that to the cloud path:

Prompt with PHI → HTTPS to vendor → tenancy in vendor cloud → inference on shared GPU → response back → logs retained by vendor for N days → potentially reviewed by vendor abuse team → potentially included in future training data depending on BAA terms.

Every arrow in that chain is a compliance surface. Every arrow is something you have to trust a vendor to have configured correctly, forever, across every product update they ship. In the on-prem architecture there are no arrows leaving the building.

This is one of the core reasons private AI infrastructure is a fundamentally different security posture than cloud AI, even cloud AI with a BAA. You're not trusting a vendor's promises. You're eliminating the transaction that requires the promise.

The honest caveat: on-prem is not automatic HIPAA compliance

This is the part most vendors selling private AI won't say clearly, so we will.

Running the model locally does not make you HIPAA compliant. It removes one large category of risk — third-party data disclosure — but it does not remove the other categories.

You still need:

  1. Access controls. Not everyone in the practice should be able to query the model. Role-based authentication tied to your existing directory service (Entra ID, Google Workspace, on-prem AD) means the front desk can draft appointment reminders but can't pull oncology summaries.

  2. Audit logging. Who queried what, when. HIPAA's audit control requirement (§164.312(b)) applies to any system that touches PHI, including the AI. Every prompt and response needs to be logged, timestamped, and retained per your practice's retention policy — typically six years.

  3. Encryption at rest. The GPU server has a disk. That disk has cached prompts, model weights, and potentially retrieval-augmented context that includes PHI. Full-disk encryption is not optional.

  4. Physical security. An on-prem server that anyone can walk up to and pull the drives from is not secure. The hardware lives in a locked room, ideally on a monitored circuit, with the same physical controls you'd apply to your EHR server.

  5. Incident response and breach notification procedures. These are HIPAA obligations regardless of what technology stack is generating the incident.

  6. Ongoing patch management. Model runtimes have vulnerabilities. Underlying Linux OS has vulnerabilities. NVIDIA driver stack has vulnerabilities. Someone has to be responsible for keeping this patched.

Private AI moves the compliance conversation from "do we trust the vendor" to "do we run our infrastructure competently." That's a much better position to be in. But it's still a position that requires actual work.

Which specialties benefit most

Any practice with sensitive longitudinal records gains disproportionately from keeping AI in-house. Three specialties stand out:

Psychiatry. Session notes contain some of the most sensitive PHI in medicine — substance use disclosures, trauma histories, custody-relevant statements, suicidality assessments. The consequences of a data leak here are not theoretical. Private AI for note summarization and treatment plan drafting keeps that material out of any vendor's telemetry pipeline.

OB/GYN. Reproductive health records have taken on new legal dimensions since 2022. Practices in states with strong privacy protections cannot assume that data crossing state lines through a cloud vendor stays protected. On-prem inference eliminates the interstate data transfer question entirely.

Oncology. Longitudinal treatment records, genetic testing results, and family history documentation accumulate over years. The value of AI summarization here is enormous — but so is the volume of PHI that would be transmitted to a cloud vendor to make it work.

Practices in family medicine, dermatology, and orthopedics benefit too. The intensity of the compliance argument just scales with the sensitivity of the records.

What the deployment actually looks like

A typical private AI installation for a mid-sized North County clinical practice runs like this:

The GPU server lives in the same server closet as the EHR appliance and the network core. It's on a dedicated VLAN with no internet egress and no path to the guest network. Staff workstations reach it through a locked-down internal URL, authenticated against the existing directory. The model is pre-loaded with the practice's letter templates, common referral formats, and — if the practice wants retrieval-augmented generation — a vectorized index of anonymized reference materials.

SentriCraft handles the network segmentation, the server provisioning, the authentication integration, the audit logging pipeline, and the ongoing patching. Your clinical staff handles the clinical judgment about what the AI is allowed to do. Your compliance officer handles the policy layer. Nobody handles a BAA with an AI vendor, because there is no AI vendor.

We do this work as part of our broader business network and infrastructure practice for professional services firms across North County — clinical practices, law firms, wealth managers, and any organization where the data on the wire is worth more than the wire itself.

The bottom line

The AI is coming into your practice whether you plan for it or not. Somebody on your staff is already using it. The choice is between a governed deployment on infrastructure you control, and a shadow deployment on infrastructure your BAA doesn't cover.

Private AI, correctly deployed, is not a magic HIPAA wand. It's a design decision that removes the largest and most opaque compliance risk in the AI stack, and leaves you with the risks you already know how to manage.

Get the architecture right at the start and the question stops being "can we use AI safely." It becomes "what do we want it to do next."

Ready for a network that just works?

Book a Free Consult