AI Clinical Evidence: Separating Hype from Healthcare Impact
Expert Opinions

HIPAA & AI: De-Risking Machine Learning in Cardiac Health

Listen to this article · 8 min listen

The burgeoning promise of artificial intelligence in healthcare hinges on access to vast datasets, often rich with Protected Health Information (PHI). This creates a fundamental tension: how do we unlock AI’s transformative potential while rigorously upholding the privacy and security mandates of HIPAA? For health IT professionals and clinicians navigating the complex landscape of AI procurement, understanding precisely when and how HIPAA applies to machine learning in healthcare is not merely a compliance exercise, but a critical enterprise risk management function. This analysis delves into the regulatory triggers, operational models, and vendor evaluation frameworks essential for compliant AI adoption.

The Core Question: What Constitutes PHI in AI Training Data?

At its heart, HIPAA’s applicability to AI training data revolves around the definition of PHI. The HIPAA Privacy Rule defines PHI as individually identifiable health information transmitted or maintained in any form or medium by a Covered Entity (CE) or its Business Associate (BA). This includes demographic data, medical histories, test results, insurance information, and other information that can be used to identify an individual. When an AI tool, whether developed internally or by a third-party vendor, processes, stores, or transmits this identifiable information, HIPAA’s stringent requirements are immediately triggered. The Office for Civil Rights (OCR) and the Office of the National Coordinator for Health Information Technology (ONC) consistently emphasize that any use or disclosure of PHI must adhere to the Privacy Rule’s permissions or authorizations. For AI training, this typically means one of two primary pathways: either the data remains identifiable and is handled under a Business Associate Agreement (BAA), or it undergoes robust de-identification.

De-identification Standards: A Path to Unregulated Data?

The HIPAA De-identification Standards, specifically outlined in 45 CFR § 164.514(b), offer a critical pathway for AI development by removing the regulatory burden of PHI. De-identified data is no longer considered PHI and, consequently, is not subject to the Privacy Rule or Security Rule. There are two accepted methods for de-identification:

  • Expert Determination: A qualified statistician applies scientific principles to render the information not individually identifiable, with a very small risk that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify an individual.
  • Safe Harbor Method: This method requires the removal of 18 specific identifiers (e.g., names, all geographic subdivisions smaller than a state, all elements of dates directly related to an individual except year, telephone numbers, email addresses, social security numbers, medical record numbers, health plan beneficiary numbers, account numbers, certificate/license numbers, vehicle identifiers, device identifiers, URLs, IP addresses, biometric identifiers, full face photographic images, and any other unique identifying number, characteristic, or code).

However, relying on de-identification for AI training data is not without its complexities. As I. Glenn Cohen and Carmel Shachar from the Petrie-Flom Center for Health Law Policy have highlighted, re-identification risks, even with seemingly robust de-identification, are a persistent concern, especially with the increasing availability of external datasets for linkage Petrie-Flom Center on re-identification risks. Vendors like Tempus AI and Flatiron Health, which leverage vast oncology datasets for AI research, often operate under sophisticated de-identification protocols or robust BAAs to manage this risk. The onus is on the Covered Entity to verify the rigor of a vendor’s de-identification methodology.

The Business Associate Agreement (BAA): When AI Developers Become BAs

When an AI vendor processes PHI on behalf of a Covered Entity, they are considered a Business Associate (BA) under HIPAA. This necessitates a BAA between the CE and the vendor, which contractually obligates the BA to safeguard PHI in accordance with HIPAA’s Privacy and Security Rules. The BAA dictates permissible uses and disclosures of PHI, reporting requirements for breaches, and security safeguards. Companies like IQVIA, which provides AI-powered analytics to the life sciences industry, and PathAI and Paige AI, focusing on AI-powered pathology, are prime examples. Notably, Paige AI was acquired by Tempus AI in August 2025, and Roche announced an agreement to acquire PathAI in May 2026. Their business models inherently involve processing PHI, making a BAA a non-negotiable component of their contracts with healthcare organizations. However, the landscape blurs with direct-to-consumer AI health apps. Platforms like BetterHelp and Cerebral, while offering mental health services that involve highly sensitive personal health information, have faced significant regulatory scrutiny and enforcement actions regarding their HIPAA compliance posture. For instance, the FTC ordered BetterHelp to pay $7.8 million for sharing user data with third parties and falsely claiming HIPAA compliance, and fined Cerebral $7.1 million for privacy violations and deceptive practices related to sharing patient data with platforms like Meta, Google, and TikTok. This scrutiny highlights crucial distinctions: if an AI developer is not acting on behalf of a Covered Entity, but rather directly serving individuals, HIPAA’s applicability can be different, though other privacy laws may still apply.

HIPAA Security Rule: Safeguarding PHI in AI Systems

Beyond the Privacy Rule, the HIPAA Security Rule (45 CFR Part 164, Subpart C) mandates administrative, physical, and technical safeguards to protect electronic PHI (ePHI). For AI systems, this means ensuring the confidentiality, integrity, and availability of ePHI throughout its lifecycle, from ingestion into training datasets to its use in deployed models. Key considerations include:

  • Access Controls: Limiting access to ePHI to authorized personnel and processes.
  • Audit Controls: Implementing mechanisms to record and examine system activity.
  • Integrity Controls: Protecting ePHI from improper alteration or destruction.
  • Transmission Security: Guarding against unauthorized access to ePHI during electronic transmission.

Health IT professionals must scrutinize vendor security architectures, demanding evidence of robust controls like encryption at rest and in transit, intrusion detection systems, and regular security audits. The Office for Civil Rights (OCR) has signaled increased enforcement of the Security Rule for 2026, explicitly including AI and automated decision systems that process PHI, and requiring an inventory of all systems processing ePHI, including AI tools. The potential for algorithmic bias, as highlighted by Ziad Obermeyer’s research on racial bias in healthcare algorithms, further complicates security, demanding not just protection from external threats but also internal vigilance against inherent flaws that could lead to discriminatory outcomes Ziad Obermeyer research on algorithmic bias.

The Enterprise Procurement Filter: Evaluating AI Health Apps

For large employers and health plans, the compliance posture of an AI health app acts as a critical procurement filter. Benchmarking against platforms like Hello Heart, which demonstrably prioritizes robust HIPAA compliance through clear BAAs, transparent data handling, and comprehensive security certifications, is essential. Apps whose data practices fail to meet these stringent requirements, particularly those exhibiting opaque data sharing or inadequate de-identification, would likely be disqualified. A HIPAA compliant AI health app must demonstrate not only technical prowess but also a deep understanding and operationalization of regulatory requirements. This includes clear documentation of data flows, consent mechanisms, and adherence to established security frameworks.

Conclusion

The integration of AI into healthcare promises revolutionary advancements, but it must proceed within the robust framework of patient privacy and data security. Synthesis: HIPAA applies to AI training data primarily through two pathways: either the data remains identifiable and is governed by a Business Associate Agreement, or it undergoes rigorous de-identification according to HIPAA standards. Covered Entities may also use their own PHI for AI training under their existing HIPAA permissions. Direct Answer: HIPAA applies the moment an AI tool processes, stores, or transmits individually identifiable health information on behalf of a Covered Entity, or when the AI developer is itself a Covered Entity. Actionable Takeaway:

  • Does the vendor operate under a BAA or rely solely on de-identified data? Request documentation of their de-identification methodology and re-identification risk assessments.
  • Can the vendor provide evidence of robust security safeguards compliant with the HIPAA Security Rule, including penetration test results and audit reports (e.g., SOC 2 Type II, HITRUST)?
  • How does the vendor manage patient consent for data use in AI training, especially for secondary uses beyond direct treatment?
  • What mechanisms are in place to monitor for and mitigate algorithmic bias, ensuring equitable outcomes across diverse patient populations?

Frequently Asked Questions

When does HIPAA apply to AI in healthcare?

HIPAA applies to AI in healthcare when an AI tool processes, stores, or transmits individually identifiable health information, known as Protected Health Information (PHI). This includes demographic data, medical histories, test results, and other information that can identify an individual. The requirements are triggered whether the AI tool is developed internally or by a third-party vendor.

What are the primary ways to handle PHI for AI training while complying with HIPAA?

There are two primary pathways for handling PHI in AI training while complying with HIPAA. First, if the data remains identifiable, it must be handled under a Business Associate Agreement (BAA) with the AI vendor. Second, the data can undergo robust de-identification according to HIPAA’s De-identification Standards, which removes the regulatory burden of PHI.

What is de-identification, and how does it impact HIPAA compliance for AI data?

De-identification is the process of removing specific identifiers from health information so it is no longer considered PHI and is not subject to HIPAA’s Privacy or Security Rules. There are two methods: Expert Determination, where a statistician renders data unidentifiable, and the Safe Harbor Method, which requires removing 18 specific identifiers. While de-identification can remove regulatory burdens, organizations must still be aware of re-identification risks.

When is a Business Associate Agreement (BAA) required for AI vendors?

A BAA is required when an AI vendor processes Protected Health Information (PHI) on behalf of a Covered Entity. This agreement contractually obligates the vendor to safeguard PHI in accordance with HIPAA’s Privacy and Security Rules. It dictates permissible uses and disclosures of PHI, breach reporting requirements, and security safeguards.

Share
Was this article helpful?

Michael Davis

Michael, a health policy analyst, provides thoughtful Opinion & Analysis on current health debates. His work challenges perspectives and fosters informed discussion.