
- Healthcare workers’ use of AI tools like ChatGPT, Copilot, and MedLM in clinical settings is widespread.
- Its rapid growth over the last two years has outpaced policies to govern its use.
- Numerous studies have shown that tools like these can be inaccurate and incomplete, and sometimes fabricate answers that put patient safety at risk.
- The use of AI is largely unregulated without widely accepted safeguards.
- Three national organizations have recently published suggested policies and safeguards to govern the use of AI in healthcare.
While the promise of AI is immense, its adoption has dramatically outpaced the development of formalized guidelines or universally accepted best practices.
For the healthcare sector, this rapid adoption is risky. The central issue is that unregulated integration into clinical environments creates significant vulnerabilities related to data accuracy, patient privacy, and clinical safety.
The Extent of AI in Healthcare Today: A Structural Shift
To grasp the urgency of the risk, look at recent data on the clinician workforce. According to the Wolters Kluwer Health Future Ready Healthcare Survey Report, real-world AI use has crossed the Rubicon. Between 2025 and 2026, the share of U.S. physicians using AI multiple times daily for work skyrocketed, tripling from 10% to 38%. Overall, broad adoption surveys indicate that physician use of AI is now widespread.
The nursing workforce is also rapidly integrating AI. Daily AI use among nurses doubled year over year, reaching 32% by early 2026. Data from Elsevier’s Global Clinician Futures Report indicate that 41% of nurses actively use AI at work, primarily using general-purpose models for patient education, professional training, and medical literature queries.
Curiously, this rapid integration isn’t driven solely by hospital administrators; patients also play a role. More than 40% of patients report using AI daily for personal health queries, and 42% frequently bring AI-generated medical summaries to their appointments. This puts pressure on frontline clinicians to interpret and integrate AI insights on the fly.
When AI Errs: Public Failures and Medical Vulnerabilities
Because this rollout has lacked rigorous systemic controls, public reports and medical studies are increasingly documenting clinical AI failures. These are not speculative theories; they are real-life operational errors that reinforce the need for governance.
- Inaccurate and Incomplete Medical Evidence: A study published in BMJ Open tested five publicly available AI chatbots with medically related questions. The researchers found that nearly half (49.6%) of the AI-generated responses were problematic, with roughly 20% deemed “highly problematic”—meaning they provided inaccurate or incomplete answers that contradicted scientific consensus. Disturbingly, the chatbots delivered these erroneous answers with high confidence, rarely including clinical caveats or disclaimers.
- Fabricated Research and Technical “Hallucinations”: When the FDA tested its internal AI assistant, Elsa, staff discovered that the tool hallucinated non-existent scientific studies and misrepresented citations, requiring tedious line-by-line verification. Similarly, Google’s Med-Gemini demonstrated the danger of blending medical terms into highly authoritative yet entirely fake outputs—such as inventing a non-existent brain structure called the “basilar ganglia” (merging the very real basal ganglia and basilar artery).
- Phantom Clinical History in Medical Scribes: On frontline clinical forums, physicians have documented critical software errors in which automated ambient AI scribes—used to listen to patient encounters and draft progress notes—inserted entirely false medical histories into patient charts. In one documented instance, an AI scribe reported in the History of Present Illness (HPI) that a patient had a history of a below-the-knee amputation, despite the patient having both limbs and the visit having nothing to do with orthopedics.
- Premature Diagnostic Alerts: Industry compliance reports document instances in which AI-enabled diagnostic tools, such as automated ECG interpretation software, have triggered severe, unverified emergency alerts—flagging a healthy 29-year-old woman as actively experiencing a heart attack and transmitting the alert across clinical communications channels before a physician could validate the tracing.
These errors reveal dangerous patterns: narrow inputs that lack complete, holistic patient context; automation bias, in which rushed staff assume that an authoritative-sounding output must be accurate; and unclear handoffs, in which no single human clinician owns the final verification of data entering the permanent Electronic Health Record (EHR).
Veterans Health Administration Audit Reveals Shadow AI Problems
A stark illustration of these administrative oversight failures occurred when the Department of Veterans Affairs Office of Inspector General (OIG) published its national report, Review of Generative Artificial Intelligence Chat Tools for Clinical Use at the Veterans Health Administration.
The OIG’s review found that more than 15,000 Veterans Health Administration (VHA) staff were actively using two general-purpose AI tools: VA GPT and Microsoft 365 Copilot Chat. Because these platforms were readily accessible, personnel routinely used them for clinical tasks involving sensitive patient information.
The critical breakdown identified by the OIG was VHA leadership’s failure to recognize its clinical impact. Under federal risk management directives, agencies must identify “high-impact” AI applications and implement mandatory risk management protocols, including rigorous pre-deployment testing and ongoing monitoring.
VHA leadership failed to classify these general-purpose chat tools as high-impact, minimizing risk during interviews by characterizing them as “analogous to search engines” and shifting the entire burden of accuracy to individual users. By contrast, the VHA correctly classified its targeted ambient AI medical scribe tool as high-impact, subjecting it to formal data controls and human-in-the-loop validation.
The OIG rejected the notion that general-purpose chat tools are harmless search engines, reminding administrators that a single confident but incorrect output can mislead a care team and waste critical minutes. Furthermore, the VHA lacked an AI-specific reporting mechanism to track and trace errors in AI-generated documentation.
For HIPAA compliance officers, the lesson is clear: formally govern general-purpose AI tools, define approved uses, and require verification before AI output enters the patient care workflow. Otherwise, your workforce may use them as “Shadow AI.” Treating an LLM as a harmless search engine overlooks the profound legal and clinical liabilities posed by unvetted information. The central concern is governance, because unvetted AI output can carry direct patient-care consequences.
Guardrails: The Health-ISAC Policies and Safeguards
To prevent the vulnerabilities highlighted in the VA OIG review, covered entities should establish clear policy boundaries. The Health Information Sharing and Analysis Center (Health-ISAC) published a framework titled Policies and Safeguards for the Safe Use of AI. This guidance provides core components of an organizational risk strategy.
The Health-ISAC framework says that AI adoption requires defined authority, documented processes, and clear limits before tools are integrated into daily workflows.
Acceptable vs. Prohibited Uses
Health-ISAC provides a clear delineation that compliance officers can incorporate into corporate policy:
- Acceptable Uses: Using approved, closed-loop AI tools to generate general, non-confidential documentation, draft job descriptions, summarize public academic research, or perform routine code debugging. Even here, a human must perform compliance and accuracy checks before the output is used.
- Prohibited Uses: Inputting confidential company data, trade secrets, or Protected Health Information (PHI) into public or open AI models is prohibited. Misrepresenting AI-authored text as entirely original clinical reasoning is similarly restricted.
Technical Safeguards Against Data Leakage
When data is submitted to external, public AI platforms, it is often retained by the vendor to train future model iterations. For a covered entity or business associate, this constitutes an impermissible disclosure under the HIPAA Privacy Rule. To defend against data loss and unauthorized data exfiltration, implement a layered defensive approach:
- Data Minimization & Anonymization: Strip all 18 HIPAA identifiers from datasets before any technical processing occurs.
- Technical Access Controls: Use browser warnings, plug-in blockers, and automated firewalls to detect and log unauthorized employee prompts to external AI platforms.
- Rigorous Cybersecurity Practices: Implement multi-factor authentication (MFA), secure API management, continuous vulnerability scanning, and routine penetration testing for all internal AI software integrations.
The HSCC Cyber Governance Framework
While policies dictate employee behavior, institutional infrastructure must be structurally secured against systemic risk. To address this, the Health Sector Coordinating Council (HSCC) released its AI Cyber Governance Framework Implementation Guide. The framework bridges the gap between high-level ethical guidelines and technical implementation for healthcare IT professionals and software business associates.
The HSCC framework treats AI tools as critical components of the third-party vendor supply chain. Deployed clinical AI models cannot be evaluated once and then forgotten; they must be managed through active risk management. The implementation guide emphasizes three core components:
- AI Asset Inventory and Mapping: Covered entities must maintain an active registry of software that uses AI algorithms within their ecosystem, including any underlying vendor patches. This mirrors upcoming regulatory shifts in HIPAA requirements that mandate comprehensive device inventorying and detailed PHI data mapping.
- Model Transparency via Model Cards: Healthcare IT departments must require “Model Cards” from their software vendors. Similar to a nutritional label for software, a model card details the origin of the training dataset, the algorithm’s known limitations, specific performance benchmarks, and potential blind spots or biases across diverse patient populations.
- Continuous Monitoring for Model Drift: An AI algorithm’s accuracy can degrade over time as it encounters real-world clinical variations—a phenomenon known as “model drift.” The HSCC framework mandates continuous validation protocols to ensure that clinical decision support tools do not gradually lose accuracy, thereby jeopardizing patient safety.
Voluntary Standards: The Joint Commission’s RUAIH Certification
As the market seeks a definitive baseline for evaluating whether a healthcare organization has successfully integrated policy and technical controls, the nation’s premier accreditation body has stepped forward. The Joint Commission launched the Responsible Use of AI in Healthcare (RUAIH) Certification.
The RUAIH certification is a voluntary program open to hospitals, critical access organizations, and large health systems. Crucially, the certification does not validate or endorse individual AI products. A healthcare practice cannot purchase a “certified” AI tool to achieve compliance. Instead, the program certifies the organization’s permanent operational governance.
The RUAIH framework is organized around five core standards:
- Governance: Establish a formal, centralized AI oversight structure (e.g., an AI Committee) with clearly defined individual accountability and standardized procurement policies.
- Effective Data Management: Enforcing strict controls over clinical data used to train or fine-tune AI systems, emphasizing data integrity, privacy compliance, and proper de-identification.
- Risk and Bias Reduction: Actively evaluating algorithms to detect inherent biases that could compromise care for specific demographics, with documented pre-deployment and operational mitigation strategies in place.
- Monitoring and Validation: Maintaining ongoing evaluation processes to confirm the safety, clinical performance, and accuracy of AI systems over time and prevent model drift.
- Transparency and Education: Providing clear disclosures about when and where AI is used in the care lifecycle, alongside mandatory, continuous training for clinicians who rely on these tools.
The Joint Commission’s program signals a major shift: responsible AI use is no longer viewed merely as an IT preference. It is officially recognized as a patient-safety, clinical-quality, and institutional-trust issue.
Action Items for Healthcare Leadership
For healthcare practice owners, business associates, compliance directors, and legal counsel, treating clinical AI as an unmanaged frontier invites regulatory enforcement actions, data breaches, and medical malpractice litigation. Organizations should immediately implement the following structural updates to ensure compliance:
- Audit and Map Deployed AI: Conduct an organization-wide assessment to identify all software tools that currently use machine learning or generative capabilities. Eliminate the use of unapproved “Shadow AI” chat platforms.
- Execute Business Associate Agreements (BAAs): Ensure that any AI vendor that processes, touches, or accesses patient data executes a comprehensive BAA. If a vendor refuses to sign a BAA, or if their public terms of service state that data is harvested for model training, clinical usage must be legally prohibited.
- Establish an AI Governance Committee: Form an interdisciplinary committee comprising IT professionals, compliance officers, clinical staff, and legal counsel to vet and authorize AI tools before deployment.
- Update Patient Safety Reporting Systems: Modify existing clinical incident reporting workflows to include specific indicators for AI-related software errors, documentation inaccuracies, or prompt failures.
- Train the Clinical Workforce: Implement mandatory training that emphasizes that AI tools are purely supplementary. Clinicians must maintain absolute oversight and independently verify all AI-generated text, summaries, and diagnostic suggestions before entering them into an EHR.
By blending the structural and behavioral controls of Health-ISAC, the technical third-party vendor tracking of the HSCC, and the governance blueprints established by The Joint Commission, healthcare organizations can safely harness the profound power of artificial intelligence while remaining firmly anchored in patient safety and regulatory compliance.

