In 2024, the Texas Attorney General settled with a healthcare AI company called Pieces over claims it had wildly oversold how accurate its software actually was. Pieces was summarizing patient charts inside real hospitals, telling doctors and nurses what a patient’s condition was, in real time, while advertising a “severe hallucination rate” of less than 1 in 100,000. Investigators weren’t convinced that the number was real. That’s not hypothetical. That’s four Texas hospitals feeding live patient data into a system whose accuracy claims turned out to be, at best, unverifiable.
This is the part that gets lost in most conversations about the ethical challenges of generative AI in healthcare. It’s easy to talk about bias and privacy in the abstract. It’s harder to sit with the fact that these systems are already inside emergency rooms, insurance departments, and diagnostic workflows, making calls that affect real people, right now, while the rules are still being written.
The Promise Nobody’s Arguing About
Nobody’s disputing that generative AI can help. It can draft clinical notes so doctors spend less time typing and more time looking at patients. It can flag drug interactions a tired resident might miss at 3 a.m. It’s been shown to cut inpatient prescribing errors by as much as 20% in some systems. That’s real, and it matters.
But every technology this powerful drags a shadow behind it. And in healthcare, the shadow isn’t an inconvenience; it’s a wrong diagnosis, a denied claim, a drug interaction nobody caught. The upside is well documented. The costs deserve equal airtime.
What Happens When the AI Is Wrong?
When a generative AI model is fed a small, planted error in a patient’s chart, it doesn’t always catch it; it often repeats it or builds on it. A recent study published in Nature Communications Medicine tested six leading language models against 300 clinical vignettes, each rigged with one fake detail: a lab value that didn’t exist, a symptom that was never there. The models parroted or expanded on that planted error in up to 83% of cases. A carefully worded prompt cut the error rate in half. It didn’t come close to eliminating it.
That number should stop you for a second. This isn’t AI inventing things out of thin air in a vacuum; it’s AI absorbing whatever’s already in the chart, mistakes included, and running with it. In a clinical workflow where a physician is skimming an AI-generated summary between patients, that’s the exact scenario where a hallucination turns into a missed diagnosis.
And clinicians know it. Surveys now show that over 90% of doctors who’ve used these tools have personally encountered a hallucination, and roughly 85% believe those hallucinations are capable of causing real patient harm. This isn’t a fringe worry. It’s the daily experience of the people using these tools.
The Bias Problem Isn’t Theoretical Anymore
Here’s the uncomfortable truth: generative AI models learn from the data they’re trained on, and healthcare data has never been evenly collected across race, gender, income, or geography. Feed a model decades of clinical notes shaped by unequal access to care, and it doesn’t just reflect that inequality; it can amplify it.
This isn’t abstract anymore, either. Insurance companies are currently facing class action lawsuits alleging their AI tools were used to override physicians’ own medical necessity determinations, essentially letting an algorithm second-guess a doctor at the bedside. Separately, a major commercial insurer is being sued over claims that its AI-driven fraud-detection tool carries racial bias. Whether or not either case succeeds in court, the underlying pattern is the one bioethicists have been warning about for years: an algorithm trained on unequal data, deployed at scale, can quietly widen the very disparities healthcare is supposed to close.
Who’s Actually Liable When AI Makes the Call?
Right now, liability in AI-driven healthcare is genuinely unsettled, and that ambiguity is itself one of the biggest ethical risks in the field.
When a human doctor makes a mistake, there’s a clear chain: the doctor, the hospital, maybe the insurer. When an AI model contributes to a bad outcome, who’s actually responsible? The company that built the model? The hospital that deployed it? The doctor who trusted its output? Courts and regulators are still working this out in real time, and the Pieces settlement is a preview of how messy it gets. The company faced consumer protection claims, not medical malpractice claims, because the legal categories built for human error don’t map cleanly onto software.
There’s also a newer, stranger wrinkle: AI-generated hallucinations showing up in legal proceedings about healthcare. In one 2025 federal case, a whistleblower lawsuit involving Intermountain Healthcare was dismissed after it came out that an expert witness’s report was riddled with AI hallucinations, fabricated deposition testimony, and quotes from documents that didn’t exist. The judge tossed the case with prejudice against the party who submitted it. Even the paperwork around healthcare AI isn’t immune to the same hallucination problem the AI itself is supposed to solve.
Your Medical Data Is Training Someone’s Model
Every time a generative AI tool summarizes a chart, drafts a note, or flags a risk, it’s typically processing protected health information to do it. Some of that data may end up shaping future versions of the model. Patients rarely get asked directly, and most consent forms weren’t written with this in mind.
Add to that a subtler issue researchers have started calling a challenge to “patient agency”: when AI-generated content starts filling gaps in medical records or research datasets, it’s substituting synthetic data for actual lived patient experience. The person whose case that data is supposed to represent may never know their story got smoothed over, or partly invented, by a model trying to fill in blanks.
The Black Box Problem: Nobody Can Explain the Diagnosis
Ask a doctor why they made a call, and they can walk you through their reasoning, the symptoms, the labs, and the differential diagnosis. Ask a large language model why it flagged a particular risk, and the honest answer is often: nobody fully knows. These systems are genuinely opaque, even to the engineers who built them.
That’s a problem for informed consent. It’s a problem for malpractice review. And it’s a problem for the doctor standing in the room who has to decide whether to trust a recommendation they can’t actually interrogate.
Is Healthcare AI Regulated Yet?
Regulation exists, but it’s fragmented and lagging behind deployment. Some states have started passing their own AI-specific healthcare laws, and agencies like the FTC have brought enforcement actions over deceptive accuracy claims, the Pieces case again being the clearest example. But there’s no single, comprehensive federal framework governing generative AI in clinical settings, and researchers reviewing the field describe current solutions as “fragmented,” warning that technical fixes like explainable AI mean little without matching governance, legal clarity, and professional training.
Separately, there’s a mental-health dimension regulators are only starting to grapple with. AI chatbot companies now face more than a dozen personal injury and wrongful death lawsuits tied to what clinicians are informally calling “AI psychosis”, cases where extended, validating conversations with a chatbot appear linked to worsening delusions or mania in vulnerable users. The American Psychological Association has already issued a formal warning against relying on AI chatbots for psychological treatment. That’s not the same category as a diagnostic hallucination, but it belongs in the same conversation: these are tools that can shape a person’s mental state, deployed with far less oversight than a prescription drug would ever get.
So, Where Does This Leave Patients and Doctors?
None of this means generative AI doesn’t belong in healthcare. It means the ethical challenges of generative AI in healthcare aren’t edge cases to patch later; they’re already showing up as settlements, lawsuits, and dismissed court cases. Bias, hallucinations, unclear liability, and thin regulation aren’t competing problems to rank by severity. They’re the same problem, wearing different clothes: a technology moving faster than the systems meant to hold it accountable.
The honest version of the conversation isn’t “AI good” or “AI bad.” It’s asking, every time a new tool shows up in a chart or a claims department: who checked this, who’s responsible if it’s wrong, and did the patient get any say in the matter at all. Right now, too often, the answer to at least one of those questions is still “nobody.”