Why No One Tracks Patient Safety Events From Everyday Health Software.
Sometime this week, you may look up a lab result, log a symptom in an app, or ask an AI about your health. If that software gets something wrong and the mistake hurts you, who is required to report it? Take your time; the answer is no one.
For many of us, this software is no longer an accessory to care. It is how care happens between visits, and at its best it is extraordinary. A newly diagnosed cancer patient can read her pathology report the moment it posts, ask an AI to translate its language at two in the morning, find a clinical trial across the country, and walk into her oncologist’s office with sharper questions than any patient a decade ago could have formed. A family chasing a rare disease diagnosis can use these same tools to shorten an odyssey that once took years. These technologies can be lifesaving, and patient communities know it better than anyone, because we have been early adopters of every one of them.
Celebrating what this innovation does well does not require turning a blind eye to what it breaks; both are a matter of patient safety. Yet nearly all of this software sits outside FDA safety review, and none of it comes with any duty to count the people it fails. That arrangement was written into law ten years ago, it has aged badly, and right now, for a few more weeks, there is something you can do about it.
What Congress traded away in 2016
In December 2016, the 21st Century Cures Act removed five categories of health software from FDA device review: administrative software, wellness apps, electronic health records, software that displays health data, and certain software that advises clinicians. Back then, this software mostly scheduled appointments, counted steps, and displayed charts, so the trade looked reasonable. Congress attached one condition: every two years, FDA must report on the risks and benefitsof what Congress excluded. That report is the check engine light for the whole system.
There was a fatal flaw in this plan: the warning light was never wired to a sensor for real patient outcomes. Hospitals must report when an infusion pump fails, because a pump is a medical device. When records software shows the wrong chart or a chatbot gives dangerous advice, no company, hospital, or developer owes a report to FDA or any safety regulator.
We’re only studying the benefit, not the risk.
FDA has published this report four times, and every edition reached the same two conclusions: benefits outweigh risks, and, quietly, harms “may be underrepresented” because nobody has to report them. Statisticians would point this out as a classic case of survivorship bias. During World War II, the military mapped bullet holes on returning bombers and proposed armor where the damage clustered, until someone pointed out that the fatal hits were on the planes that never came back, which is exactly why no one was studying them. Four FDA reports essentially say that “the benefits outweigh the risks,” even though we don’t really study or report on digital harms for what got defined as administrative software, wellness apps, software that advises clinicians, and so on. Patients harmed most, misdiagnosed, denied care, or steered wrong by a confident algorithm, often never learn software was the reason, so no record of them exists to study. A report built on missing data will always conclude everything is fine.
Meanwhile, traffic through this carve-out has exploded over the last decade. Digital tools that once displayed information now interprets it: wearables estimate blood pressure under a “wellness” label, and OpenAI’s new ChatGPT Health invites people to connect their medical records to a general-purpose AI. ECRI, the independent nonprofit that has studied health technology safety for over fifty years, named misuse of AI chatbots the number one health technology hazard for 2026. FDA’s own advisory committee devoted its November 2025 meeting to generative AI risks in mental health devices, while the same technology reaching the same patients through “wellness” apps gets no scrutiny at all. Through it all, nobody built the measurement system: there is still no dedicated front door for a patient to report that software hurt them.
Real patient harms, found by accident
Harms are surfacing anyway, through case reports, lawsuits, journalism, and academic studies, which is another way of saying…. by accident. It’s not a full picture of the health tech ecosystem, so we only see the acute examples that make the news. One such example: after consulting ChatGPT about cutting salt from his diet, a 60-year-old man replaced sodium chloride with sodium bromide and spent three weeks hospitalized with bromide poisoning, including hallucinations; we know only because his doctors published the case. An eating disorder nonprofit’s wellness chatbot gave weight loss advice to people seeking help for eating disorders until an advocate posted screenshots. Families allege in court that chatbots validated their children’s suicidal thoughts instead of interrupting them, and similar suits keep coming.
Quietest and possibly biggest place we’re not tracking patient safety issues: prediction algorithms. Congress expressly excluded software that predicts health care use, determines benefit eligibility, and manages “population health,” and that software now decides real care at enormous scale. A landmark Science study found a population health algorithm of a type applied to roughly 200 million Americans a year systematically underestimated Black patients’ needs. STAT’s “Denied by AI” investigation showed UnitedHealth using a model to forecast how long elderly patients should need rehabilitation, cutting off care when recovery outlasted the forecast; a lawsuit alleges nine in ten appealed denials were reversed, and a Senate investigation found denial rates surged as insurers adopted these tools. Every one of these is a patient safety event, known only because someone went looking.
Whose job is it to study the harm?
Right now we have a system where it’s everyone’s job to surface harm and weigh whether risks outweigh benefits, which in practice means no one’s job. Medicare regulators have told insurers that an algorithm alone cannot justify denying care. Consumer protection authorities have fined health apps for deceptive practices. Hospitals run confidential internal safety programs, but those are built for clinicians, closed to patients, and sealed from public view. Each of these groups watches its own yard, and not one of them is assigned the question that matters here: is this software hurting patients, and how often? FDA’s report to Congress, due every two years, is the only place federal law requires anyone to put the whole picture together. After all, the FDA cannot regulate these products. But we can make the one required accounting to stop being written with missing data, and to tell Congress what it will take to get real numbers.
Counting harms is necessary for healthy innovation
Innovation and safety usually get framed as opponents, and we understand the worry: nobody wants lifesaving tools buried under paperwork, least of all the patients waiting for them. But the choice in front of FDA is not speed versus caution. Patients are asking for an accurate check engine light on the patient safety dashboard, not a speed limit.
Unmeasured harm is how promising technology destroys its own market, especially when a market is inflating into a bubble. When nobody measures harm, small failures compound quietly until they surface the expensive way: a viral screenshot, a lawsuit, a Senate investigation. By then the response is never calibrated. NEDA did not adjust its chatbot; it shut the program after the screenshots surfaced. Illinois did not tune AI therapy; it banned it. Hospitals did not patch the sepsis model; they spent years unwinding a tool that independent researchers found missed most cases. Once trust breaks, it breaks for everyone, because buyers cannot tell careful companies from careless ones. Economists call this a market for lemons: when quality is invisible, the careless undercut the careful, and eventually nobody trusts the category at all. Markets can only price risk that someone measures; capital is pouring into health AI on the assumption that safety risk is near zero, and every uncounted harm is a liability sitting off the books until it surfaces all at once. Valuations built on assumptions no one is permitted to test are the anatomy of an asset bubble. Investors who want health AI to be a durable market rather than a boom and a crash should be first in line asking for harm data.
Industries that scaled to touch everyone treat failure data as an asset. Aviation became both safe and enormous by counting every incident obsessively. Cybersecurity built shared public databases of vulnerabilities, and naming flaws openly built the trust that online commerce runs on. Health AI is halfway there: Stanford researchers run MedHELM, an open evaluation testing leading AI models against a clinician-built map of 121 real medical tasks, with every score public. Serious, evidence-based effort measures what these models can do before they reach patients. Almost nothing measures what happens after. Benchmarks without harm reporting are half a safety system, like licensing drivers on the written test and never keeping crash records.
So notice what patients are actually asking for, because it is not premarket review of every app. Our comment asks FDA for the lightest instrument in the regulatory toolbox: count the harms. A reporting channel does not slow a single product launch. All it threatens is the ability to claim a product is safe while making sure no one can check. If a product’s business model depends on nobody counting its harms, that is not innovation worth protecting. Companies building genuinely safe tools should want this data to exist, because it is how they prove they are different.
How do we stop flying blind?
Right now, FDA is writing the 2026 edition of this report and accepting public comments through August 13. Its request for input lists patients first among the groups it wants to hear from; we intend to hold them to it. We are filing a comment asking FDA to answer the question at the center of all this: how can anyone keep evaluating whether these tools are safe when no one is required to report harm from them? [LINK: TLC comment letter or announcement post]
If you are part of a patient organization, file a short comment supporting these recommendations and add one example from your community. If health software has ever failed you, describe what happened in a comment of your own. Until a real reporting channel exists, the public docket is the only place these experiences get counted. You are the missing data, and every comment filed puts one of the missing planes on the map.
Congress meant this report to keep a light on. For ten years, the room has been dark and the reports have called it tidy. August 13 is the next chance to change that.
Sources
- 21st Century Cures Act, Pub. L. 114-255, Section 3060
- FDA, Reports on Non-Device Software Functions (2018–2024)
- Survivorship bias, Wikipedia
- Healio on ChatGPT Health (Jan. 2026)
- AHCJ, ECRI Top 10 Health Technology Hazards for 2026
- Federal Register, FDA Digital Health Advisory Committee meeting (Sept. 12, 2025)
- Annals of Internal Medicine: Clinical Cases, bromism case report (2025)
- NPR on the Tessa chatbot (June 2023); CNN (June 2023)
- CNN, Raine v. OpenAI (Aug. 2025); CNN, subsequent suits (Nov. 2025)
- Obermeyer et al., Science (2019)
- STAT, “Denied by AI” series; STAT (Mar. 2023); Senate PSI report (Oct. 2024)
- Wong et al., JAMA Internal Medicine (2021)
- Illinois IDFPR, Wellness and Oversight for Psychological Resources Act (2025)
- Akerlof, “The Market for Lemons,” overview
- CVE Program
- MedHELM leaderboard, Stanford CRFM
- Bedi et al., MedHELM, Nature Medicine (2025)
Discover more from Light Collective
Subscribe to get the latest posts sent to your email.