Two claims about AI in medicine travel together, and they sound like they can’t both be true. One says machines already read scans as well as specialists. The other says chatbots fumble the basics of diagnosis. Both hold more truth than you’d expect. And the reason matters if you ever sit in an exam room while software weighs in on your case.
Here’s the catch. “AI” in a hospital isn’t one thing. A tool the FDA cleared to flag one finding on a CT scan has almost nothing in common with a chatbot you ask about a rash. Lump them together and you get hype on one side and panic on the other.
Pull them apart and the picture gets clearer. It also gets more interesting. Diagnostics, bedside decision support and drug discovery each tell a different story about what these systems can do today, and what they still can’t.
Why AI in Medicine Means Two Very Different Things
Start with the number that surprises most people. The FDA has authorized more than 1,600 AI-enabled medical devices for marketing in the US, and it tracks every one on a public list of AI-enabled medical devices. These aren’t lab demos. They’re regulated products in clinical use.
Now the second number. In April 2026, a Mass General Brigham team tested 21 large language models, including Claude, DeepSeek, Gemini, GPT and Grok. The models failed to produce an appropriate differential diagnosis more than 80% of the time, as Euronews reported on the JAMA Network Open study.
Same field. Opposite headlines.
The gap comes down to scope. FDA-cleared imaging tools are narrow and task-specific. They do one job, like flagging a possible stroke on a CT or triaging a mammogram. A general chatbot tries to reason about any patient with any complaint. That’s a far harder problem, and the evidence says nobody has solved it yet.
If you want the broader vocabulary, our explainer on narrow vs generative vs agentic AI shows why these categories behave so differently. For medicine, the short version works fine. Narrow AI is in clinics now. General AI as a diagnostician isn’t ready.
Most coverage blurs that line. Once you see it, a lot of confusing news starts to make sense.
AI in Diagnostics: Narrow Tools That Stay in Their Lane
Imaging shows the narrow approach at its clearest. A scan is a structured, repeatable input. A model can learn what one specific pattern looks like, then flag it for a human reader.
What an FDA-cleared AI device actually does
Picture a radiology department on a busy night. A narrow tool might check each incoming head CT for signs of a stroke and move suspicious scans up the queue. It doesn’t write the report. It doesn’t talk to the patient. It reorders the pile so a radiologist sees the urgent case sooner.
That modest job description is the point. These systems earn trust by doing one thing inside tight boundaries. Step outside those boundaries and they have nothing useful to offer.
It’s also why they can perform reliably in their lane while general chatbots stumble on open-ended cases. They aren’t reasoning about the whole patient. They’re matching one pattern they’ve seen many times before.
Why “better than radiologists” misses the point
Headlines love a machine-versus-human scoreboard. The real comparison is messier.
A model can look excellent on the data its developers trained and validated it on. Then a new hospital runs it on different scanners and different patients, and accuracy drops. External validation keeps showing that measurable slide, and it’s one of the most consistent caveats in diagnostic AI.
We looked at where each side actually wins in AI vs radiologists. The lesson carries over here. The useful question isn’t who’s smarter. It’s which tasks you can safely hand over, and who checks the work.

What Happens When You Ask a Chatbot to Diagnose You
That Mass General Brigham study deserves more than a headline. The team ran 21 models against 29 standardized clinical vignettes from the MSD Manual. That produced 16,254 responses, which they scored with a new assessment tool called PrIME-LLM. JAMA Network Open published the results on April 14, 2026.
The 80% number, read carefully
The failure rate applies to the differential diagnosis. That’s the early list of possibilities a doctor builds before the test results arrive. It’s the most open-ended step, and it’s where the models struggled most.
They did better later. At the later diagnostic stages, final-diagnosis misdiagnosis rates fell below 40%. The best models topped 90% accuracy on the final diagnosis.
So the fair reading isn’t “AI can’t diagnose.” It’s narrower than that. These models do much better once the uncertain early work is done. Trouble is, that early work is exactly what worried patients want help with at 2 a.m.
Co-author Marc Succi put it plainly: “Despite continued improvements, off-the-shelf large language models are not ready for unsupervised clinical-grade deployment.”
Notice two words there: off-the-shelf and unsupervised. He isn’t writing off the technology. He’s saying the version you open in a browser tab shouldn’t practice medicine on its own. If you’ve wondered how far to trust a chatbot’s read on your symptoms, our piece on whether you can trust an AI to diagnose you goes deeper.
Clinical Decision Support: In the Loop or On the Loop?
Clinical decision support means software that offers suggestions while a clinician keeps the decision. It sounds like the safe middle ground. Keep the doctor in charge, add AI as a second opinion, and you should get the best of both.
The evidence doesn’t back that assumption. Adam Rodman, an associate editor at the New England Journal of Medicine, spoke at the 2025 Penn Medicine Nudges in Health Care Symposium. “Just giving a human an AI system doesn’t inherently improve that human’s performance,” he said. The pair can even do worse than if “the AI system had just run by itself.”
That’s an uncomfortable result. It means workflow design matters as much as the model.
Researchers describe the choice as two setups:
- In the loop: the physician makes every final call, and the AI advises.
- On the loop: the AI acts, and the physician monitors and can override.
Today’s clinical AI mostly stays in the first camp. Moving toward the second runs into a wall that has little to do with accuracy.
The liability problem nobody designed for
Say an AI system suggests the wrong diagnosis and a patient suffers for it. Who answers for that? Right now, the clinician does.
According to Penn LDI, the University of Pennsylvania’s health policy research center, no established mechanism exists to shift legal responsibility from the doctor to the AI tool or its maker. That single gap explains a lot. A hospital has little reason to let software act alone while a human still carries all the risk.
So “AI assists, the doctor decides” isn’t only a philosophy. For now, it’s the only arrangement that makes legal sense for the person signing the chart.
Rodman had another line worth keeping on a sticky note: “[The] tech is remarkably powerful, but it is not magic… there is no superintelligence hidden behind the curtain.”
AI in Drug Discovery: Three Molecules Worth Knowing
Drug discovery is where AI in medicine gets exciting. It’s also where the claims need the closest reading. Three examples show what AI actually contributes, and where its contribution stops.
Halicin: searching 107 million molecules
In 2020, the MIT team of Regina Barzilay and James Collins trained a machine-learning model on roughly 2,500 known molecules. Then they pointed it at the ZINC15 database and screened 107 million compounds in silico, meaning on a computer rather than at a lab bench.
The model narrowed that ocean down to 23 candidates. One of them, halicin, turned out to be a new antibiotic. Across the CDC panels of multidrug-resistant pathogens the team tested, it worked against 35 of 36.
The mouse results were striking. Halicin cleared Acinetobacter baumannii skin infections within 24 hours and C. difficile gut infections within four days. The work appeared in Cell in February 2020.
No chemist could screen 107 million molecules by hand. The model’s gift is scale. It reads the whole haystack so humans only test the few needles it points to.

Abaucin: the same trick in under two hours
The sequel matters more than people realize. In 2023, the same approach screened about 7,680 compounds in under two hours and found abaucin. It’s a narrow-spectrum antibiotic that targets drug-resistant Acinetobacter baumannii specifically, as MIT News described.
Why care about a drug that does less? Because it proves the pipeline is reusable. Pick a new pathogen, retrain, rerun. Halicin came from a broad search across a giant library. Abaucin came from a fast, targeted one. Same method, very different jobs.
That reusability is the quiet story in machine learning for pharmaceutical research. The hard part is building a method that works. After that, each new run costs far less.
INS018_055: an AI-designed drug in human trials
Insilico Medicine’s INS018_055 goes a step further. The company describes it as the first drug with both an AI-discovered target and an AI-generated molecular design to reach human trials. It’s aimed at idiopathic pulmonary fibrosis.
The timeline, according to Insilico:
- February 2021: preclinical candidate selected, 18 months after the project began
- November 2021: first-in-human microdose trial
- January 2023: positive Phase I results
- February 2023: FDA Orphan Drug Designation
- June 2023: first Phase II patients dosed, in a randomized, double-blind, placebo-controlled trial of about 60 patients across roughly 40 sites in the US and China
Insilico says the path from novel target to Phase I took under 30 months, about half the time of traditional drug discovery. That’s a company talking about its own product, so hold the “half” loosely. Still, the milestones are concrete dates, not marketing adjectives.
What AI speeds up, and what it can’t skip
Look at that timeline again. AI compressed the discovery phase. It didn’t let the drug skip Phase I or Phase II. INS018_055 went through the same regulatory gauntlet as any other drug, with real patients and a placebo arm.
That’s the honest shape of AI drug discovery today. It finds better starting points faster. Proving a molecule is safe and effective in people still takes as long as it takes. For a look at how one AI lab is pitching its own role here, see our take on Anthropic and AI drug discovery.
Where AI in Medicine Still Falls Short
Put the evidence side by side and the weak points line up:
- Open-ended diagnosis. General models failed differential diagnosis more than 80% of the time, across 16,254 responses from 21 models. That’s no anecdote you can wave away.
- Generalization. Models that shine on their home data tend to slip at new hospitals with different scanners and patients.
- Human-AI teamwork. Adding a doctor to an AI system doesn’t automatically beat either one alone. The workflow decides.
- Liability. Nothing yet moves legal responsibility onto the tool or its maker, so clinicians stay firmly in charge.
- Clinical trials. AI can speed up discovery, but every candidate still has to prove itself in human studies.
Look at how few of these come down to raw accuracy. Liability and workflow gaps hold back more autonomous clinical AI as much as any error rate does. Most explainers skip that part, and it’s the part hospitals actually wrestle with.
AI in Medicine Today: Confirmed vs. Still Projected
The scorecard, as of 2026.
Confirmed and in use:
- More than 1,600 FDA-authorized AI-enabled devices, with imaging tools built for narrow, specific tasks
- AI-found antibiotic candidates, halicin and abaucin, that showed real activity against drug-resistant bacteria
- An AI-designed drug, INS018_055, that reached Phase II trials with human patients
Still projected or unresolved:
- General-purpose AI as an autonomous diagnostician
- Clinical decision support that acts “on the loop” instead of advising “in the loop”
- A legal framework that lets responsibility follow the software
- Reliable performance when a diagnostic model moves to a hospital it has never seen
The first list is shorter than the hype suggests. It’s also far more solid than the skeptics admit.
The Question Worth Asking Next Time
Here’s our read. The most useful things AI in medicine has done so far are quiet ones. It sorts scans. It searches chemical space. It shortens the list. The loud version, a machine that hears your symptoms and tells you what’s wrong, isn’t here yet, and the 2026 data says so without much ambiguity.
So if a clinic tells you AI helped with your care, don’t ask whether the AI is smart. Ask what job it had, and who checked its work. Those two answers tell you nearly everything.
The bigger question is one this site keeps circling back to. When machines get good at a narrow slice of human expertise, which parts should stay human on purpose? In medicine, it might be the moment someone sits with real uncertainty and still takes responsibility for the call.