AI in Medicine: What Actually Works in Diagnosis and Drugs?

Artificial Intelligence Published: 11 min read Pravesh Garcia
AI in Medicine What Actually Works in Diagnosis and Drugs
Rate this post

Two claims about AI in medicine travel together, and they sound like they can’t both be true. One says machines already read scans as well as specialists. The other says chatbots fumble the basics of diagnosis. Both hold more truth than you’d expect. And the reason matters if you ever sit in an exam room while software weighs in on your case.

Here’s the catch. “AI” in a hospital isn’t one thing. A tool the FDA cleared to flag one finding on a CT scan has almost nothing in common with a chatbot you ask about a rash. Lump them together and you get hype on one side and panic on the other.

Pull them apart and the picture gets clearer. It also gets more interesting. Diagnostics, bedside decision support and drug discovery each tell a different story about what these systems can do today, and what they still can’t.

Why AI in Medicine Means Two Very Different Things

Start with the number that surprises most people. The FDA has authorized more than 1,600 AI-enabled medical devices for marketing in the US, and it tracks every one on a public list of AI-enabled medical devices. These aren’t lab demos. They’re regulated products in clinical use.

Now the second number. In April 2026, a Mass General Brigham team tested 21 large language models, including Claude, DeepSeek, Gemini, GPT and Grok. The models failed to produce an appropriate differential diagnosis more than 80% of the time, as Euronews reported on the JAMA Network Open study.

Same field. Opposite headlines.

The gap comes down to scope. FDA-cleared imaging tools are narrow and task-specific. They do one job, like flagging a possible stroke on a CT or triaging a mammogram. A general chatbot tries to reason about any patient with any complaint. That’s a far harder problem, and the evidence says nobody has solved it yet.

If you want the broader vocabulary, our explainer on narrow vs generative vs agentic AI shows why these categories behave so differently. For medicine, the short version works fine. Narrow AI is in clinics now. General AI as a diagnostician isn’t ready.

Most coverage blurs that line. Once you see it, a lot of confusing news starts to make sense.

 

AI in Diagnostics: Narrow Tools That Stay in Their Lane

Imaging shows the narrow approach at its clearest. A scan is a structured, repeatable input. A model can learn what one specific pattern looks like, then flag it for a human reader.

What an FDA-cleared AI device actually does

Picture a radiology department on a busy night. A narrow tool might check each incoming head CT for signs of a stroke and move suspicious scans up the queue. It doesn’t write the report. It doesn’t talk to the patient. It reorders the pile so a radiologist sees the urgent case sooner.

That modest job description is the point. These systems earn trust by doing one thing inside tight boundaries. Step outside those boundaries and they have nothing useful to offer.

It’s also why they can perform reliably in their lane while general chatbots stumble on open-ended cases. They aren’t reasoning about the whole patient. They’re matching one pattern they’ve seen many times before.

Why “better than radiologists” misses the point

Headlines love a machine-versus-human scoreboard. The real comparison is messier.

A model can look excellent on the data its developers trained and validated it on. Then a new hospital runs it on different scanners and different patients, and accuracy drops. External validation keeps showing that measurable slide, and it’s one of the most consistent caveats in diagnostic AI.

We looked at where each side actually wins in AI vs radiologists. The lesson carries over here. The useful question isn’t who’s smarter. It’s which tasks you can safely hand over, and who checks the work.

Clinician reviewing an AI-flagged CT scan in a dim radiology reading room

What Happens When You Ask a Chatbot to Diagnose You

That Mass General Brigham study deserves more than a headline. The team ran 21 models against 29 standardized clinical vignettes from the MSD Manual. That produced 16,254 responses, which they scored with a new assessment tool called PrIME-LLM. JAMA Network Open published the results on April 14, 2026.

The 80% number, read carefully

The failure rate applies to the differential diagnosis. That’s the early list of possibilities a doctor builds before the test results arrive. It’s the most open-ended step, and it’s where the models struggled most.

They did better later. At the later diagnostic stages, final-diagnosis misdiagnosis rates fell below 40%. The best models topped 90% accuracy on the final diagnosis.

So the fair reading isn’t “AI can’t diagnose.” It’s narrower than that. These models do much better once the uncertain early work is done. Trouble is, that early work is exactly what worried patients want help with at 2 a.m.

Co-author Marc Succi put it plainly: “Despite continued improvements, off-the-shelf large language models are not ready for unsupervised clinical-grade deployment.”

Notice two words there: off-the-shelf and unsupervised. He isn’t writing off the technology. He’s saying the version you open in a browser tab shouldn’t practice medicine on its own. If you’ve wondered how far to trust a chatbot’s read on your symptoms, our piece on whether you can trust an AI to diagnose you goes deeper.

Clinical Decision Support: In the Loop or On the Loop?

Clinical decision support means software that offers suggestions while a clinician keeps the decision. It sounds like the safe middle ground. Keep the doctor in charge, add AI as a second opinion, and you should get the best of both.

The evidence doesn’t back that assumption. Adam Rodman, an associate editor at the New England Journal of Medicine, spoke at the 2025 Penn Medicine Nudges in Health Care Symposium. “Just giving a human an AI system doesn’t inherently improve that human’s performance,” he said. The pair can even do worse than if “the AI system had just run by itself.”

That’s an uncomfortable result. It means workflow design matters as much as the model.

Researchers describe the choice as two setups:

  • In the loop: the physician makes every final call, and the AI advises.
  • On the loop: the AI acts, and the physician monitors and can override.

Today’s clinical AI mostly stays in the first camp. Moving toward the second runs into a wall that has little to do with accuracy.

The liability problem nobody designed for

Say an AI system suggests the wrong diagnosis and a patient suffers for it. Who answers for that? Right now, the clinician does.

According to Penn LDI, the University of Pennsylvania’s health policy research center, no established mechanism exists to shift legal responsibility from the doctor to the AI tool or its maker. That single gap explains a lot. A hospital has little reason to let software act alone while a human still carries all the risk.

So “AI assists, the doctor decides” isn’t only a philosophy. For now, it’s the only arrangement that makes legal sense for the person signing the chart.

Rodman had another line worth keeping on a sticky note: “[The] tech is remarkably powerful, but it is not magic… there is no superintelligence hidden behind the curtain.”

AI in Drug Discovery: Three Molecules Worth Knowing

Drug discovery is where AI in medicine gets exciting. It’s also where the claims need the closest reading. Three examples show what AI actually contributes, and where its contribution stops.

Halicin: searching 107 million molecules

In 2020, the MIT team of Regina Barzilay and James Collins trained a machine-learning model on roughly 2,500 known molecules. Then they pointed it at the ZINC15 database and screened 107 million compounds in silico, meaning on a computer rather than at a lab bench.

The model narrowed that ocean down to 23 candidates. One of them, halicin, turned out to be a new antibiotic. Across the CDC panels of multidrug-resistant pathogens the team tested, it worked against 35 of 36.

The mouse results were striking. Halicin cleared Acinetobacter baumannii skin infections within 24 hours and C. difficile gut infections within four days. The work appeared in Cell in February 2020.

No chemist could screen 107 million molecules by hand. The model’s gift is scale. It reads the whole haystack so humans only test the few needles it points to.

Millions of molecules narrowing to a few antibiotic candidates in AI drug discovery screening

Abaucin: the same trick in under two hours

The sequel matters more than people realize. In 2023, the same approach screened about 7,680 compounds in under two hours and found abaucin. It’s a narrow-spectrum antibiotic that targets drug-resistant Acinetobacter baumannii specifically, as MIT News described.

Why care about a drug that does less? Because it proves the pipeline is reusable. Pick a new pathogen, retrain, rerun. Halicin came from a broad search across a giant library. Abaucin came from a fast, targeted one. Same method, very different jobs.

That reusability is the quiet story in machine learning for pharmaceutical research. The hard part is building a method that works. After that, each new run costs far less.

INS018_055: an AI-designed drug in human trials

Insilico Medicine’s INS018_055 goes a step further. The company describes it as the first drug with both an AI-discovered target and an AI-generated molecular design to reach human trials. It’s aimed at idiopathic pulmonary fibrosis.

The timeline, according to Insilico:

  1. February 2021: preclinical candidate selected, 18 months after the project began
  2. November 2021: first-in-human microdose trial
  3. January 2023: positive Phase I results
  4. February 2023: FDA Orphan Drug Designation
  5. June 2023: first Phase II patients dosed, in a randomized, double-blind, placebo-controlled trial of about 60 patients across roughly 40 sites in the US and China

Insilico says the path from novel target to Phase I took under 30 months, about half the time of traditional drug discovery. That’s a company talking about its own product, so hold the “half” loosely. Still, the milestones are concrete dates, not marketing adjectives.

What AI speeds up, and what it can’t skip

Look at that timeline again. AI compressed the discovery phase. It didn’t let the drug skip Phase I or Phase II. INS018_055 went through the same regulatory gauntlet as any other drug, with real patients and a placebo arm.

That’s the honest shape of AI drug discovery today. It finds better starting points faster. Proving a molecule is safe and effective in people still takes as long as it takes. For a look at how one AI lab is pitching its own role here, see our take on Anthropic and AI drug discovery.

Where AI in Medicine Still Falls Short

Put the evidence side by side and the weak points line up:

  • Open-ended diagnosis. General models failed differential diagnosis more than 80% of the time, across 16,254 responses from 21 models. That’s no anecdote you can wave away.
  • Generalization. Models that shine on their home data tend to slip at new hospitals with different scanners and patients.
  • Human-AI teamwork. Adding a doctor to an AI system doesn’t automatically beat either one alone. The workflow decides.
  • Liability. Nothing yet moves legal responsibility onto the tool or its maker, so clinicians stay firmly in charge.
  • Clinical trials. AI can speed up discovery, but every candidate still has to prove itself in human studies.

Look at how few of these come down to raw accuracy. Liability and workflow gaps hold back more autonomous clinical AI as much as any error rate does. Most explainers skip that part, and it’s the part hospitals actually wrestle with.

AI in Medicine Today: Confirmed vs. Still Projected

The scorecard, as of 2026.

Confirmed and in use:

  • More than 1,600 FDA-authorized AI-enabled devices, with imaging tools built for narrow, specific tasks
  • AI-found antibiotic candidates, halicin and abaucin, that showed real activity against drug-resistant bacteria
  • An AI-designed drug, INS018_055, that reached Phase II trials with human patients

Still projected or unresolved:

  • General-purpose AI as an autonomous diagnostician
  • Clinical decision support that acts “on the loop” instead of advising “in the loop”
  • A legal framework that lets responsibility follow the software
  • Reliable performance when a diagnostic model moves to a hospital it has never seen

The first list is shorter than the hype suggests. It’s also far more solid than the skeptics admit.

The Question Worth Asking Next Time

Here’s our read. The most useful things AI in medicine has done so far are quiet ones. It sorts scans. It searches chemical space. It shortens the list. The loud version, a machine that hears your symptoms and tells you what’s wrong, isn’t here yet, and the 2026 data says so without much ambiguity.

So if a clinic tells you AI helped with your care, don’t ask whether the AI is smart. Ask what job it had, and who checked its work. Those two answers tell you nearly everything.

The bigger question is one this site keeps circling back to. When machines get good at a narrow slice of human expertise, which parts should stay human on purpose? In medicine, it might be the moment someone sits with real uncertainty and still takes responsibility for the call.

Frequently Asked Questions
How is AI used in medicine today?
Mostly through narrow, task-specific tools. The FDA has authorized more than 1,600 AI-enabled medical devices, such as software that flags a possible stroke on a CT scan or helps triage mammograms. AI also speeds up early drug discovery by screening huge libraries of molecules for promising candidates.
Is AI accurate in medical diagnosis?
It depends on the tool and the task. Narrow, FDA-cleared tools do one job within tight limits. General chatbots are far weaker at open-ended diagnosis: a 2026 Mass General Brigham study in JAMA Network Open found 21 large language models failed to produce an appropriate differential diagnosis more than 80% of the time, though the best models topped 90% accuracy on the final diagnosis.
How does AI help discover new drugs?
AI searches chemical space at a scale no lab can match. MIT researchers screened 107 million compounds on a computer to find the antibiotic halicin, and a later run screened about 7,680 compounds in under two hours to find abaucin. Insilico Medicine used AI to pick both the target and the molecular design of INS018_055, which reached Phase II trials.
Can AI replace doctors or radiologists?
Not on current evidence. Off-the-shelf language models still struggle with early diagnosis, pairing a doctor with AI doesn't automatically improve results, and there's no established way to shift legal responsibility from the clinician to the software. Today's clinical AI assists a physician who makes the final call.
What are the risks or limitations of AI in healthcare?
The main ones are weak open-ended diagnosis in general-purpose models, accuracy that drops when a model moves to a new hospital or scanner, poorly designed human-AI workflows, and unresolved liability. AI can speed up drug discovery, but every candidate still has to pass human clinical trials.
How long does AI drug discovery take compared to traditional methods?
Insilico Medicine says INS018_055 went from a novel target to Phase I in under 30 months, which it describes as about half the time of traditional drug discovery. That's the company's own figure, and the drug still had to go through standard Phase I and Phase II trials.
What is clinical decision support software?
It's software that offers suggestions to a clinician while the clinician keeps the decision. Researchers describe two designs: 'in the loop', where the physician makes every final call, and 'on the loop', where the AI acts and the physician monitors and can override it.
Is AI used in hospitals right now, or is it still experimental?
Both. More than 1,600 FDA-authorized AI devices are real, regulated products in clinical use. General-purpose AI acting as an unsupervised diagnostician is still experimental, and researchers say off-the-shelf chatbots aren't ready for that role.