Prompt20
All posts
healthcaremedical-imagingclinical-decision-supportdrug-discoverydiagnosticsregulationverticalevergreen

AI in Healthcare: What It Actually Does

A concepts-first tour of where AI is real in medicine and where it's marketing. Clinical decision support, medical imaging and radiology triage, ambient scribes and documentation, drug discovery, diagnostics, patient triage chatbots — plus the parts that stay hard: regulatory clearance, liability, validation on real populations, bias in training data, and why 'FDA-cleared' doesn't mean what you think. Built to outlast the vendor churn.

By Prompt20 Editorial · 30 min read

Here is the honest one-line version: AI in healthcare is already doing real, boring, valuable work — measuring things, flagging things, and typing things — and it is mostly not doing the thing the headlines promise, which is replacing a doctor's judgment. The gap between those two is not about model quality. A model can score better than clinicians on a benchmark and still be years from a clinic, because medicine gates deployment on evidence and liability, not on how impressive the demo looked.

That single fact explains almost everything about this field. If you understand why a slightly-better spreadsheet ships in a hospital faster than a genius diagnostic model, you understand medical AI. This post separates the clinically validated uses from the marketing, then explains the machinery — regulatory clearance, liability, population validation, and dataset bias — that decides which side of that line any given product lands on. Where I name a product category, treat it as a snapshot; the categories outlast the vendors.

Table of contents

Key takeaways

  • The winners are narrow and measurable. AI is genuinely deployed where the task is well-defined and the output is checkable: imaging triage, ambient documentation, sepsis and deterioration alerts, retinal screening. Broad "AI doctor" systems are not.
  • Deployment is gated by evidence, not accuracy. A model that beats clinicians on a test set can still be undeployable, because medicine requires prospective validation on the actual population, integration into workflow, and someone willing to hold the liability.
  • "FDA-cleared" is weaker than it sounds. Most medical AI reaches the market through clearance pathways that ask "is it substantially similar to an existing device?" — not "does it improve patient outcomes?" Clearance is a floor, not a gold star.
  • Bias is a data problem you can't model your way out of. A tool validated on one hospital's population can silently fail on another's. The failure is invisible until you audit outcomes by subgroup.
  • The scribe is the killer app. The least glamorous use — turning a conversation into a clinical note — is the one with the clearest ROI, because it attacks documentation burnout without touching a diagnosis.
  • Liability is the real bottleneck. Until it's clear who is responsible when the model is wrong, most systems stay in "assistive" mode with a human legally on the hook.

The mental model: assist, don't decide

Sort every medical AI product into one question: does a licensed human stay legally responsible for the output? Almost everything shipping today answers "yes." The AI reads the scan, but the radiologist signs the report. The AI drafts the note, but the physician attests to it. The AI flags a deteriorating patient, but a nurse decides what to do.

This is not a temporary limitation waiting for better models. It is the structural equilibrium of a field where being wrong has a body count and a courtroom. "Assistive" tools clear regulatory and legal hurdles because they leave a human accountable. "Autonomous" tools — where the machine's output is the decision — face a vastly higher bar, and only a handful have cleared it, in extremely narrow settings like diabetic retinopathy screening where the task is bounded and the downside of a miss is a referral, not a death.

So when a vendor says their model "diagnoses X better than doctors," the correct follow-up is not "how much better?" It is "who signs the chart?" If the answer is still a human, you're buying a faster human, not a replacement — and that's fine, that's often exactly what a health system needs.

There is a deeper reason "assist" is the equilibrium and not a way-station. The value of a decision in medicine is not just its accuracy; it is its accountability. A diagnosis that turns out wrong is not merely a data point — it triggers a chain of consequences (a wrong treatment, a delayed one, a lawsuit, a regulatory inquiry) that a legal system has spent a century learning to attach to a named, licensed, insured human being. A model has no license to revoke, no board to answer to, and no malpractice policy. Until the surrounding institutions — insurers, licensing boards, courts, hospital risk departments — build machinery to absorb a machine's mistakes, the machine cannot be the endpoint of a decision. It can only feed a human who is that endpoint. This is why "assistive" is not a marketing hedge or a phase we grow out of; it is the shape the field takes when you hold accountability constant and let the technology improve underneath it. Better models make the human faster and better-informed. They do not, by themselves, move the accountability.

A useful corollary: the more autonomous a tool claims to be, the narrower its safe domain has to be. The handful of genuinely autonomous systems that exist work in tasks so bounded that the entire decision space can be enumerated, the failure mode is a recoverable "refer to a human," and the population is well-characterized. Autonomy and breadth trade off against each other. Anyone selling you both at once — broad and autonomous — is selling a demo, not a deployed product.

The six categories of medical AI, honestly sorted

Almost every product in this space falls into one of six categories, and they are not equally mature. Sorting them by how close each is to real, defensible deployment is more useful than any vendor's roadmap, because the ranking is driven by structural properties of the task, not by how hard anyone is working.

  1. Medical imaging and diagnostics. The most mature. A bounded input (an image), a checkable output (a flag, a measurement, a reordered worklist), and an expert who can verify the result on the spot. Deployment is real and growing, almost entirely in an assistive posture.
  2. Clinical documentation and ambient scribes. The category with the clearest and fastest return, because it attacks an administrative burden rather than a clinical judgment, and the output is proofread by the clinician at the point of use. Low clinical risk, high economic payoff.
  3. Clinical decision support and early warning. Real and widely deployed, but chronically limited by alert fatigue rather than model accuracy. The bottleneck is calibration and workflow, not intelligence.
  4. Drug discovery and molecular modeling. Genuine scientific value, but it speeds only the cheap, early front of a decade-long pipeline. It changes research productivity, not this year's pharmacy shelf.
  5. Operational and administrative automation. Coding, billing, prior authorization, scheduling, capacity planning. Quiet, unglamorous, and often the highest-ROI use in the building, because the error surface is financial rather than clinical and the tolerance for automation is correspondingly higher.
  6. Patient-facing chatbots and symptom checkers. The loudest category and the least defensible. Fluency is mistaken for competence, the population is unbounded, the output feeds directly into a frightened non-expert with no verification layer, and nobody has solved who is liable when it's wrong.

Notice the pattern. Maturity tracks three things: how bounded the task is, how immediately checkable the output is, and how low the clinical stakes of an error are. Categories 1 through 3 satisfy those constraints; 4 and 5 sidestep clinical stakes entirely by living in research and back-office workflows; category 6 violates all three at once, which is exactly why it generates the most hype and the least trustworthy product. Keep this ordering in your head and most vendor pitches sort themselves.

Where it's real

Medical imaging and triage

Radiology is the flagship because the task fits the technology. An image goes in; a bounded, checkable output comes out — a bounding box around a suspected bleed, a flag on a nodule, a worklist reordered so the likely stroke gets read first. Crucially, the output is verifiable by the same expert who would have done the job. That verifiability is what makes it deployable.

The dominant real-world use is triage and prioritization, not autonomous reads. The model doesn't tell you the answer; it changes the order in which a human looks, or draws attention to a region. This is a smaller claim than "AI reads your X-ray," and it's exactly why it works: it improves throughput and catch rates without asking anyone to trust the machine's final judgment.

Ambient documentation (the actual killer app)

The single most successful category is the least cinematic: ambient clinical scribes that listen to a visit and draft the note. This works because the underlying capability — turn speech into structured, summarized text — is mature, and because the output is low-stakes and immediately checked. The physician reads and edits the draft; errors are caught at the point of use, not months later in a bad outcome.

The economics are also unambiguous. Documentation burden is a leading driver of clinician burnout, and every minute a doctor spends typing is a minute not spent with a patient. A tool that reliably saves that time sells itself. This is the same transformer-and-transcription stack behind consumer assistants; if you want the mechanics of the conversational layer, see how AI chatbots work. The healthcare twist is that the note becomes part of a legal and billing record, so accuracy and the "human attests" step matter more, not less.

Clinical decision support and early warning

Hospitals run background models that watch the electronic health record for patterns — early signs of sepsis, patient deterioration, readmission risk. These clinical decision support systems are useful precisely because they're modest: they raise an alert, a human investigates. The hard part here is rarely the model; it's alert fatigue. A sepsis model with too many false positives gets ignored, silenced, or clicked past, and then it might as well not exist. The engineering problem is calibration and workflow integration, not raw predictive power.

Diagnostics and screening

In narrow, high-volume screening — retinal images for diabetic eye disease, certain dermatology and pathology tasks — AI genuinely helps, especially where specialists are scarce. These are the cases closest to "autonomous," and they share a profile: a single well-defined question, a large labeled dataset, and a fallback (refer to a human) that makes a miss recoverable. The further a task drifts from that profile, the more it stays assistive.

Drug discovery and the back office

Two more real categories rarely make the patient-facing headlines. Drug discovery uses models to predict protein structure, screen candidate molecules, and prioritize experiments — compressing the early, cheap part of a pipeline whose expensive, slow part (clinical trials) AI cannot shortcut. It's real value, but it moves the front of a decade-long process, so don't expect it to visibly change your pharmacy this year. And the administrative back office — coding, billing, prior-authorization paperwork, scheduling — is where a lot of quiet ROI lives, because errors are financial rather than clinical and the tolerance for automation is higher.

Use case What it actually outputs Human still decides? Why it ships
Imaging triage Reordered worklist, region flag Yes (radiologist) Output verifiable by the same expert
Ambient scribe Draft clinical note Yes (physician attests) Low stakes, checked at point of use
Sepsis / deterioration alert A flag to investigate Yes (clinician) Modest claim, fits existing workflow
Retinal / narrow screening Refer / no-refer Sometimes (semi-autonomous) Bounded task, recoverable miss
Drug discovery Ranked candidates Yes (scientists, trials) Speeds cheap front of pipeline
Coding / billing Structured claims Partly Financial, not clinical, error surface

Why imaging was the early win, and language is the risky frontier

It is worth pausing on why medical imaging arrived first, because the reasons generalize to every future category and explain why the newest, most impressive capability — large language models reasoning over a clinical picture — is also the most dangerous to deploy.

Imaging won because the task has four properties that make machine learning tractable and safe at once. First, it is narrow: "is there a nodule in this region of this chest CT?" is a single, well-posed question, not an open-ended request for judgment. Second, it is labeled: decades of radiology practice have produced huge archives of images with expert annotations and, crucially, downstream outcomes to check them against. Third, it is measurable: you can state performance as sensitivity and specificity on a held-out set, and a radiologist can look at the model's output and immediately see whether it is right. Fourth, the failure is legible — a false positive is a second look, a false negative is a miss that the same expert would understand. A task that is narrow, labeled, measurable, and legible is a task you can validate, regulate, and insure. That is the whole reason imaging shipped.

Language-model clinical reasoning inverts every one of those properties, which is why it is simultaneously the most seductive demo and the least deployable product. The task is broad — "what's wrong with this patient?" has no bounded answer space. The training signal is weakly labeled — text about medicine is abundant, but it is not the same as validated ground truth about outcomes, and the internet's medical text is a mix of guidelines, folklore, marketing, and error. Performance is hard to measure in any way that predicts real-world safety: a model can produce a fluent, well-organized, entirely wrong answer, and no automatic metric reliably catches it. And the failure mode is illegible — a confidently phrased mistake looks exactly like a confidently phrased correct answer, so the error hides inside the fluency instead of announcing itself.

This is the crux of the whole field's near-term shape. The capability that captured the public imagination — a system that talks like a clinician — is precisely the capability whose errors are hardest to detect and whose task is hardest to bound. "Passes the medical licensing exam" is a headline about the easy version of the problem: multiple-choice questions with a known correct answer, drawn from the well-trodden center of medical knowledge. Practicing medicine is the hard version: open-ended, messy, full of missing information and atypical presentations, and unforgiving of confident wrongness. The distance between those two is not a matter of a few more model generations. It is the distance between a benchmark and a patient, and the rest of this post is about the machinery built to keep them apart.

Where it's mostly marketing

The hype clusters around a few recurring claims. The general "AI doctor" chatbot that takes symptoms and returns a diagnosis is the biggest one. Symptom-checkers are old, and dressing them in a language model makes them more fluent, not more correct — arguably more dangerous, because fluency reads as confidence. A model that hallucinates a plausible-sounding differential diagnosis is worse than a blank page, because it anchors both patient and clinician. The same output that would be a charming error in a trivia bot is a liability event in a clinic.

"Personalized medicine, powered by AI" is a real research direction and a heavily abused marketing phrase. Genuine personalization requires longitudinal, high-quality data that most systems don't have in a usable form, plus validation that the personalization actually improves outcomes rather than just producing more granular guesses. "Predicts disease years in advance" deserves particular skepticism: prediction on a retrospective dataset is easy; prospective, actionable prediction that changes what a doctor does — and doesn't just generate anxiety and unnecessary tests — is rare and hard-won.

The tell, in every case, is the same: does the claim come with prospective evidence on a real population and a plan for who's accountable, or just a benchmark number and a demo? Benchmarks measure the easy 90%. Medicine lives and dies in the last 10%, on the patients who don't look like the training set.

Why medicine gates on evidence, not model quality

This is the core of the post, so slow down here. Four gates decide whether a good model becomes a deployed tool.

Regulatory clearance — and what "cleared" doesn't mean

Most medical AI reaches market through clearance pathways built for devices, where the central question is often "is this substantially equivalent to something already on the market?" — not "does this improve patient outcomes in a trial?" That means "FDA-cleared" (or its equivalent elsewhere) tells you a product met a safety-and-similarity bar, not that it makes patients healthier. Many cleared tools were validated on retrospective data, sometimes from a handful of sites, sometimes without published prospective outcomes at all.

This isn't a scandal; it's a category error the marketing exploits. Clearance is necessary and it is a floor. It is not evidence of clinical benefit, and it says little about whether the tool works on your population. Read "cleared" the way regulators intend it, as covered in AI regulation explained: a permission to sell, not a verdict on efficacy.

Prospective validation on the real population

A model trained and tested on one health system's data is a hypothesis, not a product. The population that walks into a rural clinic differs from an academic medical center's — different demographics, different disease prevalence, different scanners, different labeling habits. Distribution shift is the quiet killer: performance that looked excellent in the paper degrades silently in the field, and nobody notices until an outcome audit — if one ever happens. This is why the durable question about any medical model is "validated where, on whom, measured how, and does it still hold at my site?"

Bias you can't model your way out of

If the training data underrepresents a group, or encodes historical inequities in care, the model inherits and can amplify them — a risk-score that reads "cost of care" as a proxy for "severity of illness" will systematically under-serve populations who historically received less care. You cannot fix this purely with a cleverer architecture, because the flaw is in what the data means, not in how well the model fits it. The only reliable defense is auditing outcomes by subgroup after deployment, which most buyers don't budget for. It's the same failure mode you see across training data debates, with mortal stakes.

Liability — the bottleneck that keeps humans in the loop

Finally, the question that decides the field's shape: when the model is wrong, who pays? A physician who follows a bad AI recommendation may still be liable; a vendor typically disclaims responsibility for clinical decisions; a hospital sits in between. Until that allocation is settled — by statute, case law, or contract — the rational posture is to keep a licensed human on the hook and label the AI "assistive." This is not timidity. It's the same logic that makes assistive tools clear regulatory gates faster: leaving a human accountable is how you make a novel technology insurable. Liability, not model quality, is why the "autonomous AI doctor" stays perpetually five years away.

The evidence bar: what clinical validation actually requires

If you take one durable idea from this post, make it this: in medicine, a demo is worth almost nothing and validation is worth almost everything, and the two are separated by a wide, expensive, deliberately unglamorous gap. Understanding what fills that gap is what lets you tell a real product from a slide deck.

A demo shows a model producing a good answer on a case someone chose. Validation asks a harder, colder question: on a defined population, measured against a pre-specified endpoint, does using this tool change what happens to patients — and does that hold up when someone who wasn't rooting for the result checks? The ladder from one to the other has rungs, and most products stop partway up:

  • A benchmark score — the model does well on a test set. This is the floor. It tells you the model learned the training distribution; it tells you nothing about a real clinic.
  • Retrospective validation — the model is run against historical data it didn't train on. Better, but the past is a forgiving judge: the cases are already resolved, the population is whatever the archive happened to contain, and it's easy to tune toward a flattering number.
  • Prospective validation — the tool runs on real patients in real time, and its performance is measured going forward. This is where most impressive models quietly falter, because the live population never matches the archive exactly.
  • A comparative or randomized study — some patients' care involves the tool and some doesn't, and outcomes are compared. This is the gold standard, it is expensive and slow, and it is rare. Most deployed medical AI has never faced it.

The reason "passes the medical exam" is not evidence of safety is that a licensing exam sits below the first rung. It is a benchmark with multiple-choice answers drawn from textbook medicine. It measures recall of consensus knowledge, not the ability to act safely under uncertainty with a real, atypical patient and incomplete data. A tool can ace every exam ever written and still be unvalidated in the only sense that matters, which is prospective evidence that using it helps and doesn't harm. When you evaluate a claim, locate it on this ladder. The height of the rung, not the size of the number, is the signal — and it should be read alongside how regulators actually license these tools, covered in AI regulation explained.

Hallucination and liability in a life-critical setting

Everywhere else on this blog, a model that invents a plausible-but-false answer is an annoyance you learn to work around. In medicine it is a category of harm. It is worth being precise about why the same failure mode changes character when the stakes are a body instead of a paragraph.

A language model does not know things; it produces text that is statistically likely given its training and the prompt. That process is indifferent to truth — it optimizes for plausibility, and plausibility and correctness usually coincide, which is exactly what makes the exceptions dangerous. When they diverge, the output is a hallucination: fluent, well-formed, confident, and wrong. In a clinical setting three things make this uniquely hazardous. First, the reader is often not equipped to catch it — a worried patient, or a clinician outside their specialty, has no independent basis to reject a confident false claim. Second, fluency reads as authority — the more coherent the wrong answer, the more it anchors the next decision, a bias that operates on experts too. Third, the error compounds down a chain — a hallucinated finding becomes a note, becomes a referral, becomes a test, becomes a treatment, and each step launders the original mistake into something that looks more established.

You cannot eliminate this behavior, but you can engineer around it, and the techniques matter more here than anywhere else: grounding outputs in retrieved source documents, constraining the model to structured tasks with checkable outputs, forcing citations, and — above all — keeping a qualified human as the mandatory verification step. The general playbook is worth internalizing; see how to reduce AI hallucinations. But note the honest limit: mitigation reduces frequency, it does not deliver a guarantee, and medicine is a domain where a rare confident error is not an acceptable tax. This is the technical fact that underwrites the legal posture. Because the failure is undetectable from the output alone, the only defensible design keeps a licensed human accountable — which folds hallucination straight back into the liability problem. The machine can be wrong invisibly, so a human who can be held responsible has to own the result. Model quality lowers the error rate; it never removes the need for the accountable human, and that is why it never removes the human.

Bias and equity in clinical data

The bias problem in medical AI deserves its own treatment because it is the failure most likely to be invisible to the people deploying the tool and most damaging to the people it fails. A model can post excellent aggregate numbers and still be quietly harming a subgroup, and nothing in the top-line metric will tell you.

The mechanism is that a model learns the world its data describes, including that world's inequities. If a group was historically under-diagnosed, the labels the model trains on encode that under-diagnosis as ground truth, and the model learns to reproduce it. If a group received less care, and a risk score is trained on cost of care as a proxy for severity of illness, the model will systematically rate that group as healthier than they are — not from malice or a coding bug, but because the proxy meant something different for them than the modelers assumed. This is the crucial and counterintuitive point: you cannot fix this with a better architecture, more parameters, or more compute, because the flaw is in what the data means, not in how well the model fits it. A more powerful model fits the biased signal more faithfully. It can make the problem worse.

The failure is also structurally invisible. Distribution shift and subgroup bias do not announce themselves; performance simply degrades for some patients while the average stays fine, and the average is what gets reported. The only reliable defense is to audit outcomes by subgroup after deployment — to measure, on your actual patients, whether the tool performs equitably across the populations you serve — and that is precisely the step most buyers never budget for and most vendors never volunteer. It is the same failure family that appears across AI bias and fairness and in the training-data debates, except here the cost of getting it wrong is measured in missed diagnoses and worse outcomes for the people already least well served.

Privacy, HIPAA, and the data that trains medical AI

Every capability in this post runs on data that is, by definition, among the most sensitive a person generates, and the rules governing that data shape what medical AI can and cannot do far more than the models do. This is not a compliance footnote; it is a load-bearing constraint on the whole field.

Health data lives under a stricter regime than ordinary personal data — in the United States, chiefly HIPAA, with analogous frameworks elsewhere — and that regime imposes real friction on the AI pipeline at three points. Training is constrained because you cannot freely pool identifiable records to build a model; the datasets have to be de-identified, governed, and often confined to a single institution, which is one structural reason models trained at one hospital struggle to generalize to another. Deployment is constrained because sending a patient's information to a third-party model provider is a data-sharing event with legal weight, which is why serious clinical tools run under specific contractual arrangements rather than a consumer API, and why "just paste the chart into a chatbot" is a genuine privacy breach, not a shortcut. Retention and secondary use are constrained because data collected to treat a patient cannot be silently repurposed to train a product without crossing consent and governance lines.

There is a specific trap worth naming for anyone tempted to use a general-purpose consumer chatbot in a clinical context: the same chatbot privacy questions that matter for ordinary use — where does the text go, is it logged, is it used for training, who can see it — become legal exposure when the text is protected health information. De-identification is also weaker than it sounds; rich clinical narratives can be surprisingly re-identifiable when combined with other data. The durable lesson is that data governance is not the boring part of medical AI you get to skip. It is one of the main reasons the field moves at the pace it does, and a tool's answer to "where does the data go and under what agreement" tells you as much about whether it is deployable as any accuracy number.

Workflow integration and the human-in-the-loop reality

A model that is accurate, validated, and compliant can still be worthless, and the reason is the least glamorous in the whole field: it doesn't fit the way clinicians actually work. Workflow integration is where more good medical AI dies than at any regulatory gate, and it is invisible in every demo because a demo has no workflow.

Consider what "in the loop" really demands. The output has to arrive at the moment the decision is made, inside the system the clinician already uses, without adding clicks to a day that is already a war against clicks. It has to be legible — a number with no explanation is an interruption, not help. And it has to earn trust calibrated to its reliability, which is a two-sided failure: a tool trusted too little is ignored, and a tool trusted too much invites automation bias, where a human rubber-stamps the machine and stops thinking. The single most common way a clinically sound model fails in the field is alert fatigue: fire too many low-value alerts and clinicians learn to dismiss all of them, including the rare true one, and the tool becomes worse than nothing because it trained its users to ignore it. This is a calibration and design problem, not an intelligence problem, and no amount of model quality solves it — a more accurate model that still cries wolf too often gets silenced just the same.

The "human in the loop" is also not a rubber stamp you can wave at the liability problem to make it go away. For the human to be a real safeguard, they need the information, the time, and the authority to actually override the machine — and if the system is designed so overriding is slow, or the human is too rushed to engage, or the interface nudges toward acceptance, then "human oversight" is a fiction that provides legal cover without providing safety. Genuine human-in-the-loop design treats the clinician as the decision-maker the tool serves, not as a formality standing between the model and the patient. Getting that relationship right is harder, and matters more, than another point of accuracy.

Hype versus real: a field guide

Pulling the threads together, here is the compressed field guide — the recurring claims and what's actually underneath them. The pattern is consistent enough to be predictive.

The claim What's real underneath The tell it's hype
"AI reads your scans" Assistive triage and flagging; a radiologist still signs No mention of who signs the report
"AI doctor / symptom checker" A more fluent symptom checker Fluency sold as diagnostic accuracy
"FDA-cleared, so it's proven" Met a safety/similarity bar "Cleared" used as a synonym for "effective"
"Beats doctors on diagnosis" Beats them on a curated test set A benchmark with no prospective, real-population evidence
"AI-powered personalized medicine" A real research direction No validation that personalization improves outcomes
"Predicts disease years early" Retrospective prediction is easy No evidence it changes what a clinician does
"Autonomous AI diagnosis" A few narrow, bounded screening tasks Breadth and autonomy claimed together
"Saves clinicians hours" (scribes) Genuinely true and the clearest ROI Almost the only claim that usually holds up

The single meta-tell across the whole table: real medical AI comes with prospective evidence on a defined population and a clear answer to who is accountable. Hype comes with a benchmark number, a polished demo, and silence on both. When you can't find the evidence and the accountability, you've found the marketing.

How to evaluate a medical AI claim

You don't need a medical degree to smell-test a product. Ask, in order:

  1. Who signs the chart? If a human still attests, it's assistive — judge it as a productivity tool, not a clinician.
  2. Cleared for what, and validated how? Distinguish a safety clearance from published prospective outcome evidence. Ask for the population and the endpoints.
  3. Validated on whom — and does that match my patients? A number from one academic center is not a guarantee at a community clinic.
  4. What's the false-positive cost? In alerting systems, over-triggering causes alert fatigue and silent failure. Calibration beats raw accuracy.
  5. Who's liable when it's wrong? If the contract disclaims all clinical responsibility, you're the accountability layer. Price that in.
  6. Is there a subgroup audit? If nobody's measuring performance across populations after deployment, bias is going undetected by design.

Answer those and you'll correctly sort most of the market without ever looking at a benchmark. The benchmark was never the point.

The honest near-term picture

Strip away both the hype and the reflexive cynicism and a clear, durable picture remains — one that has been roughly stable for years and is likely to stay stable, because it is set by the structure of the field rather than by the pace of model improvement.

Medical AI will keep getting genuinely better at the things it is already good at, and it will keep spreading, mostly invisibly, through the parts of healthcare where it is safe: reading and flagging images, drafting the notes, watching the monitors for early warnings, ranking drug candidates, and grinding through the administrative paperwork that consumes so much of a clinician's day. This is not a small thing. Attacking documentation burden and administrative waste alone would be a meaningful improvement to a system where burnout is a workforce crisis and paperwork is a tax on every encounter. The value here is real, and it is accumulating faster than the headlines, precisely because it is boring.

At the same time, the thing the headlines keep promising — an autonomous system that replaces a clinician's judgment across the open-ended reality of practice — is not arriving on the current trajectory, and the reason is not that the models are not smart enough. It is that the field gates deployment on evidence, liability, validation, equity, privacy, and workflow, and none of those gates opens when a model gets more capable. A smarter model still needs prospective validation on your population. It still needs someone to hold the liability when it is wrong. It still needs to not silently fail a subgroup, to respect the data it runs on, and to fit into a clinician's day. Those are institutional and structural problems, and they move on institutional time.

So the durable expectation is asymmetric: expect steady, compounding, unglamorous improvement in the assistive layer, and expect the autonomous layer to stay narrow, bounded, and slow to widen — advancing one carefully validated task at a time, in settings where a miss is recoverable. The most useful posture toward any medical AI claim is neither the booster's nor the skeptic's, but the auditor's: assume the boring uses are quietly real, assume the dramatic ones are marketing until shown prospective evidence and a named accountable party, and judge every product by who signs the chart. That test has outlasted a decade of vendor cycles. It will outlast this one too.

FAQ

Does "FDA-cleared" mean an AI tool improves patient outcomes? No. Clearance generally means a product met a safety bar and is considered substantially similar to an existing device — often on the basis of retrospective data. It's permission to sell, not proof of clinical benefit. Ask separately for prospective outcome evidence on a population like yours.

Will AI replace doctors? Not on the current trajectory, and the reason isn't model quality. Medicine gates deployment on evidence and liability. Almost every shipping system is "assistive," with a licensed human legally responsible for the output. Until it's clear who pays when the model is wrong, autonomous diagnosis stays confined to a few narrow, bounded tasks.

What's the most valuable medical AI use today? Ambient clinical documentation — tools that listen to a visit and draft the note. It has the clearest return on investment because it attacks documentation burnout, produces a low-stakes output that's checked immediately by the clinician, and never touches the diagnosis itself.

Why can a model that beats doctors on a test still fail in a clinic? Because a test set is not a population. Real patients differ from the training data in demographics, disease prevalence, and equipment, so performance degrades under distribution shift — often silently, until an outcome audit catches it. Benchmarks measure the easy majority of cases; medicine is decided by the hard minority.

Is AI good at diagnosing from symptoms I type in? Be skeptical. Language models make symptom-checkers more fluent, not more accurate, and fluent wrongness is dangerous because it reads as confidence. A confidently hallucinated diagnosis anchors both patient and clinician toward the wrong conclusion. Treat these tools as prompts for a conversation with a professional, not answers.

How does bias get into medical AI, and can better models fix it? It enters through the data: underrepresented groups, or historical inequities encoded as proxies (like using past healthcare spending as a stand-in for illness severity). You can't fix it with a better architecture because the flaw is in what the data means, not how well the model fits it. The only reliable defense is auditing outcomes by subgroup after deployment. The same failure family shows up across AI bias and fairness generally, but with mortal stakes here.

Does "passes the medical licensing exam" mean an AI is safe to practice medicine? No, and the gap is larger than it sounds. A licensing exam is a benchmark of multiple-choice recall drawn from consensus, textbook medicine — it sits below even retrospective validation on the evidence ladder. Practicing medicine is open-ended, works from incomplete information, is full of atypical presentations, and punishes confident wrongness. A model can ace every exam and still have no prospective evidence that using it helps real patients and doesn't harm them, which is the only thing "safe to practice" could mean.

Is it safe to paste medical records or symptoms into a general chatbot like ChatGPT or Claude? Treat it as a privacy exposure, not a shortcut. A patient's health information is legally protected, and pasting it into a consumer chatbot is a data-sharing event — the same chatbot privacy questions (where does the text go, is it logged, is it used for training) become legal risk when the text is protected health information. Serious clinical tools run under specific data agreements for exactly this reason. For your own questions, strip identifying details and treat any answer as a prompt for a conversation with a professional, never a diagnosis.

Why do good medical AI tools fail even after they're approved and accurate? Usually because they don't fit the clinician's workflow. If the output doesn't arrive at the moment of decision, inside the existing system, without adding clicks, it gets ignored — and if an alerting tool fires too many false positives, clinicians learn to dismiss all its alerts, including the true ones. This is alert fatigue, and it's a calibration and design problem that no amount of model accuracy solves. Fit and calibration kill more deployed medical AI than any regulatory gate.