Agentic AI Arrives Where Mistakes Are Measured in Patients
Medical artificial intelligence has traditionally been judged like a diagnostic test. A model is trained on a dataset, benchmarked against clinicians on a held-out sample, and scored on sensitivity, specificity, or some other measure of correctness. That entire approach rests on an assumption: the model produces an output, and a human being interprets it.
Agentic AI breaks that assumption. These systems are built to pursue goals across multiple steps — retrieving records, ordering tests, drafting documentation, scheduling follow-ups, and revising a plan as new information arrives. A paper published in Nature Medicine on 2 October 2026, titled "The missing links in agentic AI autonomy," takes aim at the gap this creates. Its subject is not model accuracy but trust: specifically, the operational and decisional trust that must exist before autonomous systems can be allowed to participate meaningfully in care.
The framing is a deliberate departure from the way the field usually argues about AI safety. Most discussions treat trust as a downstream consequence of performance — build a better model, and confidence will follow. This paper suggests the relationship runs the other way as well: without structures that make trust measurable and earned, even excellent models cannot be safely delegated to.
Two Kinds of Trust, Not One
The paper's most useful move is to split trust into two categories. The distinction matters because the two fail in very different ways, and conflating them hides where the real problems sit.
Operational trust
Operational trust is about whether a system does what it claims to do, reliably, in the environment where it is actually deployed. Can it reach the right data sources? Does it fail gracefully when a record is missing, a query times out, or an interface changes? Does it degrade in ways that a supervising human can see, or does it quietly produce confident nonsense? These look like engineering questions, but in a clinical setting they carry clinical weight. A system that silently loses a lab result is not merely buggy; it is potentially harmful.
Decisional trust
Decisional trust is narrower and harder to establish. It concerns whether the judgments a system makes — which action to take, in what sequence, and critically, when to stop — deserve to be relied upon. A system can be operationally flawless and still make decisions a clinician would reject outright. Conversely, a system with shaky plumbing may reason well within a narrow window where nothing breaks. Treating these as one variable, as much of the current conversation does, obscures the fact that fixing one does nothing for the other.
Why Autonomy Changes the Risk Calculus
A conventional clinical AI tool sits inside a workflow and waits. It answers a question and the answer is reviewed. Agentic systems do something different: they hold a thread. They make intermediate choices, some of which are consequential, and those choices compound. The person nominally in charge may see only the summary at the end.
That compounding is what makes the trust question urgent rather than philosophical. Every additional step a system takes without a checkpoint multiplies the number of places where an error can enter and the number of downstream decisions that inherit it. Human oversight that works for a single prediction does not automatically scale to a chain of ten actions taken over an afternoon.
Healthcare sharpens this further because the domain is unusually unforgiving of the failure modes agentic systems exhibit most readily: overconfidence, plausible-sounding fabrication, and difficulty recognizing the boundary of their own competence. A chatbot that invents a citation is embarrassing. A system that invents a medication interaction and acts on it is something else entirely.
The Missing Links
If the paper's argument is that autonomy requires more than capability, the natural next question is what is actually absent. Several gaps recur across the terrain the paper covers.
- Evaluation that matches the task. Static benchmarks measure single answers, not sequences of decisions. There is no mature practice for scoring whether an agent's plan was reasonable, only whether its final output matched a reference.
- Auditability of reasoning. When a system takes an action, the record of why it did so is often thin or reconstructed after the fact. Without trustworthy traces, oversight becomes theater.
- Well-defined escalation. Agentic systems need clear, reliable points at which they hand control back to a human. Too many handoffs make the system useless; too few make it dangerous, and the right threshold is context-dependent.
- Governance that anticipates autonomy. Regulatory frameworks built around devices and diagnostics do not map cleanly onto systems whose behavior changes with each deployment context.
- Shared definitions. Without agreed vocabulary for what counts as operational versus decisional trust, institutions cannot compare notes, and each deployment reinvents its own standards in isolation.
What Earned Autonomy Would Actually Look Like
The paper's implicit prescription is that autonomy should be granted incrementally and evidenced, not assumed from benchmark scores. That implies a fairly demanding set of institutional habits.
- Deployments that begin with narrow, reversible actions and expand scope only as evidence accumulates.
- Monitoring that tracks decision quality over time, not just uptime and latency.
- Clear accountability: a named human or committee responsible when an automated action causes harm.
- Documentation rigorous enough that an independent reviewer can reconstruct what happened and why.
None of this is glamorous. It is the connective tissue that turns a capable system into a trustworthy one, and it is precisely the layer that tends to be skipped when the underlying model is impressive enough to generate enthusiasm.
The Uncomfortable Conclusion
The most challenging implication of the Nature Medicine paper is that the missing links in agentic AI autonomy may not be technical at all. Better models will make operational trust easier to achieve — faster, more reliable, less brittle. But decisional trust depends on questions that no architecture solves: who is accountable, what counts as acceptable risk, how much oversight is enough, and who gets to decide.
Those are institutional and clinical questions, and healthcare has decades of hard-won experience answering them for human practitioners. The work ahead is to extend that experience to systems that act rather than merely answer. Until that happens, the honest description of most agentic AI in medicine is not autonomous. It is unsupervised, which is a different thing entirely.
This article is based on reporting by Nature Medicine. Read the original article.
Originally published on nature.com








