Artificial intelligence is moving into hospitals and clinics at a pace that legal and regulatory frameworks have not matched. The tools arriving now do more than display data or flag anomalies for a physician to review. The next generation is expected to make diagnoses, construct treatment plans, and manage patients with little or no human involvement. That shift creates a problem that medicine has not faced before: when an AI system causes harm, the existing rules for assigning responsibility may not apply to anyone.
The Accountability Gap That Advanced AI Opens
Healthcare has traditionally operated with clear lines of responsibility. Clinicians are held to a standard of care. Hospitals must organize safe processes. Device manufacturers must deliver products that work as intended. Regulators and licensing bodies can intervene when any of these parties fall short. Courts can determine liability when patients are harmed.
AI tools complicate every one of these lines simultaneously.
Current systems still operate largely as assistants. A sepsis risk model, for example, combines defined physiological variables through a transparent formula and produces an alert that a clinician must review before any treatment follows. The reasoning is visible. The human remains in the loop. Liability, if something goes wrong, can still be traced through familiar channels.
The tools being developed now work differently. Advanced deep-learning systems can arrive at a diagnosis by detecting patterns across thousands of variables, through processes that even the physicians using them cannot retrace. These are what researchers call “black-box” systems: the output may be correct, but the path to it is opaque. When a decision made jointly by a clinician and an opaque AI model leads to patient harm, existing medical liability frameworks do not address who bears responsibility. Researchers writing on this topic have identified the result as a liability gap: a situation in which someone is harmed, but no party has clearly broken an identifiable rule.
The consequences of that gap are already visible in behavior. Physicians cite legal uncertainty as a repeated barrier to adopting AI tools, even ones that could benefit patients. Hospitals may avoid deploying AI in functions where it could genuinely help, precisely because the legal risk is undefined. And vendors of AI systems have less incentive to monitor safety after a product reaches the market if they can reasonably expect that blame will be difficult to assign.
A Seven-Level Framework for Grading AI Capability
A group of researchers, including Kyle Lam from Imperial College London, Mindy Nunez Duffourc from Maastricht University, Jiankai Sun from Stanford University, Eric Topol from the Scripps Research Translational Institute, and Jianing Qiu from Mohamed bin Zayed University of Artificial Intelligence, has proposed a structured approach to this problem. Their framework defines seven levels of medical AI capability based on three properties: autonomy, automation, and operational scope.
Autonomy refers to how independently a system reasons without human input. Automation describes what tasks the system can execute on its own. Operational scope defines the boundaries within which the system is permitted to function.
The same tool can sit at different levels depending on where and how it is deployed. A retinal imaging system used by a specialist ophthalmology service to double-check a clinical decision carries different risk than the identical tool used in a primary-care clinic to autonomously decide whether to refer a patient to a specialist. The technology is the same. The accountability context is not.
The framework draws on precedents from other high-stakes domains. Aviation and autonomous vehicle regulation already use graded levels to specify what tasks a system must perform, how independently, and at what point a human must take over. The researchers argue that a comparable structure for medical AI would give regulators, courts, policymakers, and professional healthcare bodies a shared vocabulary for the tools that are coming.
At the lower end of the scale, Level 0 covers informational systems with no autonomy, such as algorithms that manage electronic health records or stream data from wearable devices. Their behavior is transparent and their outputs are independently verifiable. Conventional liability standards apply without difficulty. Level 1 covers assistive tools that perform a single defined task under clinician oversight, such as flagging suspicious cardiac rhythms or highlighting potential lung nodules on a scan. Level 2 covers decision-support systems that synthesize multiple data streams, including symptoms, medical history, imaging, and clinical guidelines, to produce diagnoses, treatment plans, or triage risk scores. In both cases, the clinician retains final authority.
Why This Framework Matters Beyond Medicine
The deeper issue here is not specific to healthcare. It is about what happens to accountability structures when a consequential decision is made by a system whose reasoning cannot be fully audited by the humans responsible for the outcome.
Medicine makes this problem unusually visible because the stakes are immediate and the existing liability frameworks are well-developed. But the same structural tension appears wherever autonomous AI systems are embedded in professional judgment: in legal analysis, in financial advice, in infrastructure management. The question of who is responsible when an opaque system contributes to harm is one that every sector deploying advanced AI will eventually have to answer.
What the proposed framework offers is a way to make that question tractable. By grading AI systems according to what they can do, how independently they do it, and in what context, it becomes possible to assign regulatory expectations and liability standards before harm occurs rather than after. That is the logic behind aviation’s approach to automation levels, and it is the logic the researchers are applying to medicine.
The absence of such frameworks does not make AI safer. It makes the consequences of failure harder to address, and it gives cautious institutions a reason to avoid tools that could help patients.
In Short
Medical AI is advancing from passive assistant to active decision-maker, and the legal frameworks designed to protect patients have not kept pace. When an opaque AI system contributes to a harmful outcome, existing liability rules may leave no party clearly responsible. A proposed seven-level framework, grounded in autonomy, automation, and operational scope, offers a way to match regulatory expectations to actual AI capabilities. The goal is not to slow adoption but to make accountability legible enough that adoption can proceed safely.
Based on reporting from Nature: Machine Learning.