Tuesday, July 28, 2026
No menu items!
HomeNatureWhen physicians and AI work together, who is accountable? How to lay...

When physicians and AI work together, who is accountable? How to lay out medical liability

Artificial intelligence is fast becoming embedded in hospitals and healthcare clinics. Yet, with the benefits of AI tools come fresh patient-safety risks — and questions about who is responsible when things go wrong.

Conventionally, responsibilities in health care are clearly delineated. Clinicians must provide a set standard of treatment. Institutions must organize safe treatment processes. Manufacturers must deliver non-defective medical devices. Regulators and bodies involved with licensing and credentials can intervene if any of these parties fall short of expected standards. And courts can assess which party is liable if patients are harmed.

An incoming wave of medical AI tools is set to blur these lines.

Current AI tools are being used mainly as assistants, backed up by human checks, much as with other medical devices. For instance, when AI models trigger an alert that a patient is at risk of sepsis, a clinician must review and confirm the physiological rationale before any treatment.

But in the next few years, AI-driven clinical tools are expected to advance from synthesizing data to acting on it, with less and less human input. They will make diagnoses, devise treatment plans and make patient-management decisions that clinicians play little or no part in. Whereas the outputs of many existing tools are understandable — sepsis models, for instance, work by combining defined physiological variables according to a transparent formula1 — more-advanced tools can involve black-box deep-learning processes. Clinicians will no longer be able to fully follow the tools’ reasoning, even for algorithms that provide some explanations for their judgements2.

Such tools land between accountability rules. Their black-box reasoning makes it hard to determine when they are defective. And when decisions are made jointly by humans and an opaque AI model, it becomes difficult to assess responsibility when people are harmed while receiving health care. Existing medical-liability frameworks do not address this crossover point3.

This legal uncertainty means that people who are harmed could fall into liability gaps in which no one has clearly broken a rule3. Hospitals are likely to avoid AI for potentially useful functions because of concerns about legal risks — concerns that are repeatedly cited by physicians as a barrier to safe AI adoption46 and that are seen by some as grounds not to use black-box tools already available. And AI vendors might shirk their obligation to monitor safety once a tool reaches the market, confident that it will be hard to attribute blame when things go wrong.

Clear liability frameworks are needed.

Here we outline how to achieve that, by defining seven levels of AI capabilities on the basis of three factors: autonomy, automation and operational scope. With aircraft and self-driving cars, regulators already use graded levels to specify exactly what tasks systems must perform, how independently, and when humans are expected to take over79. A similar set of levels for medical AI systems would help regulators, policymakers, governments, courts and professional health-care bodies to plan for the incoming tools.

Three properties, seven levels

Three properties dictate how AI is used in the clinic.

First, autonomy: how independently the system reasons, without the need for human input. Existing sepsis risk models are low-autonomy AI tools1. Deep-learning systems, by contrast, might arrive at diagnoses from patterns across thousands of variables, making inferences that physicians cannot retrace10.

Second, automation: what tasks a medical AI system can execute on its own. Such a system might range from a low-automation robot that a surgeon controls, to a system that sends chemotherapy prescriptions directly to a pharmacy.

Third, operational scope: the boundaries in which the system is allowed to run. A retinal-imaging system deployed by a specialist ophthalmology service to double-check clinical decisions is less risky than the same tool used in a primary-care clinic to make an autonomous decision about whether to refer a patient to a specialist.

A woman undergoes an AI-assisted breast cancer scan while a medical professional monitors the test on a computer.

A scanner being trialled in Krakow, Poland, uses AI to help to detect breast cancer. Credit: Klaudia Radecka/NurPhoto via Getty

The following seven levels could form a basis for defining risks around medical AI tools. A given tool might sit at different levels depending on where and how it is used, and by whom.

Level 0: Informational. An AI system with no autonomy, variable automation and narrow operational scope. Such systems store, move and display data according to rules set by humans. Examples include algorithms for managing electronic health records or streaming data from wearable devices that track vital signs. Algorithmic behaviour is transparent and any output is independently verifiable.

This is well-trodden ground. Regulators can assess whether such tools are reliable and offer sufficient cybersecurity and data protection. Courts can apply conventional standards to assess liability.

Level 1: Assistive. AI systems with low autonomy, low-to-moderate automation and narrow operational scope. Such systems can perform a single, clearly defined task, overseen by a clinician — such as flagging suspicious rhythms on a heart monitor11 or highlighting possible lung nodules on a scan12. The tools’ algorithmic reasoning might be opaque, but the output can be verified through clinical review of patient data, or by ‘explainability techniques’ that check correlations between the model’s input and output data.

These tools are regulated as diagnostic aids, similar to existing computer-aided detection systems for mammography and colonoscopy. Hospitals are held responsible for ensuring that staff understand their limitations. Users are expected to accurately judge whether alerts are meaningful.

Level 2: Decision support. AI tools with low autonomy, low-to-moderate automation and moderate operational scope. These systems combine multiple streams of data — perhaps including symptoms, previous medical history, imaging data and current guidelines. Existing tools at this level can produce a diagnosis and treatment plan, draft detailed discharge summaries and generate risk scores for triage decisions. As with level 1 tools, they are designed to be advisory — their reasoning might be opaque but is verifiable to some extent, and clinicians are expected to make the final decision.

Regulators should ask for robust evidence that these systems improve care. Hospitals must train clinicians to challenge outputs instead of accepting them by default, and courts must decide whether a clinician’s reliance on a particular recommendation was reasonable in the circumstances. There’s a risk that clinicians under time pressure will over-rely on advisory tools. There will be no easy way to determine how much a clinician was swayed by a machine’s recommendation.

Level 3: Supervised automation. AI systems with moderate levels of autonomy and automation, and moderate operational scope. The system, rather than the clinician, is the decision maker. These tools are currently rare (see ‘Classifying current medical AI’). In a tightly defined domain, they can dynamically respond to treatment needs — for example, devices for insulin delivery automatically adjust doses according to continuous glucose readings, using embedded algorithms and predefined limits13. The clinician selects which patients are placed on the system and intervenes only when something seems wrong.

Classifying Current Medical AI. A stacked horizontal bar chart classifies medical artificial intelligence (AI) device use in the five countries with the most device approvals. Levels range from Level 0 (assistive use of AI), to Level 3 (supervised automation).

Source: Analysis by K. Lam et al.

Regulators should focus on the quality of instructions for using and maintaining these systems and on information about the risks inherent in providing care. Health-care organizations should furnish guidance for informed-consent protocols that respond to the new types of risk inherent in AI-driven treatment, and providers should discuss these risks with patients.

With these systems, physicians can be criticised for enrolling an unsuitable patient, failing to respond to device alarms or not understanding system limitations. Organizations might be judged on how they selected, validated and monitored the technology. Manufacturers might be liable if flawed designs or updates contribute to harm. But conventional liability assessments can be unfit for purpose, if physicians and organizations acted reasonably, and harm is caused by an algorithmically determined decision that is not obviously wrong.

Level 4: AI-triggered human oversight. AI tools with moderate-to-high levels of autonomy and automation, and broad operational scope. These systems manage whole episodes of therapy in areas such as intensive care. They continuously monitor patients and adjust treatments — for instance balancing ventilator and sedation settings against several physiological targets at once. They trigger back-up protocols such as summoning a rapid-response team when something seems wrong. Clinicians are a safety net and act only when the system sounds the alarm. Level 4 deployments are yet to break through in medicine, although the technical components needed are emerging.

For these systems, regulators should require strong evidence that including a human in the loop would reduce safety. Given that clinicians would have no realistic opportunity to review the actions of these systems, courts might struggle to determine whether an adverse outcome reflects negligence, a hidden defect or simply the inherent risks of medicine.

RELATED ARTICLES

Most Popular

Recent Comments