Blog

Medical researchers reviewing diagnostic imaging and clinical evidence together

Artificial Intelligence in Medicine: From Prediction to Clinical Evidence

AI Science & Applications

Artificial Intelligence in Medicine: From Prediction to Clinical Evidence

Artificial intelligence is becoming part of medicine, but its scientific value depends on a question more demanding than whether a model can produce an impressive answer: does it improve care safely for the population in which it is used?

AI systems can analyse medical images, estimate clinical risks, assist with documentation, identify patterns in patient records and help researchers search large bodies of literature. These are different tasks and require different forms of evidence. A model that classifies an image is not evaluated in the same way as a conversational system that generates a clinical summary.

From retrospective accuracy to clinical utility

Many medical AI studies begin with retrospective data: historical cases are divided into training and test sets, and model predictions are compared with known outcomes. This can establish technical performance, but not necessarily clinical usefulness.

A test set may not represent another hospital, scanner, population or period. Disease prevalence influences predictive value, and changes in clinical practice can create distribution shift. Strong evaluation therefore includes external validation, prospective testing and analysis of performance across relevant patient groups.

Decision support is not decision replacement

AI can detect statistical patterns, but clinical decisions combine evidence with examination, patient preferences, uncertainty and professional responsibility. A model may support attention by identifying cases that deserve review, yet its output can also create automation bias if users assume the recommendation is more reliable than it is.

Human oversight must be meaningful. A clinician needs enough information, time and authority to challenge the system. Simply placing a person at the end of an automated process does not guarantee safe supervision.

Bias can enter at multiple stages

Medical data reflect how care was delivered, which patients had access and how outcomes were recorded. Labels may contain diagnostic uncertainty. Underrepresented populations may receive less reliable predictions. Bias can also emerge after deployment when the tool changes which patients are tested or treated.

Evaluation should therefore examine data provenance, subgroup performance, missingness and the consequences of false positives and false negatives. A single average accuracy score can hide clinically important differences.

Governance is part of scientific quality

The World Health Organization identifies autonomy, safety, transparency, accountability, inclusiveness and sustainability as core principles for AI in health. It also warns that plausible language from large language models can contain serious errors and should not be adopted without rigorous evidence and oversight.

Medical AI is most credible when its intended use is narrow enough to evaluate, its limitations are documented and its performance is monitored in practice. The future is unlikely to be medicine without clinicians. It is more plausibly medicine in which carefully validated computational tools help clinicians see patterns, manage information and allocate attention—while responsibility remains human.

What a convincing clinical study should measure

Sensitivity and specificity describe only part of a system’s behaviour. Calibration matters too: outcomes should occur at approximately the rate the model predicts. A clinically useful evaluation also records downstream tests, treatment delays, alert burden, workflow failures and the different consequences of false positives and false negatives.

The strongest evidence compares an AI-supported pathway with credible normal care in the intended population. External and prospective validation are more informative than another retrospective benchmark. Hospitals replace scanners, laboratories alter assays and populations change, so deployed performance must be monitored for drift. Model updates require version control and, when behaviour changes materially, renewed validation.

A practical scientific checklist

  • Define the population, task, setting and user before choosing a model.
  • Compare against a credible clinical baseline, not only another algorithm.
  • Validate externally and prospectively across relevant patient groups.
  • Measure calibration, clinical utility, workflow effects and possible harms.
  • Document privacy controls, accountability and conditions for withdrawal.

Health data are unusually sensitive. Data minimisation, secure access and a clearly stated purpose should be designed into the system. Explanations must suit their audience, but explanation alone does not determine who is accountable for procurement, validation, use and incident response.

Further reading: WHO guidance on ethics and governance of AI for health.

Leave your thought here

Your email address will not be published. Required fields are marked *