Medical AI has produced a large body of published results showing performance comparable to or better than clinicians on specific tasks. Many of those results are real.
The deployment record is more mixed, and the reasons are consistent and knowable in advance. A workforce that understands them catches problems before a system reaches a patient.
The gap between the paper and the ward
A reported accuracy figure describes performance on a particular dataset, assembled in a particular way, evaluated against a particular reference standard.
Clinical performance is what happens when the system meets an unselected population, different equipment, different practice, and clinicians who respond to its output in ways the evaluation did not model.
The model was not wrong in the paper. It was answering a question about a dataset, and the hospital is asking a different question.
The demand pressure is real: the World Economic Forum's Future of Jobs Report 2025 found 92 percent of employers in medical and healthcare services say AI and big data skills are growing in importance for their workforce. The risk is that the capability being hired is model building rather than model evaluation.
What the direction covers
The scope: AI in radiology and pathology, clinical decision support, image analysis and validation.
The word validation is doing the most work in that list, and it is the part most often under-resourced.
Four capabilities.
Medical imaging and its physics. What an image actually is, how acquisition parameters affect it, and why images from different equipment differ in ways that are invisible to a human reader and highly visible to a model.
Model development for clinical data. Including the specific issues: label quality, class imbalance, data leakage in ways peculiar to clinical datasets.
Clinical validation. Study design for evaluating a clinical AI system, which is epidemiology rather than machine learning.
Deployment and monitoring. Integration into workflow, ongoing performance monitoring, and knowing what to do when performance drifts.
Distribution shift, in plain terms
The central technical problem, stated without jargon.
A model learns from the data it was trained on. That data came from particular hospitals, particular equipment, particular populations and particular clinical practices.
Deployed somewhere else, several things differ at once.
The population differs. Different age distribution, different disease prevalence, different comorbidity. Prevalence alone changes the predictive value of any output, which is the base rate point from the epidemiology direction and it applies with full force here.
The equipment differs. Different scanners, protocols and settings produce systematically different images. Models are known to pick up on equipment characteristics rather than pathology, and such a model can appear excellent in validation and fail entirely on a different machine.
Practice differs. Who gets tested, when, and how the reference standard was established all vary between institutions.
Time passes. Practice changes, equipment is replaced, populations shift. A model that was accurate at deployment can degrade without anyone changing anything, which is why monitoring is a permanent staffing requirement rather than a project task.
The workforce implication: local validation before deployment and continuous monitoring after it are both necessary, and both need people. Organisations that buy a system without budgeting for either have bought an unmonitored clinical intervention.
Where this sits in the domain
Medical AI, imaging and diagnostics is the eighth of ten directions in Astra Trainer's medicine and healthtech domain. It depends on public health and epidemiology for validation methodology, on health informatics for the data layer, on pathophysiology for clinical plausibility, and on medical devices for the regulatory framework, since clinical AI is frequently regulated as a device.
For partners staffing from the technical side, the AI, data and computing domain covers artificial intelligence and machine learning, data science and analytics, and cybersecurity. The combination that actually works is technical capability plus clinical validation capability, and organisations usually have the first. You can see the ten directions here.
Retrospective is not prospective
A distinction that decides whether evidence means anything, and it is routinely blurred in commercial claims.
Retrospective evaluation runs a model on historical data where outcomes are already known. It is cheap, fast and necessary, and it systematically overstates real-world performance. The dataset was curated, poor-quality cases were often excluded, and the model may exploit patterns that exist only because of how the data was assembled.
Prospective evaluation runs the model on patients as they present, with performance assessed against what actually happened afterward. Slower, more expensive, and far more informative.
Clinical impact evaluation asks the question that matters: does using this system change what happens to patients. A model can be accurate and change nothing, because clinicians already knew, or because the finding does not alter management, or because the workflow does not act on it.
Each step is a higher bar and each is considerably less common than the one before. A team that cannot articulate which of these an evidence claim rests on cannot evaluate a procurement decision.
What clearance does and does not mean
Regulatory clearance or approval means a device met the requirements of a defined pathway in a particular jurisdiction, on the evidence submitted, for a stated intended purpose.
It does not automatically mean the system improves patient outcomes, that it will perform as described in your population, or that it has been compared against current practice.
This is not a criticism of regulators. Regulatory pathways are designed to establish defined standards, and different pathways require different levels of evidence. It is a warning about how clearance is used in sales conversations, where it is frequently presented as though it settled the clinical question.
An organisation needs someone who can read what a clearance actually covers, including the stated intended purpose, and compare it against the use being proposed. That is a specific and scarce capability.
The human in the loop, which is not a solution by itself
The standard reassurance is that a clinician reviews the output, so errors are caught.
This is weaker than it sounds, for reasons that are well documented in human factors research.
Automation bias. People tend to defer to confident automated recommendations, including when they are wrong, and including when the person would have decided correctly unaided.
Deskilling. Capability that is not exercised degrades. A reviewer who has not independently assessed cases for a long time is a weaker check than they were.
Volume defeats review. Where a system's value is processing more than a human could, genuine independent review of every output is not possible by construction.
Responsibility becomes unclear. If the system is usually right and the clinician is accountable, the clinician carries risk for a system they cannot inspect. That is a governance problem that needs settling before deployment rather than after an incident.
Human oversight is necessary and it is not sufficient, and designing it to actually work requires human factors capability, which links to the usability point in the medical devices direction.
The roles, named
Clinical AI validation scientists. The scarcest and most important role here, and one that barely exists as a job title.
Medical imaging scientists and physicists. Understand acquisition and what makes images differ.
Clinical data scientists with health-specific methodological grounding.
Clinical safety officers for AI systems.
Regulatory specialists for software and AI devices.
Implementation and workflow specialists, because integration decides whether anything changes.
Monitoring and performance analysts, post-deployment, which is a permanent function.
Who can be trained into it
Radiographers and imaging technologists. Underused and well placed. They understand acquisition, artefacts and the practical reality of imaging, and adding validation methodology produces exactly the capability that is missing.
Medical physicists. Already hold quantitative and imaging expertise.
Data scientists in healthcare. Need epidemiological methodology and clinical grounding rather than more modelling technique.
Clinicians with analytical interest. Hold the plausibility judgement that catches problems no metric shows.
Clinical scientists in pathology and laboratory medicine, as digital pathology grows.
Quality and regulatory staff extending into software and AI devices.
Regulation and accountability. Clinical AI systems are frequently regulated as medical devices, with the obligations described in the medical devices direction, and in several jurisdictions their clinical deployment requires formal clinical safety assessment. Clinical accountability for decisions remains with registered professionals. Training builds validation and evaluation capability. It does not confer regulatory approval, clinical safety qualification, or authority to deploy a system, and no model output should be treated as clinical advice.
What to take from this
Reported accuracy and clinical performance are different quantities, and the gap is predictable rather than mysterious.
Distribution shift across population, equipment, practice and time is the central problem, which makes local validation and continuous monitoring permanent staffing requirements.
Retrospective, prospective and clinical impact evidence are three different bars, and most claims rest on the first.
Clearance means a defined pathway was satisfied for a stated intended purpose. It is not evidence of patient benefit and it is frequently presented as though it were.
And human oversight is necessary and insufficient, because automation bias and volume both undermine it by design.
Why do medical AI systems underperform after deployment?
Distribution shift. A model learns the population, equipment, practice patterns and reference standards it was trained on, and all four differ elsewhere. Prevalence alone changes the predictive value of any output.
Does regulatory clearance mean a system works?
It means a defined pathway was satisfied on the evidence submitted for a stated intended purpose. It is not equivalent to evidence of improved patient outcomes in your population.
Is human review enough of a safeguard?
No. Automation bias leads people to defer to confident systems, unexercised skills degrade, and where the system's value is volume, genuine independent review of every output is impossible by construction.
Who should be trained for this?
Radiographers and imaging technologists are underused and well placed, since they understand acquisition and artefacts. Data scientists in healthcare need epidemiological methodology rather than more modelling technique.
Where does this fit in the domain?
Eighth of ten directions in Astra Trainer's medicine and healthtech domain, depending on epidemiology for validation method and medical devices for the regulatory frame. You can see them here.
