Astra Trainer
Future Industries

The Biostatistics Gap Behind Most Health Data Mistakes

Aleksandr Mikhailov
Founder, Astra Trainer
Updated
8 min read

Epidemiology sounds like a specialised public health subject. It is better understood as the accumulated knowledge of how population health data fools people, written down by a field that has been fooled repeatedly and learned from it.

That makes it directly relevant to anyone working with health data, which is now a much larger group than the people who call themselves epidemiologists.

The discipline that exists to stop you fooling yourself

Health data has properties that break ordinary analytical intuition.

People are not randomly assigned to exposures. Sick people seek care, so appearing in a dataset is not independent of being unwell. Measurement happens because someone decided to measure, and that decision carries information. Outcomes take years. And the population that generated the data is not the population the conclusion will be applied to.

The methods exist because the obvious analysis is wrong often enough to have caused real harm, repeatedly, over a century.

The World Economic Forum's Future of Jobs Report 2025 found that 92 percent of employers in medical and healthcare services say AI and big data skills are growing in importance. That demand is being met largely by people trained on data that behaves nothing like health data.

What the direction covers

The scope: how disease spreads, prevention, population health, biostatistics, health policy and inequality.

Four capabilities.

Study design. Cohort, case-control, cross-sectional, randomised. Which question each design can answer and which it cannot. Most analytical disputes resolve at this level rather than at the statistics.

Measures of association and impact. Relative and absolute risk, odds ratios, numbers needed to treat. Understanding why a large relative change can be a trivial absolute one, which is the most commonly exploited gap in health reporting.

Bias and confounding. The systematic ways a study can be wrong, and what can be done about each.

Population health and prevention. Screening, vaccination, surveillance, and the reasoning behind population-level intervention.

Four errors that keep being made

Confounding read as causation. Two things move together because a third drives both. In health data the third thing is very often age, deprivation or baseline severity, and all three are strongly associated with almost every outcome.

Selection effects in routine data. Electronic records contain people who sought care and were tested. A model trained on that data learns about the tested population, not the general one. This is not fixable by adding data, because more data from the same source has the same selection built in.

Base rate neglect in test performance. The most consequential error in this list. A test with excellent sensitivity and specificity produces mostly false positives when the condition is rare, because the number of people without the condition is so much larger. Anyone evaluating a diagnostic, a screening program or a predictive alert needs this, and it is routinely absent from product claims.

Competing risks and immortal time. Subtler problems around how time is handled, both of which produce impressive-looking effects from nothing. They appear frequently in analyses of routine data by people who have not encountered them.

Where this sits in the domain

Public health and epidemiology is the fourth of ten directions in Astra Trainer's medicine and healthtech domain, and it is the methodological backbone of the ones above it: clinical research and clinical trials, medical AI and imaging, and precision medicine all depend on it.

For partners whose people are arriving from data science rather than from health, it is usually scoped alongside the AI, data and computing domain, where data science and analytics covers the statistical machinery. The pairing matters because the machinery without the epidemiology produces confident errors. You can see the ten directions here.

Health inequality as technical content

The site's own description of this direction includes inequality, and it belongs there as method rather than as commentary.

Three reasons it is technical.

It is a confounder almost everywhere. Socioeconomic position is associated with exposure, with access to care, with how completely data is recorded and with outcome. An analysis that omits it will attribute its effects to whatever else is in the model.

It shapes who is in the data. Groups with poorer access to care are underrepresented in clinical datasets. Models trained on those datasets perform worse for exactly the groups already receiving less, and the failure is invisible in aggregate performance metrics.

Aggregate performance hides subgroup failure. A model can perform well overall and badly for a subgroup, and a headline metric will not show it. Stratified evaluation is a methodological requirement rather than an ethical extra.

So a workforce program that teaches this properly produces people who evaluate systems in a way that catches problems before deployment. One that treats it as a values statement produces people who agree it matters and do not know how to check.

The roles, named

Epidemiologists and public health analysts in agencies and health systems.

Biostatisticians, in trials, research and health systems.

Health data scientists working with population and routine data.

Surveillance and outbreak response staff.

Health economists and outcomes researchers, where the same methods support reimbursement and value arguments.

Regulatory and HTA analysts assessing evidence submitted by manufacturers.

Clinical trial statisticians and methodologists, linking to the clinical research direction.

Policy analysts in government and insurers, who make decisions that rest on these methods whether or not they understand them.

Who can be trained into it

Data scientists and statisticians moving into health. The largest group and the one with the biggest gap relative to consequence. They have the technical machinery and lack the domain-specific failure modes, which is precisely what this direction is.

Clinical staff moving into analytical roles. Hold the clinical understanding and need the methods. The complementary conversion, and generally the more reliable one, because clinical intuition catches implausible results.

Policy and government analysts. Frequently interpreting epidemiological evidence with no formal training in it.

Health service managers working with performance and outcome data that carries all of the above problems.

Journalists and communications staff in health organisations, where the relative-versus-absolute risk distinction alone changes the quality of output.

Insurance and payer analysts working with claims data.

What this is and is not. This direction builds methodological capability for analysing population health data. It is not clinical training and does not support decisions about individual patients, and the distinction between population-level and individual-level inference is itself one of the things the discipline exists to enforce. Public health practice in official roles also carries its own professional registration and qualification requirements in many jurisdictions.

What to take from this

Epidemiology is the recorded history of how health data misleads, which makes it useful to far more people than public health.

Confounding, selection effects, base rates and time handling cause more wrong conclusions than any modelling choice, and none of them is fixed by more data.

Base rate neglect is the most consequential single error, because it makes a good test look useless or useful depending on a fact about the population that product claims routinely omit.

Health inequality is a confounder, a data-availability problem and a subgroup-performance problem, which makes it method rather than commentary.

And data scientists entering health are the group whose training gap is largest relative to what their work decides.

Frequently asked questions
Why do data scientists need epidemiology?

Because health data breaks ordinary analytical intuition. People are not randomly assigned to exposures, appearing in a dataset depends on having sought care, and measurement itself carries information. The methods exist because the obvious analysis is frequently wrong.

What is the most consequential error?

Base rate neglect. A test with excellent sensitivity and specificity produces mostly false positives when the condition is rare, and product claims routinely omit the prevalence that determines this.

Is health inequality a technical topic?

Yes. It confounds almost every analysis, it determines who appears in clinical datasets, and it produces subgroup failures that aggregate performance metrics hide. Stratified evaluation is a methodological requirement.

Who converts into these roles best?

Clinical staff moving into analysis, because clinical intuition catches implausible results, and data scientists moving into health, who have the machinery and need the domain failure modes.

Where does this fit in the domain?

Fourth of ten directions in Astra Trainer's medicine and healthtech domain, and the methodological backbone under clinical research, medical AI and precision medicine. You can see them here.

The methods exist because the obvious answer is wrong
Ten directions across medicine and healthtech, including public health and epidemiology alongside clinical research, medical AI and precision medicine. Scoped with your own specialists, in five-minute lessons.
Written by Aleksandr Mikhailov
Founder, Astra Trainer · Published · Updated
Continue reading