Most people learn CPR on a manikin. That manikin, in most cases, has a flat chest. A 2024 study examining 20 CPR training models sold worldwide found that three quarters were described as male or had no sex specified. Only one offered a breast overlay. This is not a minor detail about training equipment. It is a visible symptom of a much deeper structural problem, one that now threatens to be encoded permanently into healthcare artificial intelligence.
The Gap That Starts Before the Algorithm
The consequences of underrepresentation in medical training are measurable. A US study of 19,331 out-of-hospital cardiac arrests found that 39% of women who collapsed in public received bystander CPR, compared with 45% of men. That six-percentage-point gap may partly reflect hesitation rooted in unfamiliarity: people trained on male-shaped manikins may feel less confident performing CPR on a female body. The timing matters enormously. A separate US study found that people who received bystander CPR four to five minutes after a witnessed cardiac arrest had 27% lower odds of surviving to hospital discharge than those who received it within one minute. In England, fewer than one in 12 patients for whom ambulance services attempted resuscitation survived 30 days, according to figures reported in 2024.
The manikin problem is fixable with better equipment. The deeper problem is what happens when the same historical bias enters the data that trains AI systems.
How Biased Records Become Biased Algorithms
Healthcare AI learns by finding patterns in medical records. The problem is that medical records do not just reflect biology. They reflect clinical decisions, and those decisions have historically been shaped by assumptions about sex and gender.
An experimental study of medical students and residents illustrates the mechanism clearly. When coronary heart disease symptoms were presented alongside psychological stress, women received fewer coronary heart disease diagnoses and fewer cardiology referrals than men. Their symptoms were more likely to be interpreted as psychogenic. An AI system trained on records shaped by those decisions would not know that a bias was present. It would simply learn the pattern and reproduce it.
This is not a hypothetical risk. Reviews of AI in medicine have flagged that unbalanced datasets produce uneven performance. Assessing that unevenness is itself difficult: a 2024 review of 692 AI-enabled medical devices approved by the US Food and Drug Administration found that demographic information and details of performance studies were often missing from public documents.
The data problem runs deep. Females have been consistently underrepresented in medical research for decades. Following the thalidomide tragedy of the late 1950s and early 1960s, the US Food and Drug Administration in 1977 recommended excluding women who could become pregnant from early drug trials. That policy was not reversed until 1993. Male animals were also favoured in laboratory research, and many studies failed to report the sex of cells used in experiments.
The effects on treatment are documented. Women clear the sleeping drug zolpidem more slowly than men, and in 2013 regulators recommended halving the starting dose for women. A 2020 analysis found that women experienced adverse drug reactions nearly twice as often as men, with differences in how drugs move through the body explaining many of them.
Old Assumptions, New Technology
Representation in medical research has improved. The 2016 Sex and Gender Equity in Research guidelines encourage researchers to consider and report sex and gender. A 2026 review of 574 papers found that 61% included both sexes. Yet only 44% of those papers analysed results by sex. Inclusion without analysis leaves the underlying knowledge gap intact.
The distinction between biological sex and gender adds another layer of complexity. Sex-related biology can influence how drugs are metabolised. Gendered experiences can affect access to healthcare and how symptoms are interpreted. Many health databases collapse both into a single binary category, limiting what researchers can actually conclude from the data.
Symptom presentation illustrates why precision matters. An analysis of more than one million people with acute coronary syndromes found that 74% of women and 79% of men experienced chest pain, but women had higher odds of reporting additional symptoms including neck or jaw pain, fatigue, and shortness of breath. An Australian study of more than 202,000 people admitted with stroke found that among those under 70 arriving by ambulance, women were less likely than men to be assessed as having a stroke by paramedics and less likely to be managed under the pre-hospital stroke protocol.
Here is what most coverage of healthcare AI misses: the problem is not that AI systems are poorly designed. The problem is that they are trained on records that were produced under conditions of systematic underrepresentation. A technically sound algorithm applied to historically skewed data will produce historically skewed outputs. The bias does not disappear because the technology is new.
A study of 6.9 million people in Denmark found that women were older at their first hospital diagnosis for most conditions examined. That pattern could reflect later disease onset, different healthcare-seeking behaviour, or slower diagnosis. Disentangling those causes requires careful research design. AI systems trained on the raw pattern, without that context, cannot make the distinction.
In Short
Decades of medical research that treated the male body as the default have produced datasets that AI systems now learn from. When those systems identify patterns in health data, they may be learning not just about disease but about how patients were historically treated differently based on sex and gender. Improving representation in future research is necessary but not sufficient. The more urgent task is understanding which patterns in existing data reflect biology and which reflect old assumptions, before those assumptions are embedded in the diagnostic and treatment tools of the next generation.
Based on reporting from The Conversation AI.