Do new AI models reproduce gender and racial stereotypes in medicine?
Relying on AI in health care carries a risk that existing gender and racial stereotypes will be reflected in the medical content it generates. This is according to Flinders University researchers who, to determine whether new AI models reproduce gender and racial stereotypes in medicine, asked o3-mini and DeepSeek-R1 — two next-generation reasoning large language models (LLMs) — to describe fictional patients with common medical conditions.
The researchers found that these models frequently reproduced gender and racial stereotypes, indicating, the researchers say, that advancements in AI reasoning do not inherently improve representational fairness. “Large language models have the potential to transform health care but risk exacerbating health disparities if they perpetuate biases,” said Joshua Docking, lead researcher from Flinders University’s College of Medicine and Public Health.
As noted by the Flinders researchers, previous research has demonstrated potential gender and racial biases in clinical vignettes generated by GPT-4; this has included over-representation of Black patients in stereotypical medical conditions. However, next-generation reasoning LLMs have emerged since then, offering improved reasoning capability and demonstrating superior benchmark performance, Flinders said.
“Whether these advances reduce representational bias in health care remains unknown, so this study evaluated whether reasoning LLMs exhibit racial and gender biases in generated clinical content,” Docking said.
Research lead author Professor Michael Sorich, Flinders University’s Professor in Clinical Pharmacology, added: “Our results show comparable or higher rates for o3-mini (78% race, 56% gender) and DeepSeek-R1 (89% race, 67% gender), indicating no improvement in representation with the newer reasoning models.”
The models generated 36,000 unique clinical vignettes for this research, and o3-mini and DeepSeek-R1 reasoning was found to frequently misrepresent the distribution of gender and race in medical conditions, mirroring issues previously observed in GPT-4, which met the threshold for significant misrepresentation in 67% for gender and 67% of conditions for race.
Like GPT-4, both o3-mini and DeepSeek-R1 were found to over-represent Black populations in stereotypically associated conditions such as sarcoidosis, systemic lupus erythematosus, pre-eclampsia and essential hypertension, with even higher median misrepresentation of 44% and 31%, respectively, compared to 15% in the earlier-generation GPT-4 software.
“This persistent pattern may reflect underlying bias, though the new models may also default to generating prototypical cases rather than representative samples due to patterns in their training data,” Sorich said. This was supported by qualitative analysis of DeepSeek-R1’s reasoning traces, which revealed that the model explicitly invoked disease-demographic associations when selecting patient demographics, without referencing quantitative epidemiological data.
“Consistently over-representing certain demographic groups, particularly for conditions that in practice affect diverse populations, risks reinforcing narrowed demographic assumptions in clinical contexts where understanding disease prevalence across populations is an important component of diagnostic reasoning,” Sorich added.
Previous findings that LLM outputs can skew towards gender stereotypes in health care align with consistent exaggeration of the majority gender, the researchers said. “Despite having enhanced reasoning capabilities, the clinical outputs of o3-mini and DeepSeek-R1 still exhibit racial and gender disease stereotyping in common medical conditions,” Sorich said.
Advancements in LLM capabilities do not guarantee parallel improvements across all dimensions, including fairness and representation in health care, the researchers said. “Awareness of these demographic defaults is essential for the safe integration of LLMs into clinical workflows, and continuous monitoring of potential biases should accompany their adoption,” Sorich added.
The study was published in the Journal of Medical Internet Research (doi: 10.2196/82256).
How agentic orchestration can help solve health care's workforce challenge
As healthcare organisations look to address workforce shortages and improve access to care, they...
Social media behind almost half of racism and discrimination complaints, Ahpra says
Practitioner conduct on social media accounted for 44% of notifications, particularly alleged...
New ACMA SMS rules now in force — could patients be missing vital updates?
With the SMS Sender ID Register having commenced this month, the GM of a multichannel messaging...
