Study Reveals Persistent Race and Gender Bias in Medical AI Models
Posted on 17 Aug 2026
Hospitals are turning to artificial intelligence (AI) to draft clinical materials and support decision-making, yet algorithmic bias can entrench inequities in care. When generated content reflects racial or gender stereotypes, it may distort diagnostic reasoning and undermine patient trust. Persistent bias also threatens quality initiatives that rely on accurate depictions of patient populations. A new study shows that next-generation reasoning AI still reproduces these stereotypes in common medical scenarios.
Researchers at Flinders University evaluated two reasoning large language models (LLMs), o3-mini and DeepSeek-R1, to determine whether stronger reasoning capabilities reduce representational bias in generated clinical content. Each model was prompted to create clinical vignettes for common conditions, and researchers compared the resulting race and gender distributions with known disease prevalence patterns. The study builds on earlier findings of demographic bias in GPT-4.
Across 36,000 unique vignettes, both models frequently misrepresented race and gender, echoing the earlier GPT-4 results. In the previous analysis, significant misrepresentation was found for 67% of conditions for both race and gender. The new study assessed whether newer reasoning models corrected these distortions, but the findings indicate that substantial bias remains.
For o3-mini, significant misrepresentation occurred in 78% of conditions for race and 56% for gender. DeepSeek-R1 showed misrepresentation in 89% of conditions for race and 67% for gender. Like GPT-4, both models overrepresented Black populations in conditions stereotypically associated with those groups, including sarcoidosis, systemic lupus erythematosus, pre-eclampsia, and essential hypertension. Median misrepresentation reached 44% for o3-mini and 31% for DeepSeek-R1, compared with 15% for GPT-4. The findings suggest that advances in reasoning capability have not translated into more representative clinical vignette generation.
Qualitative review of DeepSeek‑R1’s reasoning traces showed the model invoking disease‑demographic associations to select patient characteristics without citing quantitative epidemiology. The study also observed exaggeration of the majority gender, consistent with previously reported gender stereotyping in clinical text generation. The authors note that advances in capability do not guarantee improvements in fairness or representation. The research is published in the Journal of Medical Internet Research.
Related Links
Flinders University