Even the latest generation of reasoning-capable AI models reproduce racial and gender stereotypes when asked about medical scenarios, according to new research from Australian scientists. The findings, reported on August 7, 2026, are a sobering reminder that scaling up model intelligence does not automatically eliminate the systemic biases embedded in training data. For more breaking AI news on the intersection of artificial intelligence and healthcare equity, the study raises urgent questions about deploying these tools in clinical settings.

Testing the Next Generation of Reasoning Models

Researchers asked two next-generation reasoning large language models, OpenAI's o3-mini and DeepSeek-R1, about fictional patients to assess whether newer models have overcome the well-documented biases of their predecessors. Reasoning models, which spend additional computation "thinking" through problems before answering, are generally considered more capable than standard chatbots at complex tasks. The assumption has been that this added capability might also make them more reliable in sensitive domains like medicine.

The results suggest otherwise. The study found that both models reproduced racial and gender stereotypes when presented with clinical scenarios, indicating that the leap from standard to reasoning architectures has not resolved the fundamental problem of biased training data.

A Persistent Problem Across Generations

The findings are consistent with a growing body of evidence that medical bias in AI is stubbornly persistent. A systematic literature review published in Nature in April 2026 examined bias, representation, and clinical fidelity in AI-generated images used for medical education. That review found widespread underrepresentation of certain demographic groups and stereotyped portrayals when they did appear.

In May 2026, The Lancet published a scoping review of fairness metrics for AI-based clinical prediction models. That analysis found that while researchers have proposed numerous mathematical frameworks for measuring fairness, the metrics are applied inconsistently across studies, making it difficult to assess whether real progress is being made.

The Australian study adds a new dimension by focusing specifically on reasoning models, the category of AI that many believed would be more robust against bias because of their structured problem-solving approach.

Why Reasoning Does Not Eliminate Bias

The persistence of bias in reasoning models points to a structural issue in how these systems learn. Large language models are trained on vast corpora of text, much of which reflects the biases, stereotypes, and inequities present in the real world. Medical literature itself has historically overrepresented white male patients in clinical trials and underrepresented women and racial minorities.

Reasoning models improve performance on logic and multi-step problems by generating intermediate reasoning steps, but those reasoning steps are still produced by the same underlying model trained on the same data. If the model has learned associations between certain demographic groups and certain medical conditions or personality traits, it can reproduce those associations even while "reasoning" through a clinical scenario.

This means that the clinical fidelity of a model, its ability to produce accurate, equitable medical outputs, depends not just on reasoning architecture but on the representativeness and quality of the data it was trained on.

The Stakes for Healthcare AI

The findings arrive at a moment when AI is being rapidly embedded in healthcare infrastructure. Hospitals and health systems are deploying AI for tasks ranging from clinical documentation and medical imaging analysis to patient triage and treatment recommendation. The HCPLive network reported in August 2026 on concerns that AI could worsen healthcare disparities if biased models are deployed without adequate oversight.

If reasoning models still reproduce stereotypes about race and gender when asked about patients, the risk is that these tools could reinforce existing disparities in diagnosis, treatment recommendations, and patient communication. A model that systematically associates certain conditions with certain demographic groups, or that treats patients differently based on race or gender, could cause real harm in clinical settings.

The Eurasia Review, reporting on the study, noted that the persistence of bias across model generations suggests that the AI industry's approach to addressing medical bias may need to shift from model architecture improvements to data governance and representational equity.

What Needs to Change

The research points to several necessary changes. First, training data for medical AI must be audited for representational equity, ensuring that women, racial minorities, and other underrepresented groups are adequately reflected. Second, fairness testing should be a mandatory step before deploying AI in clinical settings, not an afterthought. Third, the medical community needs standardized benchmarks for evaluating whether models treat patients equitably across demographic groups.

The Lancet scoping review's finding that fairness metrics are applied inconsistently highlights the need for standardized evaluation frameworks. Without agreement on how to measure bias, it is impossible to hold AI developers accountable or to track whether progress is being made.

A Call for Caution

The Australian study serves as a warning against assuming that more capable models are inherently safer. As reasoning models become the default for high-stakes applications in medicine, law, and finance, the medical bias findings suggest that raw capability and equity are separate problems that require separate solutions.

For healthcare organizations considering AI deployment, the message is clear: even the most advanced reasoning models available today can reproduce harmful stereotypes, and that risk must be actively managed through testing, oversight, and a commitment to equitable data.

Stay Ahead of AI

As AI becomes embedded in healthcare, understanding its limitations is just as important as celebrating its capabilities. To keep up with the latest AI developments in medical research, model safety, and healthcare technology, follow our ongoing AI industry coverage.

Read more AI news →