Skip to main navigation Skip to search Skip to main content

Towards reliable use of artificial intelligence to classify otitis media using otoscopic images: addressing bias and improving data quality

  • Yixi Xu
  • , Al Rahim Habib
  • , Graeme Crossland
  • , Hemi Patel
  • , Chris Perry
  • , Kris Bock
  • , Tony Lian
  • , William B. Weeks
  • , Rahul Dodhia
  • , Juan Lavista Ferres
  • , Narinder Pal Singh
  • Microsoft USA
  • The University of Sydney
  • Westmead Hospital
  • Children’s Health Queensland
  • Royal Darwin Hospital
  • University of Queensland
  • Microsoft Pty Limited

Research output: Contribution to journalArticlepeer-review

Abstract

Ear disease contributes significantly to global hearing loss, with recurrent otitis media being a primary preventable cause in children, impacting development. Artificial intelligence (AI) offers promise for early diagnosis via otoscopic image analysis, but dataset biases and inconsistencies limit model generalizability and reliability. This retrospective study systematically evaluated three public otoscopic image datasets (Chile; Ohio, USA; Türkiye) using quantitative and qualitative methods. Two counterfactual experiments were performed: (1) obscuring clinically relevant features to assess model reliance on non-clinical artifacts, and (2) evaluating the impact of hue, saturation, and value on diagnostic outcomes. Quantitative analysis revealed significant biases in the Chile and Ohio, USA datasets. Counterfactual Experiment I found high internal performance (AUC > 0.90) but poor external generalization, because of dataset-specific artifacts. The Türkiye dataset had fewer biases, with AUC decreasing from 0.86 to 0.65 as masking increased, suggesting higher reliance on clinically meaningful features. Counterfactual Experiment II identified common artifacts in the Chile and Ohio, USA datasets. A logistic regression model trained on clinically irrelevant features from the Chile dataset achieved high internal (AUC = 0.89) and external (Ohio, USA: AUC = 0.87) performance. Qualitative analysis identified redundancy in all the datasets and stylistic biases in the Ohio, USA dataset that correlated with clinical outcomes. In summary, dataset biases significantly compromise reliability and generalizability of AI-based otoscopic diagnostic models. Addressing these biases through standardized imaging protocols, diverse dataset inclusion, and improved labeling methods is crucial for developing robust AI solutions, improving high-quality healthcare access, and enhancing diagnostic accuracy.

Original languageEnglish
Article numbere0338867
Number of pages13
JournalPLoS One
Volume21
Issue number5
DOIs
Publication statusPublished - May 2026
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being
  2. SDG 10 - Reduced Inequalities
    SDG 10 Reduced Inequalities

Fingerprint

Dive into the research topics of 'Towards reliable use of artificial intelligence to classify otitis media using otoscopic images: addressing bias and improving data quality'. Together they form a unique fingerprint.

Cite this