Mitigating Bias in Radiology Machine Learning: 1. Data Handling

Pouria Rouzrokh; Bardia Khosravi; Shahriar Faghani; Mana Moassefi; Diana V Vera Garcia; Yashbir Singh; Kuan Zhang; Gian Marco Conte; Bradley J Erickson

doi:10.1148/ryai.210290

Mitigating Bias in Radiology Machine Learning: 1. Data Handling

Radiol Artif Intell. 2022 Aug 24;4(5):e210290. doi: 10.1148/ryai.210290. eCollection 2022 Sep.

Authors

Pouria Rouzrokh¹, Bardia Khosravi¹, Shahriar Faghani¹, Mana Moassefi¹, Diana V Vera Garcia¹, Yashbir Singh¹, Kuan Zhang¹, Gian Marco Conte¹, Bradley J Erickson¹

Affiliation

¹ Radiology Informatics Laboratory, Department of Radiology, Mayo Clinic, 200 1st St SW, Rochester, MN 55905.

Abstract

Minimizing bias is critical to adoption and implementation of machine learning (ML) in clinical practice. Systematic mathematical biases produce consistent and reproducible differences between the observed and expected performance of ML systems, resulting in suboptimal performance. Such biases can be traced back to various phases of ML development: data handling, model development, and performance evaluation. This report presents 12 suboptimal practices during data handling of an ML study, explains how those practices can lead to biases, and describes what may be done to mitigate them. Authors employ an arbitrary and simplified framework that splits ML data handling into four steps: data collection, data investigation, data splitting, and feature engineering. Examples from the available research literature are provided. A Google Colaboratory Jupyter notebook includes code examples to demonstrate the suboptimal practices and steps to prevent them. Keywords: Data Handling, Bias, Machine Learning, Deep Learning, Convolutional Neural Network (CNN), Computer-aided Diagnosis (CAD) © RSNA, 2022.

Keywords: Bias; Computer-aided Diagnosis (CAD); Convolutional Neural Network (CNN); Data Handling; Deep Learning; Machine Learning.