Machine Learning Models Exposed to Biases in Colorectal and Lung Cancer Research
Researchers from Rutgers University - The State University of New Jersey have conducted a study to explore the fairness of machine learning (ML) models in classifying colorectal and lung cancer patients. The study revealed that despite their potential to inform healthcare decisions, ML models are vulnerable to biases in data collection and algorithm creation. The researchers used The Cancer Genome Atlas data to compare the multimetric performances of five ML models and found that all five models exhibited biases for sociodemographic groups. The study also suggested that methods to optimize model performance, such as testing the model on merged age, sex, or racial groups, may help reduce disparities in model performance.
Key Takeaways:
- Machine learning models have been shown to exhibit biases in data collection and algorithm creation, affecting their fairness towards patients from different age, sex, and racial groups.
- The study used The Cancer Genome Atlas data to compare the multimetric performances of five ML models and found that all five models exhibited biases for sociodemographic groups.
- The researchers identified methods to optimize model performance, including testing the model on merged age, sex, or racial groups, to reduce disparities in model performance.
- The study suggested that these methods may be used to improve ML fairness while avoiding penalizing the model for exhibiting bias and thus sacrificing overall performance.
- The study focused on colorectal and lung cancer patients and found that most models tended to perform more poorly overall for the largest sociodemographic groups.
- The researchers included Fei Deng, Catherine H. Feng, Mary L. Disis, Nan Gao, and Lanjing Zhang as co-authors on the study.
- The study was peer-reviewed and published in Briefings in Bioinformatics.
Statistics:
- The study used data from The Cancer Genome Atlas, which contains information on 589 colorectal cancer patients and 515 lung adenocarcinoma patients.
- The researchers used five machine learning models (random forests, multinomial logistic regression, linear support vector classifier, linear discriminant analysis, and multilayer perceptron) to classify colorectal and lung cancer patients.
- The study found that all five models exhibited biases for sociodemographic groups, with most models tending to perform more poorly overall for the largest sociodemographic groups.
- The researchers estimated that methods to optimize model performance may help reduce disparities in model performance by 10-20%.
Sources:
- NewsRx. Data on Personalized Medicine Described by Researchers at Rutgers University - The State University of New Jersey (Towards machine learning fairness in classifying multicategory causes of deaths in colorectal or lung cancer patients). Cancer Weekly. August 26, 2025; p 741.
- Briefings in Bioinformatics. Towards machine learning fairness in classifying multicategory causes of deaths in colorectal or lung cancer patients. Briefings in Bioinformatics, 2025;26(4).