Artificial Intelligence Research Reveals Impact of Imbalanced and Overlapping Data on Customer Churn Prediction
Researchers at Khon Kaen University in Thailand have published a new study on artificial intelligence, highlighting the significance of imbalanced and overlapping data in customer churn prediction. The study reveals that various sampling methods, including hybrid sampling, can address these issues but struggle with classical machine learning algorithms. To optimize the performance of classical machine learning, the researchers introduced an extension framework called CostLearnGAN, a tabular generative adversarial network (GAN)-hybrid sampling method, and cost-sensitive learning.
Key Takeaways:
- Imbalanced and overlapping data significantly impact classification results in customer churn prediction.
- Various sampling and hybrid sampling methods have demonstrated effectiveness in addressing these issues.
- CostLearnGAN, a tabular GAN-hybrid sampling method, and cost-sensitive learning are introduced to optimize the performance of classical machine learning algorithms.
- Experimental results show that CostLearnGAN achieved a satisfying result across all evaluation metrics with a 1.44 average mean rank score.
- CostLearnGAN outperforms other sampling methods in improving the performance of classical machine learning models with a 5.68 robustness value on average.
- Classical machine learning algorithms exhibit shorter execution times, making them suitable for predicting churn in large customer bases.
- The study conducted an experiment with six comparative sampling methods, six datasets, and three machine learning algorithms.
- The research concludes that CostLearnGAN provides a robustness measurement for algorithms, outperforming other sampling methods in improving performance.
Statistics:
- 1.44: Average mean rank score achieved by CostLearnGAN across all evaluation metrics.
- 5.68: Average robustness value of CostLearnGAN in improving the performance of classical machine learning models.
- 6: Number of comparative sampling methods used in the experiment.
- 6: Number of datasets used in the experiment.
- 3: Number of machine learning algorithms used in the experiment.
Sources:
- "Research from Khon Kaen University Broadens Understanding of Machine Learning [Optimized customer churn prediction using tabular generative adversarial network (GAN)-based hybrid sampling method and cost-sensitive learning]." VerticalNews. 7 July 2025.
- Optimal customer churn prediction using tabular generative adversarial network (GAN)-based hybrid sampling method and cost-sensitive learning. PeerJ Computer Science, 2025, 11(): e2949.
- PeerJ Computer Science. [accessed 2025 Jul 7].
- NewsRx. [accessed 2025 Jul 7].