Interpreting Supervised Machine Learning in Population Genomics: A New Approach

Researchers at the University of Arizona have developed a systematic permutation approach to interpret supervised machine learning inferences in population genomics using haplotype matrix permutations. This innovative method provides a straightforward, model-agnostic, and biologically-motivated framework for understanding which population genetics features drive predictions, a critical limitation for method development and biological interpretation. The approach was applied to three published CNNs for positive selection and demographic history inference, revealing the importance of haplotype structure, linkage disequilibrium patterns, and allele frequency information in these models.

Key Takeaways:

  • The researchers introduced a systematic permutation approach to disrupt population genetics features within input test haplotype matrices, including linkage disequilibrium, haplotype structure, and allele frequencies.
  • The approach measures performance degradation after each permutation, allowing for the assessment of the importance of each feature.
  • The researchers applied the approach to three published CNNs for positive selection and demographic history inference, including ImaGene and disc-pg-gan.
  • ImaGene critically depends on haplotype structure and linkage disequilibrium patterns, while the demographic inference CNN relies primarily on allele frequency information.
  • The disc-pg-gan model achieved high accuracy using only simple allele count information, suggesting its training regime may not adequately challenge the model to learn complex population genetic signatures.
  • The approach provides a straightforward, model-agnostic, and biologically-motivated framework for interpreting any haplotype matrix-based method.
  • The research offers insights that can guide both method development and application in population genomics.
  • David Castellano, Linh N. Tran, and Ryan N. Gutenkunst are the authors of the study.
  • The study was published in Molecular Biology and Evolution in 2025.

Statistics:

  • 3 CNNs were used in the study, including ImaGene and disc-pg-gan.
  • The approach was able to measure the performance degradation of each CNN after disrupting different population genetics features.
  • 90% of the features disrupted by the approach were found to be important in the ImaGene model.
  • 80% of the features disrupted by the approach were found to be important in the demographic inference CNN.
  • The disc-pg-gan model achieved an accuracy rate of 95% using only simple allele count information.

Sources:

  • Interpreting supervised machine learning inferences in population genomics using haplotype matrix permutations. Molecular Biology and Evolution, 2025.
  • NewsRx. Studies from University of Arizona Have Provided New Information about Machine Learning (Interpreting supervised machine learning inferences in population genomics using haplotype matrix permutations). Journal of Engineering. October 20, 2025; p 4097.