Merck & Company Reports Breakthrough in Biotechnology Research with Deep Learning Models

Research conducted by Merck & Company has led to a significant advancement in the field of medicinal chemistry, with the development of predictive models that enable faster discovery of safer and more efficacious therapeutics. The study highlights the potential of deep learning models, particularly graph neural networks, in improving model performance for drug discovery. By leveraging large internal data sets and publicly available data sets, the researchers identified factors that contribute to the higher performance of predictive models built using graph neural networks compared to traditional methods. This breakthrough has the potential to revolutionize the field of biotechnology and improve the discovery of new medicines.

Key Takeaways:

  • The research found that graph neural networks outperform traditional methods such as XGBoost and random forest in building predictive models for drug discovery.
  • The study identified factors that contribute to the higher performance of predictive models built using graph neural networks, including the ability to handle complex biological data.
  • The researchers developed a scaling relationship that explains 81% of the variance in model performance across various assays and data regimes, which can be used to estimate the performance of models for ADMET end points.
  • The findings offer guidance for further improving model performance in drug discovery, which is critical for the rapid development of new medicines.
  • The study used a combination of internal and publicly available data sets, demonstrating the potential of collaborative research in advancing biotechnology.
  • The research has been peer-reviewed and published in the Journal of Chemical Information and Modeling.

Statistics:

  • 81%: The percentage of variance in model performance explained by the developed scaling relationship.
  • 3: The number of different property spaces tested in model extrapolation tasks.
  • 4: The number of tasks used to assess model performance, including random, temporal, and reverse-temporal data ablation tasks.
  • 15: The number of data sets used in the study, including several publicly available data sets.
  • 2025: The year in which the research was conducted and published.

Sources:

  • "Data Scaling and Generalization Insights for Medicinal Chemistry Deep Learning Models." Journal of Chemical Information and Modeling, 2025.
  • American Chemical Society (ACS) - www.acs.org
  • Journal of Chemical Information and Modeling - www.pubs.acs.org/journal/jcisd8
  • Merck & Company Inc. - Modeling & Informatics, South San Francisco, California 94080, United States - www.merck.com