Language Models in the Natural Sciences: Unraveling the Chemical Understanding of Transformer CLMs
Scientists at the University of Bonn have published a study in the journal Patterns that demonstrates the limitations of chemical language models (CLMs) in understanding biochemical relationships. The researchers found that transformer CLMs are trained to recognize patterns in data and make predictions based on statistical correlations, rather than acquiring knowledge of biochemical principles. This suggests that these models are not capable of understanding the underlying chemistry of a problem.
Key Takeaways:
- The study focused on transformer CLMs, which are a type of AI algorithm trained on vast quantities of text data, similar to ChatGPT and Google Gemini.
- CLMs are used in pharmaceutical research to predict active molecules based on amino acid sequences of target proteins, but the University of Bonn researchers found that these models lack a deeper chemical understanding.
- The researchers systematically manipulated the training data to test the CLMs' ability to predict active compounds, finding that the models relied on statistical correlations rather than biochemical principles.
- The models did not learn to distinguish between functionally important and unimportant sequence parts, and instead simply repeated what they had read before.
- The results suggest that CLMs may not be capable of understanding the underlying chemistry of a problem, but may instead recognize similarities in text-based molecular representations.
- The study's findings do not discredit the results of CLMs, but rather highlight the need for careful interpretation of their predictions.
Statistics:
- 50-60% amino acid sequence similarity was considered sufficient for enzymes to be considered similar by the CLMs.
- The researchers could randomize and scramble the sequences at will, as long as sufficient original amino acids were retained.
- 10% of the amino acid sequence was found to be sufficient for enzymes to perform their task.
- The study found that the CLMs were not capable of learning to distinguish between functionally important and unimportant sequence parts.
Sources:
- "Unraveling learning characteristics of transformer models for molecular design" by Jannik P. Roth, Jurgen Bajorath, published in Patterns (https://doi-org.sdpl.idm.oclc.org/10.1016/j.patter.2025.101392)
- University of Bonn press release (https://www.cell.com/patterns/fulltext/S2666-3899(25)00240-5)