Artificial Intelligence Enhances Systematic Literature Review with Machine Learning

Researchers at Intelligent Medical Objects have conducted a new study examining the application of deep learning algorithms in systematic literature reviews, finding that machine learning models significantly outperform conventional ones in extracting relevant data elements. The study involved collecting and annotating a large dataset of full-text articles and applying Natural Language Processing (NLP) approaches to automatically extract data elements. The results show that deep learning algorithms, particularly Long Short-Term Memory (LSTM) models, achieved superior performance in recognizing targeted data elements compared to conventional machine learning algorithms.

Key Takeaways:

  • The researchers collected and annotated 239 full-text articles for 12 important variables, including study cohort, lab technique, and disease type, for proper systematic literature review summary of Human papillomavirus (HPV) Prevalence, Pneumococcal Epidemiology, and Pneumococcal Economic Burden.
  • The annotated corpora contain 4,498, 579, and 252 annotated entity mentions for HPV Prevalence, Pneumococcal Epidemiology, and Pneumococcal Economic Burden tasks respectively.
  • Deep learning algorithms achieved superior performance in recognizing the targeted systematic literature review data elements, compared to conventional machine learning algorithms.
  • LSTM models have achieved 0.890, 0.646, and 0.615 micro-averaged F1 scores for three tasks respectively, demonstrating their superiority in data element extraction tasks.
  • While Bidirectional Encoder Representations from Transformers (BERT) models are known to generally achieve superior performance on many NLP tasks, they did not exhibit improvement in this study's three tasks.

Statistics:

  • The researchers collected and annotated 239 full-text articles.
  • The annotated corpora contain 4,498, 579, and 252 annotated entity mentions for HPV Prevalence, Pneumococcal Epidemiology, and Pneumococcal Economic Burden tasks respectively.
  • The study involved three Classic Named Entity Recognition (NER) algorithms: Conditional Random Fields (CRF), Long Short-Term Memory (LSTM), and Bidirectional Encoder Representations from Transformers (BERT).
  • The LSTM model achieved 0.890, 0.646, and 0.615 micro-averaged F1 scores for three tasks respectively.

Sources:

  • NewsRx. New Machine Learning Research Reported from Intelligent Medical Objects (Use of deep learning-based NLP models for full-text data elements extraction for systematic literature review tasks). Health & Medicine Week. June 27, 2025; p 2557.
  • Jingcheng Du et al. Use of deep learning-based NLP models for full-text data elements extraction for systematic literature review tasks. Scientific Reports, 2025,15(1):1-8. (Scientific Reports - http://www.nature.com/srep/index.html).
  • Nature Portfolio. Scientific Reports.