Automated Benefit Extraction in Translational Science Research
Researchers at the National Center for Advancing Translational Sciences have developed an automated approach to extracting benefits from large research datasets, using Natural Language Processing (NLP) and Artificial Intelligence (AI). The study, which analyzed over 1 million research outputs from the Clinical and Translational Science Award (CTSA) program, found that clinical and community benefits were the most frequently extracted benefits from publications and projects, while economic and policy benefits were less frequently identified. The automated pipeline enabled a data-driven approach to applying the Translational Science Benefits Model (TSBM) at the scale of the entire CTSA program outputs.
Key Takeaways:
- The Clinical and Translational Science Award (CTSA) program, funded by the National Center for Advancing Translational Sciences (NCATS), has supported over 65 hubs, generating 118,490 publications from 2006 to 2021.
- Measuring the impact of these outputs remains challenging, as traditional bibliometric methods fail to capture patents, policy contributions, and clinical implementation.
- The study developed an NLP-driven pipeline that automates the extraction of TSBM benefits from research outputs using Latent Dirichlet Allocation (LDA) topic modeling.
- The analysis corpus comprised 1,296 projects, 127,958 publications, and 352 patents, spanning CTSA hub grants awarded from 2006 to 2023.
- Clinical and community benefits were the most frequently extracted benefits from publications and projects, reflecting the patient-centered and community-driven nature of CTSA research.
- Economic and policy benefits were less frequently identified, prompting the inclusion of patent data to better capture commercialization impacts.
- The Publications LDA Model proved the most effective for benefit extraction for publications and projects.
- All patents were automatically tagged as economic benefits, given their intrinsic focus on commercialization and in accordance with TSBM guidelines.
Statistics:
- 65 CTSA hubs have been supported since 2006, generating 118,490 publications.
- 1,296 projects, 127,958 publications, and 352 patents were analyzed in the study.
- Clinical and community benefits were the most frequently extracted benefits from publications and projects, accounting for 71% and 61% of all benefits, respectively.
- Economic benefits were less frequently identified, accounting for 13% of all benefits.
- The study used LDA topic modeling to analyze 1.3 million texts from the CTSA corpus.
Sources:
- Topic analysis on publications and patents toward fully automated translational science benefits model impact extraction. Frontiers in Research Metrics and Analytics, 2025;10:1596687.
- Reports Outline Library Science Study Findings from Eline Appelmans and Colleagues (Topic analysis on publications and patents toward fully automated translational science benefits model impact extraction). Information Technology Newsweekly. October 21, 2025; p 645.