Large Language Models Show Promise in Alleviating Clinical Documentation Burden
A new study from the Northwestern University Feinberg School of Medicine has shed light on the potential of large language models (LLMs) in reducing the clinical documentation burden in healthcare. The research, which was published in the Journal of Clinical Neuroscience, analyzed clinical notes from 88 patients undergoing spine surgery and found that LLMs can efficiently extract relevant data from operative reports. The study's findings have significant implications for healthcare providers, who often struggle with the time-consuming task of documenting patient information.
Key Takeaways:
- The clinical documentation burden has increased exponentially since the Health Information Technology for Economic and Clinical Health Act in 2009.
- Large language models (LLMs) offer substantial potential for alleviating this burden by analyzing clinical notes.
- The study used five GPT models (4o, o1, o3 low reasoning, o3medium reasoning, o3 high reasoning) to extract data from operative reports and found no significant performance differences among the models.
- Reviewer agreement for the presence of certain variables varied considerably (52%-99%).
- The lowest duration and cost were observed in the o3-low model, taking 20.032 seconds and costing $0.008, respectively.
- Further work is necessary to extend and validate these models in other institutions or procedures.
- The study highlights the potential of LLMs in reducing the clinical documentation burden and improving the efficiency of healthcare providers.
Statistics:
- 88 patients undergoing lumbar decompression and/or fusion surgery in 2022 for lumbar spondylolisthesis were retrospectively reviewed.
- Five GPT models were used to extract data from operative reports, with no significant performance differences observed among the models.
- Reviewer agreement for the presence of certain variables varied considerably (52%-99%).
- The o3-low model took 20.032 seconds and cost $0.008 to process.
- The study aimed to validate GPT-based LLMs for extracting clinical data from spine surgery operative reports.
Sources:
- "Automated review of spine surgery operative reports with large language models: a pilot study of GPT reasoning models." Journal of Clinical Neuroscience, 2025;142:111648.
- Rushmin Khazanchi, Northwestern University Feinberg School of Medicine, Northwestern University, 420 E Superior St, Chicago, IL 60611, United States.
- Elsevier Sci Ltd, 125 London Wall, London, England.
- Journal of Clinical Neuroscience, www.journals.elsevier.com/journal-of-clinical-neuroscience/
- Northwestern University Feinberg School of Medicine, www.feinberg.northwestern.edu