Large Language Models Show Promise in Measuring Clinical Research Outcomes
Researchers at the University of Washington have assessed the performance of a publicly available large language model (LLM) in identifying documented goals-of-care discussions in clinical notes. The study compared the performance of the LLM, Llama 3.3, using zero-shot prompting, with a task-specific BERT-based model trained on manually annotated notes. The findings suggest that the LLM can measure novel clinical research outcomes with fewer or no task-specific training data.
Key Takeaways:
- The study used a publicly available LLM, Llama 3.3, to identify documented goals-of-care discussions in clinical notes.
- The LLM was evaluated on records from a series of clinical trials enrolling adult patients with chronic life-limiting illness hospitalized over 2018-2023.
- The study compared the performance of the LLM with a task-specific BERT-based model trained on 4,642 manually annotated notes.
- The LLM achieved promising results in measuring documented goals-of-care discussions, with an area under the receiver operating characteristic curve (AUC) of 0.83.
- The study evaluated the area under the precision-recall curve (AUPRC) and maximal F score for both note-level and patient-level classification over a 30-day period.
- The researchers found that the LLM outperformed the task-specific BERT-based model in identifying documented goals-of-care discussions.
- The study highlights the potential of LLMs in measuring novel clinical research outcomes with fewer or no task-specific training data.
- The research was conducted by Robert Y. Lee, Kevin S. Li, James Sibley, Trevor Cohen, William B. Lober, Danae G. Dotolo, and Erin K. Kross at the University of Washington.
- The study's findings have implications for the use of LLMs in palliative care and clinical research workflows.
Statistics:
- 4,642 manually annotated notes used to train the task-specific BERT-based model.
- 2018-2023: Time period over which clinical trials were conducted.
- 0.83: Area under the receiver operating characteristic curve (AUC) achieved by the LLM in identifying documented goals-of-care discussions.
- 30 days: Time period over which note-level and patient-level classification was evaluated.
- 0.71: Area under the precision-recall curve (AUPRC) achieved by the LLM in identifying documented goals-of-care discussions.
Sources:
- Assessment of a zero-shot large language model in measuring documented goals-of-care discussions. Journal of Pain and Symptom Management, 2025.
- University of Washington.