Large Language Models Outperform Human Readers in Coronary CT Angiography Report Categorization

Researchers at Yonsei University College of Medicine have found that large language models (LLMs) can accurately categorize coronary artery disease (CAD) using coronary CT angiography (CCTA) reports, outperforming human readers in some cases. The study evaluated the accuracy of four LLMs and four human readers in assigning CAD-RADS 2.0 categories and modifiers based on real-world CCTA reports. The results showed that LLMs, particularly O1, demonstrated high accuracy in full CAD-RADS categorization, with O1 achieving 90.7% accuracy in the internal validation dataset and 95.7% accuracy in the external validation dataset.

Key Takeaways:

  • Large language models (LLMs) can accurately categorize coronary artery disease (CAD) using coronary CT angiography (CCTA) reports.
  • The study evaluated the accuracy of four LLMs and four human readers in assigning CAD-RADS 2.0 categories and modifiers based on real-world CCTA reports.
  • O1, an LLM, demonstrated the highest accuracy in full CAD-RADS categorization, with 90.7% accuracy in the internal validation dataset.
  • In the external validation dataset, O1 achieved 95.7% accuracy for full CAD-RADS categorization.
  • LLMs exhibited similar or higher accuracy and shorter processing times compared to human readers for full CAD-RADS 2.0 categorization.
  • The study concluded that LLMs, particularly O1, can be a viable option for CAD-RADS 2.0 categorization based on CCTA reports.
  • The research has been peer-reviewed and published in the Journal of Imaging Informatics In Medicine.

Statistics:

  • 2752 CCTA reports were generated at an academic hospital between January and September 2024.
  • 180 CCTA reports were randomly selected for the study, with 90 reports for internal validation and 90 reports for external validation.
  • 327 CCTA reports were used for external validation, derived from two independent institutions.
  • O1 demonstrated 90.7% accuracy in full CAD-RADS categorization in the internal validation dataset.
  • O1 achieved 95.7% accuracy for full CAD-RADS categorization in the external validation dataset.
  • LLMs, including O1, processed reports with average times ranging from 1.34 s to 16.61 s.
  • Human readers required 32.10 s to 55.06 s to complete the same task.

Sources:

  • Son J, Yoo WS, Kim JY, Park JH, Park HJ, Kim C, Choi BW, Suh YJ. Large Language Models Versus Human Readers in CAD-RADS 2.0 Categorization of Coronary CT Angiography Reports. Journal of Imaging Informatics In Medicine, 2025.
  • NewsRx. New Heart Disease Findings from Yonsei University College of Medicine Described (Large Language Models Versus Human Readers in CAD-RADS 2.0 Categorization of Coronary CT Angiography Reports). Cardiovascular Week. October 20, 2025; p 88.