Automated Evaluation of Electronic Health Records Using Large Language Models

Advanced technologies have been developed to summarize large amounts of clinical data stored in Electronic Health Records (EHRs), a process that can be overwhelming for healthcare providers. Recent breakthroughs in Generative AI with Large Language Models have enabled the summarization of patient records into actionable insights, reducing the cognitive burden on providers. However, the quality of these summaries needs to be evaluated, a task typically performed by human experts. This approach is not only time-consuming but also costly. To address this, researchers have introduced an automated method for evaluating real-world EHR multi-document summaries using a Large Language Model (LLM) as the evaluator, referred to as LLM-as-a-Judge. This innovative framework has demonstrated strong inter-rater reliability with human evaluators.

Key Takeaways:

  • The proposed LLM-as-a-Judge framework uses GPT-o3-mini as the evaluator and achieves high inter-rater reliability with human evaluators, with an intraclass correlation coefficient of 0.818 (95% CI 0.772, 0.854).
  • The framework completes evaluations in just 22 seconds, making it a time-efficient solution.
  • LLM-as-a-Judge excels in inter-rater reliability, particularly in evaluations requiring advanced reasoning and domain expertise, outperforming non-reasoning models, those trained on the task, and multi-agent workflows.
  • The framework is scalable and efficient, enabling the rapid identification of accurate and safe AI-generated summaries in healthcare settings.
  • Cross-task validation on the Problem Summarization task confirmed high reliability of the LLM-as-a-Judge framework.
  • The proposed method can be used to validate the quality of AI-generated summaries in various healthcare settings.

Statistics:

  • The intraclass correlation coefficient of GPT-o3-mini is 0.818 (95% CI 0.772, 0.854), indicating excellent inter-rater reliability.
  • The median score difference between LLM-as-a-Judge and human evaluators is 0, indicating high accuracy.
  • The LLM-as-a-Judge framework completes evaluations in just 22 seconds.
  • The proposed method can be applied to various healthcare settings, including those requiring advanced reasoning and domain expertise.

Sources:

  • medrxiv.org/content/10.1101/2025.04.22.25326219v2