Artificial Intelligence in Patient Education: Evaluating DeepSeek-V3 and ChatGPT-4o

New research suggests that artificial intelligence-based large language models (AI-based LLMs) are increasingly used for obtaining medical information, but their accuracy and reliability remain a topic of debate. A study published in the ANZ Journal of Surgery compared the performance of two AI-based LLMs, DeepSeek-V3 and ChatGPT-4o, in providing medical insights to patients. The results showed that DeepSeek-V3 provided statistically significantly more suitable responses compared to ChatGPT-4o.

Key Takeaways:

  • The study aimed to evaluate the appropriateness, accuracy, and readability of responses provided by DeepSeek-V3 and ChatGPT-4o AI-based LLMs for patient education regarding laparoscopic cholecystectomy (LC).
  • The 20 most frequently asked questions by patients regarding LC were presented to the DeepSeek-V3 and ChatGPT-4o chatbots, and the comprehensiveness of their responses was evaluated by two board-certified general surgeons using a Likert scale.
  • DeepSeek-V3 provided statistically significantly more suitable responses compared to ChatGPT-4o, with a 5-point rating for 19 out of 20 questions (95%), whereas ChatGPT-4o achieved a 5-point rating for only 13 questions (65%).
  • The study concluded that DeepSeek-V3 provides more suitable responses to patient inquiries regarding LC compared to ChatGPT-4o.
  • The research has been peer-reviewed and published in the ANZ Journal of Surgery.

Statistics:

  • 20 questions were presented to both DeepSeek-V3 and ChatGPT-4o chatbots for evaluation.
  • DeepSeek-V3 received a 5-point rating for 19 out of 20 questions (95%), while ChatGPT-4o achieved a 5-point rating for 13 questions (65%).
  • The Paired sample t-test and Wilcoxon signed rank test were used to analyze the data, which showed a statistically significant difference between the two models (p = 0.033).
  • Inter-rater reliability was analyzed with Cohen's Kappa test.

Sources:

  • Evaluating Artificial Intelligence in Patient Education: DeepSeek-V3 Versus ChatGPT-4o in Answering Common Questions on Laparoscopic Cholecystectomy (2025)
  • ANZ Journal of Surgery, 2025
  • ANZ Journal of Surgery can be contacted at: Wiley, 111 River St, Hoboken 07030-5774, NJ, USA
  • (Wiley-Blackwell - www.wiley.com/; ANZ Journal of Surgery - onlinelibrary.wiley.com/journal/10.1111/(ISSN)1445-2197)
  • Dogukan Dogu, Dept. of General Surgery, Sincan Training and Research Hospital, Ankara, Turkey