Artificial Intelligence in Patient Education: Evaluating DeepSeek-V3 and ChatGPT-4o
New research suggests that artificial intelligence-based large language models (AI-based LLMs) are increasingly used for obtaining medical information, but their accuracy and reliability remain a topic of debate. A study published in the ANZ Journal of Surgery compared the performance of two AI-based LLMs, DeepSeek-V3 and ChatGPT-4o, in providing medical insights to patients. The results showed that DeepSeek-V3 provided statistically significantly more suitable responses compared to ChatGPT-4o.
Key Takeaways:
- The study aimed to evaluate the appropriateness, accuracy, and readability of responses provided by DeepSeek-V3 and ChatGPT-4o AI-based LLMs for patient education regarding laparoscopic cholecystectomy (LC).
- The 20 most frequently asked questions by patients regarding LC were presented to the DeepSeek-V3 and ChatGPT-4o chatbots, and the comprehensiveness of their responses was evaluated by two board-certified general surgeons using a Likert scale.
- DeepSeek-V3 provided statistically significantly more suitable responses compared to ChatGPT-4o, with a 5-point rating for 19 out of 20 questions (95%), whereas ChatGPT-4o achieved a 5-point rating for only 13 questions (65%).
- The study concluded that DeepSeek-V3 provides more suitable responses to patient inquiries regarding LC compared to ChatGPT-4o.
- The research has been peer-reviewed and published in the ANZ Journal of Surgery.
Statistics:
- 20 questions were presented to both DeepSeek-V3 and ChatGPT-4o chatbots for evaluation.
- DeepSeek-V3 received a 5-point rating for 19 out of 20 questions (95%), while ChatGPT-4o achieved a 5-point rating for 13 questions (65%).
- The Paired sample t-test and Wilcoxon signed rank test were used to analyze the data, which showed a statistically significant difference between the two models (p = 0.033).
- Inter-rater reliability was analyzed with Cohen's Kappa test.
Sources:
- Evaluating Artificial Intelligence in Patient Education: DeepSeek-V3 Versus ChatGPT-4o in Answering Common Questions on Laparoscopic Cholecystectomy (2025)
- ANZ Journal of Surgery, 2025
- ANZ Journal of Surgery can be contacted at: Wiley, 111 River St, Hoboken 07030-5774, NJ, USA
- (Wiley-Blackwell - www.wiley.com/; ANZ Journal of Surgery - onlinelibrary.wiley.com/journal/10.1111/(ISSN)1445-2197)
- Dogukan Dogu, Dept. of General Surgery, Sincan Training and Research Hospital, Ankara, Turkey