GPT-4's Performance in Supporting Physician Decision-Making in Nephrology

According to research published in Scientific Reports, Generative Pre-trained Transformer (GPT)-4, a conversational artificial intelligence, showed potential in assisting physicians with nephrology questions. The study evaluated the performance of GPT-4 in assisting junior and senior physicians with nephrology questions, with surprising results. GPT-4 significantly improved the accuracy of junior physicians, but may have had a negative impact on senior physicians in specific subfields. Careful consideration is required when using GPT-4 to support physicians' decision-making.

Key Takeaways:

  • GPT-4 correctly answered 77.8% of nephrology questions, significantly improving junior physicians' accuracy from 53.3% to 72.2% and senior physicians' accuracy from 65.6% to 75.6%.
  • The improvement was significantly higher for junior physicians, especially in clinical categories.
  • Senior physicians showed a decreased proportion of correct answers in one of the clinical categories after using GPT-4 support.
  • GPT-4's performance varied by subfield, with negative impacts observed in specific areas.
  • The study suggests careful consideration is required when using GPT-4 to support physicians' decision-making in nephrology.
  • The research involved 45 single-answer multiple-choice questions extracted from the Core Curriculum in Nephrology articles published in the American Journal of Kidney Diseases between October 2021 and June 2023.
  • The study was conducted by researchers from St. Marianna University School of Medicine, with additional authors including Kenichiro Tanabe, Daisuke Ichikawa, and Yugo Shibagaki.

Statistics:

  • 77.8% of nephrology questions were answered correctly by GPT-4.
  • 53.3% of questions were answered correctly by junior physicians with no GPT-4 support.
  • 72.2% of questions were answered correctly by junior physicians with GPT-4 support.
  • 65.6% of questions were answered correctly by senior physicians with no GPT-4 support.
  • 75.6% of questions were answered correctly by senior physicians with GPT-4 support.
  • The improvement was significantly higher for junior physicians, especially in clinical categories (p = 0.017).

Sources:

  • GPT-4's performance in supporting physician decision-making in nephrology multiple-choice questions. Scientific Reports, 2025, 15(1):1-8. (Scientific Reports - http://www.nature.com/srep/index.html). The publisher for Scientific Reports is Nature Portfolio.
  • https://doi-org.sdpl.idm.oclc.org/10.1038/s41598-025-99774-3.