Artificial Intelligence Study Reveals Accuracy and Reliability Concerns

A recent study published in the American Journal of Orthodontics and Dentofacial Orthopedics has shed light on the accuracy and reliability of artificial intelligence (AI) language models in providing medical information. The study, conducted by Istanbul University-Cerrahpasa researchers, evaluated the performance of three popular AI language models - ChatGPT-4, Gemini, and Copilot - in answering questions related to surgically-assisted rapid palatal expansion. The study found that while none of the AI models showed statistically significant differences in accuracy, ChatGPT-4 had the highest objectively true rate, while Gemini produced answers with more balanced accuracy, and Copilot had the highest number of false answers.

Key Takeaways:

  • The study aimed to evaluate the accuracy, reliability, and comprehensibility of information provided by AI language models on surgically-assisted rapid palatal expansion.
  • 115 questions were created by 3 orthodontists and 1 oral and maxillofacial surgeon, and the accuracy of the answers was independently evaluated by the same experts via a 5-point Likert scale.
  • The study found that ChatGPT-4 had the highest objectively true rate, while Gemini produced answers with more balanced accuracy, and Copilot had the highest number of false answers.
  • Although there were no statistically significant differences between the AI models, the study suggests that the accuracy of AI-supported language models may vary according to subject matter.
  • The study highlights the importance of critically evaluating the accuracy and reliability of AI-generated information, particularly in medical contexts.
  • The researcher's findings have implications for the development and use of AI language models in healthcare, highlighting the need for more rigorous evaluation and testing.

Statistics:

  • 115 questions were created by 3 orthodontists and 1 oral and maxillofacial surgeon.
  • The study evaluated the accuracy of the responses generated by the AI language models using a 5-point Likert scale.
  • ChatGPT-4 had the highest objectively true rate.
  • Gemini produced answers with more balanced accuracy.
  • Copilot had the highest number of false answers.
  • The study found that there were no statistically significant differences between the AI models.
  • The study suggests that the accuracy of AI-supported language models may vary according to subject matter.

Sources:

  • The role of artificial intelligence in providing accurate and reliable information on surgically-assisted rapid palatal expansion: A cross-sectional study. American Journal of Orthodontics and Dentofacial Orthopedics, 2025.
  • NewsRx LLC. Reports Outline Artificial Intelligence Study Results from Istanbul University-Cerrahpasa (The role of artificial intelligence in providing accurate and reliable information on surgically-assisted rapid palatal expansion: A cross-sectional study). Journal of Engineering. October 20, 2025; p 2621.