Artificial Intelligence in Facial Plastic Surgery: A Comparative Evaluation of Multimodal Large Language Models
Facial analysis is crucial for preoperative planning in facial plastic surgery, but traditional methods can be time-consuming and subjective. A recent study from Mayo Clinic investigated the potential of Artificial Intelligence (AI) in objective and efficient facial analysis using Multimodal Large Language Models (MLLMs). The research evaluated the ability of four MLLMs to analyze facial skin quality, volume, symmetry, and adherence to aesthetic standards. The findings showed promise in analyzing qualitative features, but the MLLMs struggled with precise quantitative measurements of facial ratios. The study concluded that current general-purpose MLLMs are not yet ready to replace manual clinical assessments but may assist in general facial feature analysis.
Key Takeaways:
- Four MLLMs (ChatGPT-4o, ChatGPT-4, Gemini 1.5 Pro, and Claude 3.5 Sonnet) were evaluated for their ability to analyze facial skin quality, volume, symmetry, and adherence to aesthetic standards.
- The MLLMs showed promise in analyzing qualitative features, but they struggled with precise quantitative measurements of facial ratios.
- The study used two evaluation forms and 15 diverse facial images generated by a Generative Adversarial Network (GAN) to assess the MLLMs' ability to analyze facial skin quality, volume, symmetry, and adherence to aesthetic standards.
- The MLLMs' performance was compared to evaluations from a plastic surgeon and manual measurements of facial ratios.
- The study found that current general-purpose MLLMs are not yet ready to replace manual clinical assessments but may assist in general facial feature analysis.
Statistics:
- Mean accuracy for general analysis was:
* ChatGPT-4o: 0.61 +/- 0.49
* Gemini 1.5 Pro: 0.60 +/- 0.49
* ChatGPT-4: 0.57 +/- 0.50
* Claude 3.5 Sonnet: 0.52 +/- 0.50
- Mean accuracy for facial ratio assessments was:
* Gemini 1.5 Pro: 0.39 +/- 0.49
- Inter-rater reliability based on Cohen's Kappa values ranged from poor to high for qualitative assessments (kappa 0.7 for some questions) but was generally poor (near or below zero) for quantitative assessments.
Sources:
- NewsRx. Findings from Mayo Clinic in the Area of Artificial Intelligence Described (Facial Analysis for Plastic Surgery In the Era of Artificial Intelligence: a Comparative Evaluation of Multimodal Large Language Models). Medical Devices & Surgical Technology Week. June 29, 2025; p 299.
- Facial Analysis for Plastic Surgery In the Era of Artificial Intelligence: a Comparative Evaluation of Multimodal Large Language Models. Journal of Clinical Medicine, 2025;14(10):3484.
- Mdpi, St Alban-Anlage 66, Ch-4052 Basel, Switzerland.
- Antonio Jorge Forte, Mayo Clinic, Division of Plastic Surgery, Jacksonville, FL 32224, United States.