Artificial Intelligence Shows Promise in Detecting Depressive Symptoms
Recent advances in artificial intelligence, particularly large language models (LLMs), have demonstrated potential for mental health applications, including the automated detection of depressive symptoms from natural language. Researchers at Psychiatric University Hospital have fine-tuned a German BERT-based LLM to predict individual Montgomery-Asberg Depression Rating Scale (MADRS) scores using a regression approach. The fine-tuned model achieved a mean absolute error of 0.7-1.0 across items, with accuracies ranging from 79 to 88%, closely matching clinician ratings. This research demonstrates the potential of lightweight LLMs to accurately assess depressive symptom severity, offering a scalable tool for clinical decision-making and monitoring treatment progress, particularly in low-resource settings.
Key Takeaways:
- Researchers at Psychiatric University Hospital have fine-tuned a German BERT-based large language model (LLM) to predict individual Montgomery-Asberg Depression Rating Scale (MADRS) scores using a regression approach.
- The fine-tuned model achieved a mean absolute error of 0.7-1.0 across items, with accuracies ranging from 79 to 88%, closely matching clinician ratings.
- Fine-tuning resulted in a 75% reduction in prediction errors relative to the untrained model.
- The study demonstrates the potential of lightweight LLMs to accurately assess depressive symptom severity.
- The model has the potential to be used as a scalable tool for clinical decision-making and monitoring treatment progress, particularly in low-resource settings.
- The research was conducted by Samantha Weber, Nicolas Deperrois, Robert Heun, Laura Fruhschutz, Anna Monn, Stephanie Homan, Andrea Hafliger, Erich Seifritz, Tobias Kowatsch, and Birgit Kleim.
- The study was supported by the Schweizerischer Nationalfonds Zur Forderung Der Wissenschaftlichen Forschung.
Statistics:
- Mean absolute error of 0.7-1.0 across items
- Accuracies ranging from 79 to 88%
- 75% reduction in prediction errors relative to the untrained model
Sources:
- Weber, S., Deperrois, N., Heun, R., Fruhschutz, L., Monn, A., Homan, S., Hafliger, A., Seifritz, E., Kowatsch, T., and Kleim, B. (2025). Using a fine-tuned large language model for symptom-based depression evaluation. npj Digital Medicine, 8(1), 1-11. doi: 10.1038/s41746-025-01982-8
- Health & Medicine Week. (2025, October 31). Researchers at Psychiatric University Hospital Report Research in Health and Medicine (Using a fine-tuned large language model for symptom-based depression evaluation). p 6153.