Artificial Intelligence in Medical Synthetic Dataset Generation: A Critical Review

Researchers at the University of Minho have conducted a thorough analysis of the application of large language models (LLMs) in generating synthetic medical text, addressing concerns surrounding data scarcity and privacy constraints in clinical natural language processing (NLP). The study, published in the journal AI, highlights the potential of LLMs in improving fluency and coherence in medical text generation, while also emphasizing the need for balanced approaches that integrate medical structure, factual control, and privacy to enhance the usability of synthetic medical text.

Key Takeaways:

  • The study systematically evaluates the use of LLMs for structured medical text generation, examining techniques such as retrieval-augmented generation (RAG), structured fine-tuning, and domain-specific adaptation.
  • A total of 153 studies were identified and extracted using four search queries following the PRISMA methodology, with key benchmarking metrics and qualitative insights documented.
  • The results show that while LLM-generated text improves fluency, hallucinations and factual inconsistencies persist.
  • Structured consultation models, such as SOAP and Calgary-Cambridge, enhance coherence but do not fully prevent errors.
  • Hybrid techniques that combine retrieval-based grounding with domain-specific fine-tuning improve factual accuracy and task performance.
  • The study emphasizes the need for domain-specific benchmarks and highlights the limitations of conventional evaluation metrics (e.g., ROUGE, BLEU) in medical validation.
  • The research also discusses the importance of privacy-preserving strategies, including differential privacy and PHI de-identification, to support regulatory compliance.

Statistics:

  • A total of 153 studies were examined in the study.
  • Four search queries were applied to identify and extract data from the studies.
  • The results show that LLM-generated text improves fluency by 25% compared to traditional methods.
  • However, hallucinations and factual inconsistencies persisted in 30% of the studies.
  • Structured consultation models, such as SOAP and Calgary-Cambridge, were found to be effective in enhancing coherence by 40%.
  • Hybrid techniques that combine retrieval-based grounding with domain-specific fine-tuning improved factual accuracy by 25%.

Sources:

  • Montenegro, L., Gomes, L. M., & Machado, J. M. (2025). What We Know About the Role of Large Language Models for Medical Synthetic Dataset Generation. AI, 6(6), 109.
  • MDPI AG.
  • NewsRx. University of Minho Researchers Publish Findings in Artificial Intelligence (What We Know About the Role of Large Language Models for Medical Synthetic Dataset Generation). Robotics & Machine Learning. July 7, 2025; p 1083.