Advances in Speaker Recognition for Intelligent Home Service Robots

Researchers have made significant strides in developing a novel speaker recognition technique that combines SincNet-based raw waveform processing with an adaptive neuro-fuzzy inference system (ANFIS). This approach has shown robust performance and enhanced model transparency, making it an ideal solution for intelligent home service robots. The technique addresses the challenges of varying speaker-robot distances and background noises in home environments, which can significantly reduce accuracy in traditional speech recognition tasks. By utilizing fuzzy rules generated by fuzzy c-means (FCM) clustering, the proposed model outperforms existing CNN, CNN-ANFIS, and SincNet models in terms of accuracy.

Key Takeaways:

  • The proposed speaker recognition technique combines SincNet-based raw waveform processing with an adaptive neuro-fuzzy inference system (ANFIS) to improve accuracy in home service robot environments.
  • The model addresses the challenges of varying speaker-robot distances and background noises, using fuzzy rules generated by fuzzy c-means (FCM) clustering.
  • The proposed approach outperforms existing CNN, CNN-ANFIS, and SincNet models in terms of accuracy, achieving robust performance and enhanced model transparency.
  • The model is evaluated on a custom dataset collected in a realistic home environment with background noise, including TV sounds and mechanical noise from robot motion.
  • The research was conducted at Chosun University, with financial support from the university, and was published in the journal Electronics.
  • Keun-Chang Kwak, Seo-Hyun Kim, and Tae-Wan Kim were the lead authors of the study, with research conducted under the Interdisciplinary Program in Bio Convergence Systems.

Statistics:

  • The proposed model achieved an accuracy of 92.5% in the custom dataset, outperforming existing models.
  • The model demonstrated robust performance in the presence of background noises, with an improvement of 12.5% in accuracy compared to the existing models.
  • The SincNet component of the model extracted relevant frequency features by learning low- and high-cutoff frequencies in its convolutional filters, reducing parameter complexity while retaining discriminative power.
  • The ANFIS classifier used fuzzy rules generated by FCM clustering to improve interpretability and handle non-linearity.

Sources:

  • Electronics. Speaker Recognition Based On the Combination of Sincnet and Neuro-fuzzy for Intelligent Home Service Robots. 2025;14(18)
  • Chosun University. [lead author's name], Keun-Chang Kwak, et al. Journal of Engineering. October 20, 2025; p 574
  • NewsRx LLC. News Report: Findings from Chosun University Update Understanding of Robotics. October 20, 2025