Artificial Intelligence Researchers Develop New Approach for Analyzing Extremist Content in Kazakh Language
Kazakhstan-based researchers from Al-Farabi Kazakh National University (KazNU) have made significant strides in the realm of artificial intelligence by developing a novel approach for identifying and classifying extremist ideologies in the Kazakh language. This breakthrough comes at a critical time as online data continues to grow exponentially, highlighting the urgent need for effective tools to detect and mitigate extremist content. By harnessing the power of machine learning and natural language processing, the researchers have created a robust framework that can analyze and categorize extremist content with impressive accuracy.
Key Takeaways:
- The researchers developed a hybrid approach that combines traditional text vectorization techniques, machine learning algorithms, and a psycholinguistic analysis module (PLAM) specifically adapted for the Kazakh language.
- The experimental results demonstrate the effectiveness of the hybrid approach, with the combination of CountVectorizer + Logistic Regression + PLAM achieving the highest performance among traditional models (F1-score: 0.9305, Accuracy: 0.9308, ROC AUC: 0.9892).
- Among deep learning models, the BERT + LSTM model yielded the best results (F1-score: 0.9481, Accuracy: 0.9485, ROC AUC: 0.9918), followed by the standalone BERT model (F1-score: 0.9412, Accuracy: 0.9414, ROC AUC: 0.9901).
- The research provides an effective framework for multilingual text analysis, contributing to improved monitoring and prevention of extremist content in underrepresented languages, such as Kazakh.
- Future work will focus on refining these methods and exploring their application in other domains for robust content moderation and security in the digital space.
Statistics:
- The researchers achieved an F1-score of 0.9305 on traditional models using CountVectorizer + Logistic Regression + PLAM.
- Among deep learning models, the BERT + LSTM model yielded an F1-score of 0.9481, Accuracy of 0.9485, and ROC AUC of 0.9918.
- The standalone BERT model achieved an F1-score of 0.9412, Accuracy of 0.9414, and ROC AUC of 0.9901.
Sources:
- "Extremist Ideology Classification in Kazakh: A Multi-Class Approach Using Machine Learning and Psycholinguistic Analysis." IEEE Access, 2025, 13():140500-140518.
- IEEE (publisher) - http://ieeexplore.ieee.org/servlet/opac?punumber=6287639
- https://doi-org.sdpl.idm.oclc.org/10.1109/ACCESS.2025.3596601