Cyberbullying Detection in Social Media Using Natural Language Processing: A Study
Research conducted by a team from Fayoum University has focused on the increasing issue of cyberbullying on social media platforms. The study highlighted that social media has become a breeding ground for cyberbullying, with tweets and comments often causing significant emotional distress. To combat this issue, the researchers developed a model that combines Machine Learning (ML) classifiers with Natural Language Processing (NLP) techniques to identify and detect cyberbullying messages.
Key Takeaways:
- The study utilized a dataset of 39,870 Twitter posts and comments, categorized into five types of cyberbullying: religion, age, gender, ethnicity bullying, and non-cyberbullying.
- The proposed model aims to train ML classifiers after being processed using NLP techniques, with the goal of detecting cyberbullying messages early on.
- The implementation results showed that the Random Forest classifier achieved an accuracy rate of 94%, outperforming other classifiers such as Support Vector Machine, Logistic Regression, Naive Bayes, and K-Nearest Neighbor.
- The study concluded that the Random Forest classifier is the most effective in detecting cyberbullying messages, followed by the Support Vector Machine classifier.
- The study's findings have crucial implications for the development of effective strategies to combat cyberbullying on social media platforms.
- The researchers used a dataset of 39,870 tweets and comments to train and test their model, indicating a large and representative sample size.
- The study highlights the importance of using NLP techniques in conjunction with ML classifiers to improve the accuracy of cyberbullying detection.
Statistics:
- The study used a dataset of 39,870 Twitter posts and comments.
- The proposed model achieved accuracy rates of 94% for the Random Forest classifier, 93% for the Support Vector Machine classifier, 92% for the Logistic Regression classifier, 92% for the Naive Bayes classifier, and 73% for the K-Nearest Neighbor classifier.
- The study found that the Random Forest classifier outperformed all other classifiers in terms of accuracy.
- The study's findings have significant implications for the development of effective strategies to combat cyberbullying on social media platforms.
Sources:
- Cyberbullying detection in social media using natural language processing. Scientific African, 2025, 28(): e02713.
- Elsevier. (Publisher for Scientific African)
- https://www.journals.elsevier.com/scientific-african (Journal website)
- https://doi-org.sdpl.idm.oclc.org/10.1016/j.sciaf.2025.e02713 (Journal article DOI)