SMS Spam Detection and Classification to Combat Abuse in Telephone Networks Using Natural Language Processing

Oyeyemi, Dare Azeez and Ojo, Adebola K. (2023) SMS Spam Detection and Classification to Combat Abuse in Telephone Networks Using Natural Language Processing. Journal of Advances in Mathematics and Computer Science, 38 (10). pp. 144-156. ISSN 2456-9968

[thumbnail of Ojo38102023JAMCS108409.pdf]

Text
Ojo38102023JAMCS108409.pdf - Published Version
Download (529kB)

Official URL: https://doi.org/10.9734/jamcs/2023/v38i101832

Abstract

In the modern era, mobile phones have become ubiquitous, and Short Message Service (SMS) has grown to become a multi-million-dollar service due to the widespread adoption of mobile devices and the millions of people who use SMS daily. However, SMS spam has also become a pervasive problem that endangers users' privacy and security through phishing and fraud. Despite numerous spam filtering techniques, there is still a need for a more effective solution to address this problem [1]. This research addresses the pervasive issue of SMS spam, which poses threats to users' privacy and security. Despite existing spam filtering techniques, the high false-positive rate persists as a challenge. The study introduces a novel approach utilizing Natural Language Processing (NLP) and machine learning models, particularly BERT (Bidirectional Encoder Representations from Transformers), for SMS spam detection and classification. Data preprocessing techniques, such as stop word removal and tokenization, are applied, along with feature extraction using BERT.

Machine learning models, including SVM, Logistic Regression, Naive Bayes, Gradient Boosting, and Random Forest, are integrated with BERT for differentiating spam from ham messages. Evaluation results revealed that the Naïve Bayes classifier + BERT model achieves the highest accuracy at 97.31% with the fastest execution time of 0.3 seconds on the test dataset. This approach demonstrates a notable enhancement in spam detection efficiency and a low false-positive rate.

The developed model presents a valuable solution to combat SMS spam, ensuring faster and more accurate detection. This model not only safeguards users' privacy but also assists network providers in effectively identifying and blocking SMS spam messages.

Item Type:	Article
Subjects:	Impact Archive > Computer Science
Depositing User:	Managing Editor
Date Deposited:	01 Nov 2023 08:33
Last Modified:	01 Nov 2023 08:33
URI:	http://research.sdpublishers.net/id/eprint/3311

Actions (login required)

: View Item