Detection of hoax content using support vector machine, Naive Bayes, and random forest algorithms
Abstract
The rapid dissemination of hoax information through digital media poses a major danger to public perception and societal order. Automated detection systems are crucial to combat this issue. This research providing a side-by-side evaluation of three machine learning (ML) models: support vector machine (SVM), Naive Bayes (NB), and random forest (RF) for classifying news text as hoax or non-hoax. A total of 4,009 articles were first cleaned, broken into tokens, stripped of common stopwords, and lemmatized. Next, the term frequency–inverse document frequency (TF-IDF) approach was employed to pull out the key features. Evaluation of the models relied on metrics such as accuracy, precision, recall, and F1-score. The top results came from recorded by SVM, which reached 97% accuracy and a balanced F1-score of 96%, followed by RF at 95% accuracy. Meanwhile, NB only achieved 87% accuracy and showed a stronger tendency toward false positives. These results lead to the conclusion that SVM is the optimal algorithm for hoax news detection, thanks to its superior and well-balanced classification capability. What sets this study apart is its use of both comparative performance evaluation and statistical testing on an English-language dataset, providing a robust methodological foundation for future development and adaptation of hoax detection systems in other linguistic contexts, including Indonesian.
Keywords
Hoax detection; Machine learning; Naive Bayes; Random forest; Support vector machine; Term frequency–inverse document frequency
Full Text:
PDFDOI: https://doi.org/10.11591/eei.v15i4.11351
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Bulletin of Electrical Engineering and Informatics (BEEI)
ISSN: 2089-3191
,
e-ISSN: 2302-9285
This journal is published by the
Institute of Advanced Engineering and Science (IAES)
.