Article Open Access

Performance Evaluation of Indobert and Roberta for Fake News Detection in Indonesia using Transfer Learning

(1) * Ramzi Akbarsya Dihariyanto Mail (Department of Informatics, Universitas Pembangunan Jaya, South Tangerang, 15413, Indonesia)
(2) Lathifah Alfat Mail (Center for Urban Studies, Universitas Pembangunan Jaya, South Tangerang, 15413, Indonesia)
*Corresponding author

Abstract


The rapid spread of hoax news through digital media in Indonesia has reduced public trust in the credibility of online information. Therefore, an accurate automated system is required to detect false information efficiently. This study evaluates the performance of IndoBERT and RoBERTa models for Indonesian fake news classification using the Indonesian Fact and Hoax Political News dataset containing more than 30,000 articles. The preprocessing stage involved text cleaning and random undersampling to address class imbalance. Experimental results show that IndoBERT achieved excellent performance with 97% accuracy and consistent precision, recall, and F1-score values of 0.97. In contrast, RoBERTa produced lower performance with 51% accuracy due to majority class bias. These findings indicate that pretrained models trained on local Indonesian corpora are more effective than globally pretrained models for Indonesian fake news detection. The best-performing model was further deployed into a Flask-based web application to provide an interactive hoax detection tool for the public.

Keywords


BERT; Fake News; IndoBERT; Information Detection; Natural Language Processing (NLP); Transfer Learning

   

DOI

https://doi.org/10.33122/ejeset.v7i1.1424
      

Article metrics

Abstract views : 0 | PDF views : 0

   

Cite

   

Full Text

Download

References


Ahmed, H., Traore, I., & Saad, S. (2022). Detecting opinion spams and fake news using text classification. Journal of Information Security and Applications, 65, 103082. https://doi.org/10.1002/spy2.9

Aji, A. F., Winata, G. I., Koto, F., et al. (2022). One country, 700+ languages: NLP challenges for underrepresented languages and dialects in Indonesia. Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.500

Alnabhan, M. Q., & Branco, P. (2024). Fake news detection using deep learning: A systematic literature review. IEEE Access, 12, 114435–114459. https://doi.org/10.1109/ACCESS.2024.3435497

Buda, M., Maki, A., & Mazurowski, M. A. (2021). A systematic study of the class imbalance problem in convolutional neural networks. Neural Networks, 106, 249–259. https://doi.org/10.1016/j.neunet.2018.07.011

Chalkidis, I., Fergadiotis, M., Malakasiotis, P., Aletras, N., & Androutsopoulos, I. (2020). LEGAL-BERT: The Muppets straight out of law school. Findings of the Association for Computational Linguistics: EMNLP 2020, 2898–2904. https://doi.org/10.18653/v1/2020.findings-emnlp.261

Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21(1), 6. https://doi.org/10.1186/s12864-019-6413-7

Cistia Sukmawati, E., Suryaningrum, L., Angelica, D., & Ghaniaviyanto Ramadhan, N. (2024). Klasifikasi berita palsu menggunakan model Bidirectional Encoder Representations from Transformers (BERT). https://doi.org/10.37278/sisinfo.v6i2.934

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019. https://doi.org/10.18653/v1/N19-1423

Ellam, I. V., Okorie, K. M., & Okebanama, U. F. (2025). Fake news detection system using natural language processing: An optimized approach. European Journal of Applied Science, Engineering and Technology, 3(2). https://doi.org/10.59324/ejaset.2025.3(2).15

Ibrahim, A. A., Ali, H. U., Yakubu, I. Z., & Lawal, I. A. (2024). Fake news detection in Hausa language using transfer learning method. International Journal of Innovative Science and Research Technology, 9(10), 2259–2269. https://doi.org/10.38124/ijisrt/IJISRT24OCT1050

Jocelynne, C., Wijayakusuma, I. G. N. L., & Harini, L. P. I. (2024). Detection of political hoax news using fine-tuning IndoBERT. Journal of Applied Informatics and Computing, 9(2). https://doi.org/10.30871/jaic.v9i2.8989

Kaliyar, R. K., Goswami, A., Narang, P., & Sinha, S. (2021). FakeBERT: Fake news detection in social media with a BERT-based deep learning approach. Multimedia Tools and Applications, 80, 11765–11788. https://doi.org/10.1007/s11042-020-10183-2

Khalkia, I. (2023). Analisis berita hoax dalam bahasa Indonesia menggunakan metode multilingual BERT.

Khan, J. Y., Khondaker, M. T. I., Iqbal, A., & Afroz, S. (2021). A benchmark study of machine learning models for online fake news detection. Machine Learning with Applications, 4, 100032. https://doi.org/10.1016/j.mlwa.2021.100032

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. Proceedings of COLING 2020. https://doi.org/10.18653/v1/2020.coling-main.66

Maulud, D. H., & Abdulazeez, A. M. (2020). A review on linear regression comprehensive in machine learning. Journal of Applied Science and Technology Trends, 1(4), 140–147. https://doi.org/10.38094/jastt1457

Mujilahwati, S., Zamroni, M. R., & Sholihin, M. (2026). Hybrid deep learning approach for Indonesian hoax detection: A comparative evaluation with IndoBERT. International Journal of Advances in Applied Sciences, 15(1), 322–332. https://doi.org/10.11591/ijaas.v15.i1.pp322-332

Oshikawa, R., Qian, J., & Wang, W. Y. (2020). A survey on natural language processing for fake news detection. Proceedings of COLING 2020. https://doi.org/10.18653/v1/2020.coling-main.75

Praha, T., Widodo, W., & Nugraheni, M. (2024). Indonesian fake news classification using transfer learning in CNN and LSTM. International Journal on Informatics Visualization, 8(3). http://dx.doi.org/10.62527/joiv.8.3.2126

Prayoga, H., & Kurniawan, R. (2025). Implementasi algoritma C4.5 pada klasifikasi status gizi balita menggunakan framework Flask. Jurnal Decode. https://doi.org/10.51454/decode.v5i2.1194

Riadi, A. T., Indriani, F., Mazdadi, M. I., Faisal, M. R., & Herteno, R. (2025). Cross-temporal generalization of IndoBERT for Indonesian hoax news classification. Jurnal Teknik Informatika (JUTIF). https://doi.org/10.52436/1.jutif.2025.6.5.4757

Ridho, M. Y., & Yulianti, E. (2024). From text to truth: Leveraging IndoBERT and machine learning models for hoax detection in Indonesian news. Jurnal Ilmiah Teknik Elektro Komputer dan Informatika, 10(3). https://doi.org/10.26555/jiteki.v10i3.29450

Rizqullah, M. R., et al. (2023). Indonesian fact and hoax political news [Data set]. Kaggle. https://www.kaggle.com/datasets/linkgish/indonesian-fact-and-hoax-political-news

Roy, A., Basak, K., Ekbal, A., & Bhattacharyya, P. (2022). A deep ensemble framework for fake news detection using transformer-based models. Expert Systems with Applications, 198, 116758. https://doi.org/10.48550/arXiv.1811.04670

Sarker, I. H. (2021). Machine learning: Algorithms, real-world applications and research directions. SN Computer Science, 2(3), 160. https://doi.org/10.1007/s42979-021-00592-x

Sholikhah, I. I., Tri, A., Harjanta, J., & Latifah, K. (2023). Machine learning untuk deteksi berita hoax menggunakan BERT. Proceedings of INFEST 2023, 1–8. https://conference.upgris.ac.id/index.php/infest/article/view/3818

Shu, K., Wang, S., & Liu, H. (2020). Beyond news contents: The role of social context for fake news detection. Proceedings of WSDM. https://doi.org/10.48550/arXiv.1712.07709

Suherlan, E., Arti, S., Nabilah, S., & Hazimah, Z. (2025). Penerapan Flask framework untuk deployment model machine learning dalam mendukung analisis adaptasi mahasiswa pada pembelajaran daring. Jurnal PINTER. https://doi.org/10.21009/pinter.9.1.15

Wang, Y., Ma, F., Jin, Z., Yuan, Y., Xun, G., Jha, K., Su, L., & Gao, J. (2021). EANN: Event adversarial neural networks for multi-modal fake news detection. Proceedings of the ACM SIGKDD. https://doi.org/10.1145/3219819.3219903

Wibawa, I. G. B. S., Indrayana, I. N. E., & Ariawan, M. P. A. (2024). Penerapan metode IndoBERT untuk deteksi berita hoaks pada media digital berbahasa Indonesia. Jurnal Riset dan Aplikasi Mahasiswa Informatika (JRAMI). https://doi.org/10.30998/94zt3k10

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S. Y., & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. Proceedings of EMNLP 2020.


Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

 
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0