Comparison of Shallow and Deep Learning for Indonesian Clickbait Headline Classification
DOI:
https://doi.org/10.52465/joiser.v4i3.85Keywords:
Clickbait, Text classification, IndoBERT, Shallow learning, Deep learningAbstract
Clickbait is an increasingly prevalent phenomenon in Indonesian online news media, where headlines are crafted to attract clicks without accurately reflecting article content. This study proposes and compares eight classification models: three shallow learning algorithms — Naive Bayes, Support Vector Machine (SVM), and Logistic Regression — and five transformer-based deep learning models: IndoBERT-p1, IndoBERT-p2, XLM-RoBERTa, mBERT, and DistilBERT. The dataset used is CLICK-ID, consisting of 15,000 labeled headlines from 12 Indonesian news portals, expanded to 25,138 samples via semi-supervised pseudo labeling with a confidence threshold of 0.85. All deep learning models were trained with Focal Loss (α=0.25, γ=2.0) to address class imbalance and Automatic Mixed Precision (AMP) for GPU efficiency. Results show that IndoBERT-p1, IndoBERT-p2, XLM-RoBERTa, mBERT, and DistilBERT achieve comparable performance, with macro F1-scores ranging from 86.57% to 88.54%. Among shallow learning models, SVM performs best with 83.51% F1-score. An average ensemble of all five transformer models achieves the best overall performance at 90.14% accuracy and 89.00% F1-score, outperforming every individual model. This study contributes to Indonesian clickbait detection research by demonstrating that ensemble aggregation of diverse transformer architectures yields more reliable performance than reliance on any single model.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information System Exploration and Research

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


