Comparison of Shallow and Deep Learning for Indonesian Clickbait Headline Classification

Authors

  • Muhammad Noer Attalah Dzahkwan Universitas Amikom Yogyakarta
  • Majid Rahardi Universitas Amikom Yogyakarta

DOI:

https://doi.org/10.52465/joiser.v4i3.85

Keywords:

Clickbait, Text classification, IndoBERT, Shallow learning, Deep learning

Abstract

Clickbait is an increasingly prevalent phenomenon in Indonesian online news media, where headlines are crafted to attract clicks without accurately reflecting article content. This study proposes and compares eight classification models: three shallow learning algorithms — Naive Bayes, Support Vector Machine (SVM), and Logistic Regression — and five transformer-based deep learning models: IndoBERT-p1, IndoBERT-p2, XLM-RoBERTa, mBERT, and DistilBERT. The dataset used is CLICK-ID, consisting of 15,000 labeled headlines from 12 Indonesian news portals, expanded to 25,138 samples via semi-supervised pseudo labeling with a confidence threshold of 0.85. All deep learning models were trained with Focal Loss (α=0.25, γ=2.0) to address class imbalance and Automatic Mixed Precision (AMP) for GPU efficiency. Results show that IndoBERT-p1, IndoBERT-p2, XLM-RoBERTa, mBERT, and DistilBERT achieve comparable performance, with macro F1-scores ranging from 86.57% to 88.54%. Among shallow learning models, SVM performs best with 83.51% F1-score. An average ensemble of all five transformer models achieves the best overall performance at 90.14% accuracy and 89.00% F1-score, outperforming every individual model. This study contributes to Indonesian clickbait detection research by demonstrating that ensemble aggregation of diverse transformer architectures yields more reliable performance than reliance on any single model.

Downloads

Published

2026-07-29

Issue

Section

Articles