IndoHoaxDetector: IndoBERT, sanitised mixed corpus

Binary hoax classifier (0 = factual, 1 = hoax) fine-tuned from indobenchmark/indobert-base-p1. It accompanies the ICDEES 2026 paper IndoHoaxDetector: A Comparative Evaluation of Classical Machine Learning and IndoBERT for Indonesian Mixed Corpus Hoax Detection. Code: https://github.com/theonegareth/IndoHoaxDetector

Training data

The training data is 62,934 Indonesian documents from news portals (CNN Indonesia, Detik.com, Kompas.com, Tempo.co), the TurnBackHoax fact-checker and the Komdigi hoax registry. Fact-checker verdict tags ([SALAH], [HOAKS], ...) and the fact-check template sections were removed before training. Left in place, they let a rule with no training reach 99.64% weighted F1. Earlier versions of this repository were trained on the unsanitised text and reported 99.89%; those numbers reflected the verdict tags and are superseded.

Training

Input is the title and body, unstemmed, truncated to 128 tokens. Training ran for 3 epochs with batch size 16, AdamW at learning rate 1e-5, weight decay 0.01, 10% warm-up and fp16. The best epoch was chosen on a 10% validation split, with seeds 42, 1 and 2. Hardware was an NVIDIA RTX 3050 Ti Laptop GPU (4 GB), and one training run took 2,190 s.

Results (held-out 20%, stratified)

Metric Value
Accuracy (3 seeds) 99.71 ± 0.05%
Macro-F1 99.71%
Balanced accuracy 99.71%
ROC-AUC 99.99%
Weighted F1, unseen publishers (train CNN/Kompas/TurnBackHoax, test Detik/Tempo/Komdigi; 3 seeds) 86.03 ± 2.96%
ROC-AUC on an independent human-labelled benchmark (Pratiwi et al., 600 articles) 0.50 (chance)

Limitations

Labels come from each document's source rather than from per-document fact-checking, so the model largely learns publisher style. On unseen publishers weighted F1 falls by 13.7 points on average, to the level of a linear SVM: the model transfers well to unseen news outlets but recognises only 80.2% of hoaxes from an unseen hoax registry. On an independent benchmark labelled by human annotators it performs at chance (ROC-AUC 0.50), so do not use it as a general Indonesian hoax detector.

Short inputs lean towards “hoax”. In training, hoax documents average 43 tokens against 137 for news articles, so a single sentence, even a factual one, often scores high. Give the model a full article (title plus body) wherever possible. The model was not trained on social-media posts and should not be applied to them without in-domain data.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for theonegareth/IndoHoaxDetector

Finetuned
(156)
this model