IndoHoaxDetector: IndoBERT, sanitised mixed corpus
Binary hoax classifier (0 = factual, 1 = hoax) fine-tuned from indobenchmark/indobert-base-p1.
It accompanies the ICDEES 2026 paper IndoHoaxDetector: A Comparative Evaluation of Classical
Machine Learning and IndoBERT for Indonesian Mixed Corpus Hoax Detection.
Code: https://github.com/theonegareth/IndoHoaxDetector
Training data
The training data is 62,934 Indonesian documents from news portals (CNN Indonesia, Detik.com,
Kompas.com, Tempo.co), the TurnBackHoax fact-checker and the Komdigi hoax registry. Fact-checker
verdict tags ([SALAH], [HOAKS], ...) and the fact-check template sections were removed
before training. Left in place, they let a rule with no training reach 99.64% weighted F1.
Earlier versions of this repository were trained on the unsanitised text and reported 99.89%;
those numbers reflected the verdict tags and are superseded.
Training
Input is the title and body, unstemmed, truncated to 128 tokens. Training ran for 3 epochs with batch size 16, AdamW at learning rate 1e-5, weight decay 0.01, 10% warm-up and fp16. The best epoch was chosen on a 10% validation split, with seeds 42, 1 and 2. Hardware was an NVIDIA RTX 3050 Ti Laptop GPU (4 GB), and one training run took 2,190 s.
Results (held-out 20%, stratified)
| Metric | Value |
|---|---|
| Accuracy (3 seeds) | 99.71 ± 0.05% |
| Macro-F1 | 99.71% |
| Balanced accuracy | 99.71% |
| ROC-AUC | 99.99% |
| Weighted F1, unseen publishers (train CNN/Kompas/TurnBackHoax, test Detik/Tempo/Komdigi; 3 seeds) | 86.03 ± 2.96% |
| ROC-AUC on an independent human-labelled benchmark (Pratiwi et al., 600 articles) | 0.50 (chance) |
Limitations
Labels come from each document's source rather than from per-document fact-checking, so the model largely learns publisher style. On unseen publishers weighted F1 falls by 13.7 points on average, to the level of a linear SVM: the model transfers well to unseen news outlets but recognises only 80.2% of hoaxes from an unseen hoax registry. On an independent benchmark labelled by human annotators it performs at chance (ROC-AUC 0.50), so do not use it as a general Indonesian hoax detector.
Short inputs lean towards “hoax”. In training, hoax documents average 43 tokens against 137 for news articles, so a single sentence, even a factual one, often scores high. Give the model a full article (title plus body) wherever possible. The model was not trained on social-media posts and should not be applied to them without in-domain data.
Model tree for theonegareth/IndoHoaxDetector
Base model
indobenchmark/indobert-base-p1