Ensemble Anomaly Detection and SHAP-Based Attribution for Mapping Educational Inequality across Indonesian Provinces

Emeylia Safitri, I Gusti Ngurah Sentana Putra, Yeni Rahkmawati, Ika Nur Laily Fitriana, Nuramaliyah Nuramaliyah, Ria Faulina

Abstract


Educational inequality across Indonesia’s 38 provinces remains a challenge to equitable development. This study identifies anomalous provincial educational profiles by comparing Isolation Forest, Local Outlier Factor, and One-Class SVM across 50 parameter configurations in five analytical categories. The framework integrates ensemble majority voting, SHAP-based attribution, and leave-one-out robustness testing. Under the Silhouette-based selection criterion, Isolation Forest with contamination 0.05 ranked highest in four categories, with Silhouette Scores of 0.47-0.60 and Stability Scores of 1.0. However, sensitivity analysis showed that this criterion tends to favor configurations detecting fewer anomalies, so IF (0.05) should not be considered unambiguously superior. Highland Papua was consistently identified as anomalous across all categories. SHAP highlighted socioeconomic indicators, particularly rural child-labor participation, as important contributors to anomaly scores. Given the single-year design, 38 observations, and high feature-to-sample ratio, findings should be interpreted as exploratory descriptive mapping rather than causal or broadly generalizable inference.

Keywords


Anomaly Detection; Educational Inequality; Isolation Forest; SHAP.

Full Text:

PDF

References


S. Robiati, A. Hakim, G. Dharmawan, and C. Khotimah, “Application of K-Medoids for Regional Classification Based on Quality, Access, and Governance of Education in Indonesia,” Proceedings of The International Conference on Data Science and Official Statistics, vol. 2025, no. 1, pp. 1002–1016, Dec. 2025, https://doi.org/10.34123/icdsos.v2025i1.682.

Md. W. P. Dananjaya, N. N. K. Krisnawijaya, G. H. Prathama, I. G. N. D. Paramartha, and A. W. O. Gama, “Analisis Determinan Karakter Siswa Menggunakan Explainable Machine Learning (SHAP) dan Klasterisasi Profil Sekolah Studi Kasus Rapor Pendidikan Provinsi Bali,” Jurnal Kridatama Sains dan Teknologi, vol. 7, no. 02, pp. 936–948, Dec. 2025, https://doi.org/10.53863/kst.v7i02.1988.

T. Terttiaavini, A. Heryati, and T. S. Saputra, “Optimizing Socioeconomic Features for Poverty Prediction in South Sumatera,” TIERS Information Technology Journal, vol. 6, no. 1, pp. 16–32, Jun. 2025, https://doi.org/10.38043/tiers.v6i1.6244.

L. A. Alfia, F. Nugroho, and Y. M. Arif, “Comparative Analysis of Artificial Neural Networks, Linear Regression, Random Forest, and Support Vector Machine for Predicting Poverty Levels in Indonesia,” International Journal of Advances in Data and Information Systems, vol. 6, no. 3, pp. 789–800, Dec. 2025, https://doi.org/10.59395/ijadis.v6i3.1467.

R. G. Martinez and M. Cooray, “Enhancing Poverty Targeting with Spatial Machine Learning: An application to Indonesia,” 2025.

Z. Sidek, S. S. S. Ahmad, and N. H. I. Teo, “Unsupervised outlier detection in high-dimensional text data: a comparative analysis,” Bulletin of Electrical Engineering and Informatics, vol. 14, no. 4, pp. 2997–3005, Aug. 2025, https://doi.org/10.11591/eei.v14i4.9573.

M. M. Breunig, H.-P. Kriegel, R. T. Ng, and J. Sander, “LOF,” in Proceedings of the 2000 ACM SIGMOD international conference on Management of data, New York, NY, USA: ACM, May 2000, pp. 93–104. https://doi.org/10.1145/342009.335388.

G. Airlangga, “Advanced Machine Learning Techniques for Seismic Anomaly Detection in Indonesia: A Comparative Study of Lof, Isolation Forest, and One-Class SVM,” Jurnal Lebesgue : Jurnal Ilmiah Pendidikan Matematika, Matematika dan Statistika, vol. 5, no. 1, pp. 49–61, Apr. 2024, https://doi.org/10.46306/lb.v5i1.490.

F. T. Liu, K. M. Ting, and Z.-H. Zhou, “Isolation-Based Anomaly Detection,” ACM Transactions on Knowledge Discovery from Data, vol. 6, no. 1, pp. 1–39, Mar. 2012, https://doi.org/10.1145/2133360.2133363.

B. Schölkopf, R. Williamson, A. Smola, J. Shawe-Taylor, and J. Piatt, “Support vector method for novelty detection,” Advances in Neural Information Processing Systems, pp. 582–588, 2000.

D. Samariya and A. Thakkar, “A Comprehensive Survey of Anomaly Detection Algorithms,” Annals of Data Science, vol. 10, no. 3, pp. 829–850, Jun. 2023, https://doi.org/10.1007/s40745-021-00362-9.

M. Mutmainah and W. Yustanti, “Studi Komparasi Local Outlier Factor (LOF) dan Isolation Forest (IF) pada Analisis Anomali Kinerja Dosen,” Journal of Informatics and Computer Science (JINACS), vol. 6, no. 02, pp. 532–540, Jul. 2024, https://doi.org/10.26740/jinacs.v6n02.p532-540.

G. P. Oliveira, J. Castro Gertrudes, and R. B. Oliveira, “Spending pattern visualization using unsupervised machine learning,” in Anais do XXXVIII Simpósio Brasileiro de Banco de Dados (SBBD 2023), Sociedade Brasileira de Computação - SBC, Sep. 2023, pp. 167–178. https://doi.org/10.5753/sbbd.2023.231577.

V. Papastefanopoulos, P. Linardatos, and S. Kotsiantis, “Unsupervised Outlier Detection: A Meta-Learning Algorithm Based on Feature Selection,” Electronics, vol. 10, no. 18, p. 2236, Sep. 2021, https://doi.org/10.3390/electronics10182236.

T. Kim and C. H. Park, “Anomaly Pattern Detection in Streaming Data Based on the Transformation to Multiple Binary-Valued Data Streams,” Journal of Artificial Intelligence and Soft Computing Research, vol. 12, no. 1, pp. 19–27, Jan. 2022, https://doi.org/10.2478/jaiscr-2022-0002.

B. Chugh, N. Malik, D. Gupta, and B. S. Alkahtani, “A probabilistic approach driven credit card anomaly detection with CBLOF and isolation forest models,” Alexandria Engineering Journal, vol. 114, pp. 231–242, Feb. 2025, https://doi.org/10.1016/j.aej.2024.11.054.

S. M. Lundberg and S. I. Lee, “A unified approach to interpreting model predictions,” Advances in Neural Information Processing Systems, vol. 2017-December, no. Section 2, pp. 4766–4775, 2017.

Y. Huang, Y. Zhou, J. Chen, and D. Wu, “Applying Machine Learning and SHAP Method to Identify Key Influences on Middle-School Students’ Mathematics Literacy Performance,” Journal of Intelligence, vol. 12, no. 10, p. 93, Sep. 2024, https://doi.org/10.3390/jintelligence12100093.

M. Delprato, “Identifying the post-pandemic determinants of low performing students in Latin America through interpretable Machine Learning methods,” Engineering Applications of Artificial Intelligence, vol. 179, p. 115154, Sep. 2026, https://doi.org/10.1016/j.engappai.2026.115154.

B. Zaman et al., “Modeling education impact: a machine learning-based approach for improving the quality of school education,” Journal of Computers in Education, vol. 11, no. 4, pp. 1181–1214, Dec. 2024, https://doi.org/10.1007/s40692-023-00297-5.

P. J. Rousseeuw, “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis,” Journal of Computational and Applied Mathematics, vol. 20, pp. 53–65, Nov. 1987, https://doi.org/10.1016/0377-0427(87)90125-7.

D. L. Davies and D. W. Bouldin, “A Cluster Separation Measure,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. PAMI-1, no. 2, pp. 224–227, Apr. 1979, https://doi.org/10.1109/TPAMI.1979.4766909.

T. Caliński and J. Harabasz, “A dendrite method for cluster analysis,” Communications in Statistics, vol. 3, no. 1, pp. 1–27, Jan. 1974, https://doi.org/10.1080/03610927408827101.

J. Han, J. A. Guzman, and M. L. Chu, “Prediction of gully erosion susceptibility through the lens of the SHapley Additive exPlanations (SHAP) method using a stacking ensemble model,” Journal of Environmental Management, vol. 383, p. 125478, May 2025, https://doi.org/10.1016/j.jenvman.2025.125478.

S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees,” Nature Machine Intelligence, vol. 2, no. 1, pp. 56–67, Jan. 2020, https://doi.org/10.1038/s42256-019-0138-9.

E. Settanni and J. S. Srai, “It’s a long way to the top (if you wanna biplot): a back-to-basics perspective on the implementation of principal component biplots in R,” Quality & Quantity, vol. 60, no. 1, pp. 1173–1213, Jul. 2025, https://doi.org/10.1007/s11135-025-02266-9.

H. Abdi and L. J. Williams, “Principal component analysis,” WIREs Computational Statistics, vol. 2, no. 4, pp. 433–459, Jul. 2010, https://doi.org/10.1002/wics.101.

L. van der Maaten and G. Hinton, “Visualizing Data using t-SNE,” Journal of Machine Learning Research, vol. 9, no. 86, pp. 2579–2605, 2008.

L. McInnes, J. Healy, N. Saul, and L. Großberger, “UMAP: Uniform Manifold Approximation and Projection,” Journal of Open Source Software, vol. 3, no. 29, p. 861, Sep. 2018, https://doi.org/10.21105/joss.00861.

K. V. Nayak and S. Alam, “The digital divide, gender and education: challenges for tribal youth in rural Jharkhand during Covid-19,” DECISION, vol. 49, no. 2, pp. 223–237, Jun. 2022, https://doi.org/10.1007/s40622-022-00315-y.

N. Singh, Dr. S. Singh, Dr. B. Goswami, Dr. S. Kumar, and B. Kumar, “Bridging The Digital Divide: A Comprehensive Analysis Of ICT Infrastructure In Rural Schools Of Jharkhand, India,” Educational Administration: Theory and Practice, Jun. 2024, https://doi.org/10.53555/kuey.v30i6.7279.




DOI: http://dx.doi.org/10.30829/zero.v10i2.31161

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

 
 
✉  Contact & Indexing
Get in touch with ZERO: Jurnal Sains, Matematika dan Terapan
Email
zero_journal@uinsu.ac.id
WhatsApp · Admin Official
085270009767