Evaluation of Preprocessing Pipelines for Improving Student Risk Clustering in Educational Data

Khen Oliver Katminto, Hironimus Leong

Abstract


This study aims to build a clustering-based model to predict student risk while focusing on how preprocessing affects the clustering performance. The dataset for the study has around 1 million records. The study uses preprocessing techniques such as data cleaning, normalization, encoding, and dimensionality reduction to prepare the data, and Dask and Apache Spark are used for scalable processing. For the clustering phase, the study employs MiniBatch K-Means and DBSCAN. On the other hand, the study evaluates the final results with Silhouette Score and Davies-Bouldin Index for clustering quality, and the runtime and memory usage to evaluate computational performance. The study also compares multiple preprocessing configurations to find the best trade-off between efficiency and interpretability. The study aims to develop a type-aware, scalable preprocessing framework. The outcome should improve how universities identify at-risk students early, under an efficient and practical framework for real-world use and future work in educational data mining.


Keywords


student risk prediction; preprocessing; clustering; educational data mining; scalability

Full Text:

PDF

References


Hendrastuty, N. (2024). Penerapan data mining menggunakan algoritma K-Means clustering dalam evaluasi hasil pembelajaran siswa. Jurnal Ilmiah Informatika dan Ilmu Komputer (JIMA-ILKOM), 3(1): Article 1. https://doi.org/10.58602/jima-ilkom.v3i1.26

Julakanti, S.R. (2024). Cloud-powered data mining: Unlocking hidden patterns in cloud storage. International Journal of Intelligent Systems and Applications in Engineering, 12(23s): Article 23s.

Choque-Soto, V.M., Sosa-Jauregui, V.D., Ibarra, W. (2025). Characterization of the dropout student profile using data mining techniques. Revista de Gestão Social e Ambiental, 19(2): e011306. https://doi.org/10.24857/rgsa.v19n2-067

Griep, K. et al. (2025). Ensuring ethical, transparent, and auditable use of education data and algorithms on AutoML. Proceedings of the ACM. https://doi.org/10.1145/3639479.3639492

Lubis, A.H., Muliono, R., Khairina, N., Novita, N. (2024). The impact of K-Means on association rules mining algorithms performance. JCOSITTE. https://doi.org/10.30596/jcositte.v5i2.20907

Erdiansyah, D., Abdullah, I.N., Tallo, A.J. (2025). Determination of potential business locations using data mining clustering. Jurnal Pilar Nusa Mandiri, 21(1): Article 1. https://doi.org/10.33480/pilar.v21i1.6295

Antonio, R., Leong, H. (2023). Performance of synthetic minority over-sampling technique on support vector machine and K-nearest neighbor for sentiment analysis of metaverse in Indonesia. Proxies: Jurnal Informatika, 6(2): 160–170. https://doi.org/10.24167/proxies.v6i2.12459

Lim, A.N.P., Leong, H. (2020). Hate speech prediction using K-Means algorithm. Proxies: Jurnal Informatika, 3(2): 98–104. https://doi.org/10.24167/proxies.v3i2.12430

Qu, F. (2024). Research on online learning behavior of higher vocational students based on data mining. In: Proceedings of the 4th International Conference on Artificial Intelligence and Education (ICAIE 2023), Atlantis Highlights in Computer Sciences, Vol. 15, pp. 30–37. https://doi.org/10.2991/978-94-6463-242-2_5

Al-Badani, A.M., Shujaaddeen, A.A., Aljafare, M.M. (2025). Efficient mining of FP-growth algorithm structure and Apriori algorithm using OFIM for big data. International Journal of Applied Information Systems, 12(47): 8–15.

Mkungudza, J., Twabi, H.S., Manda, S.O.M. (2024). Development of a diagnostic predictive model for determining child stunting in Malawi: A comparative analysis of variable selection approaches. BMC Medical Research Methodology, 24(1): 175. https://doi.org/10.1186/s12874-024-02283-6

Stefanus, K., Leong, H. (2023). Comparison of random forest algorithm accuracy with XGBoost using hyperparameters. Proxies: Jurnal Informatika, 7(1): 15–23. https://doi.org/10.24167/proxies.v7i1.12464




DOI: https://doi.org/10.24167/proxies.v9i2.14948

Copyright (c) 2026 Proxies : Jurnal Informatika



View My Stats