Evaluation of Preprocessing Pipelines for Improving Student Risk Clustering in Educational Data
Abstract
This study aims to build a clustering-based model to predict student risk while focusing on how preprocessing affects the clustering performance. The dataset for the study has around 1 million records. The study uses preprocessing techniques such as data cleaning, normalization, encoding, and dimensionality reduction to prepare the data, and Dask and Apache Spark are used for scalable processing. For the clustering phase, the study employs MiniBatch K-Means and DBSCAN. On the other hand, the study evaluates the final results with Silhouette Score and Davies-Bouldin Index for clustering quality, and the runtime and memory usage to evaluate computational performance. The study also compares multiple preprocessing configurations to find the best trade-off between efficiency and interpretability. The study aims to develop a type-aware, scalable preprocessing framework. The outcome should improve how universities identify at-risk students early, under an efficient and practical framework for real-world use and future work in educational data mining.
Keywords
Full Text:
PDFReferences
Hendrastuty, N. (2024). Penerapan data mining menggunakan algoritma K-Means clustering dalam evaluasi hasil pembelajaran siswa. Jurnal Ilmiah Informatika dan Ilmu Komputer (JIMA-ILKOM), 3(1): Article 1. https://doi.org/10.58602/jima-ilkom.v3i1.26
Julakanti, S.R. (2024). Cloud-powered data mining: Unlocking hidden patterns in cloud storage. International Journal of Intelligent Systems and Applications in Engineering, 12(23s): Article 23s.
Choque-Soto, V.M., Sosa-Jauregui, V.D., Ibarra, W. (2025). Characterization of the dropout student profile using data mining techniques. Revista de Gestão Social e Ambiental, 19(2): e011306. https://doi.org/10.24857/rgsa.v19n2-067
Griep, K. et al. (2025). Ensuring ethical, transparent, and auditable use of education data and algorithms on AutoML. Proceedings of the ACM. https://doi.org/10.1145/3639479.3639492
Lubis, A.H., Muliono, R., Khairina, N., Novita, N. (2024). The impact of K-Means on association rules mining algorithms performance. JCOSITTE. https://doi.org/10.30596/jcositte.v5i2.20907
Erdiansyah, D., Abdullah, I.N., Tallo, A.J. (2025). Determination of potential business locations using data mining clustering. Jurnal Pilar Nusa Mandiri, 21(1): Article 1. https://doi.org/10.33480/pilar.v21i1.6295
Antonio, R., Leong, H. (2023). Performance of synthetic minority over-sampling technique on support vector machine and K-nearest neighbor for sentiment analysis of metaverse in Indonesia. Proxies: Jurnal Informatika, 6(2): 160–170. https://doi.org/10.24167/proxies.v6i2.12459
Lim, A.N.P., Leong, H. (2020). Hate speech prediction using K-Means algorithm. Proxies: Jurnal Informatika, 3(2): 98–104. https://doi.org/10.24167/proxies.v3i2.12430
Qu, F. (2024). Research on online learning behavior of higher vocational students based on data mining. In: Proceedings of the 4th International Conference on Artificial Intelligence and Education (ICAIE 2023), Atlantis Highlights in Computer Sciences, Vol. 15, pp. 30–37. https://doi.org/10.2991/978-94-6463-242-2_5
Al-Badani, A.M., Shujaaddeen, A.A., Aljafare, M.M. (2025). Efficient mining of FP-growth algorithm structure and Apriori algorithm using OFIM for big data. International Journal of Applied Information Systems, 12(47): 8–15.
Mkungudza, J., Twabi, H.S., Manda, S.O.M. (2024). Development of a diagnostic predictive model for determining child stunting in Malawi: A comparative analysis of variable selection approaches. BMC Medical Research Methodology, 24(1): 175. https://doi.org/10.1186/s12874-024-02283-6
Stefanus, K., Leong, H. (2023). Comparison of random forest algorithm accuracy with XGBoost using hyperparameters. Proxies: Jurnal Informatika, 7(1): 15–23. https://doi.org/10.24167/proxies.v7i1.12464
DOI: https://doi.org/10.24167/proxies.v9i2.14948
Copyright (c) 2026 Proxies : Jurnal Informatika
View My Stats




