Penerapan Random Forest, SMOTEN, dan Class Merging untuk Klasifikasi Status Pekerjaan Lulusan pada Data Tracer Study UIN Raden Mas Said Surakarta

Penulis

  • Febrianta Surya Nugraha Universitas Islam Negeri Raden Mas Said Surakarta
  • Layyin Mahfiana Universitas Islam Negeri Raden Mas Said Surakarta

DOI:

https://doi.org/10.55635/jic.v11i2.334

Kata Kunci:

class merging, ketidakseimbangan kelas, Random Forest, SMOTEN, tracer study

Abstrak

Tracer study merupakan instrumen penting untuk mengevaluasi kualitas lulusan dan relevansi kurikulum terhadap kebutuhan dunia kerja. Namun, pemanfaatan data tracer study masih terbatas, sementara karakteristik datanya umumnya menghadapi permasalahan ketidakseimbangan kelas (class imbalance) yang dapat menurunkan performa model klasifikasi. Penelitian ini bertujuan membangun model klasifikasi status pekerjaan lulusan menggunakan algoritma Random Forest dengan pendekatan SMOTEN serta membandingkan tiga strategi class merging pada data tracer study UIN Raden Mas Said Surakarta yang terdiri atas 2.258 responden dan 30 variabel. Tahapan penelitian mengikuti kerangka CRISP-DM yang meliputi persiapan data, pemodelan, dan evaluasi menggunakan Stratified 5-Fold Cross Validation. Hasil penelitian menunjukkan bahwa strategi aggressive class merging (2 kelas) memberikan performa terbaik dengan test accuracy sebesar 76,99%, balanced accuracy sebesar 69,54%, serta menurunkan overfitting gap dari 35,83% menjadi 15,06%. Analisis feature importance menunjukkan bahwa tingkat hubungan studi–pekerjaan (20,62%), IPK kelulusan (17,09%), waktu mulai mencari pekerjaan (11,95%), praktikum (9,63%), dan kerja lapangan (6,68%) merupakan faktor yang paling berpengaruh terhadap status pekerjaan lulusan. Temuan ini menunjukkan bahwa penyederhanaan kelas efektif meningkatkan performa klasifikasi pada data tracer study yang tidak seimbang, sekaligus menegaskan pentingnya kesesuaian kurikulum dan pengalaman praktis dalam meningkatkan kesiapan kerja lulusan.

Referensi

[1] H. Schomburg and U. Teichler, Higher Education and Graduate Employment in Europe: Results from Graduate Surveys from Twelve Countries. Dordrecht, The Netherlands: Springer, 2006.

[2] Kementerian Pendidikan, Kebudayaan, Riset, dan Teknologi Republik Indonesia, Indikator Kinerja Utama Perguruan Tinggi Negeri. Jakarta, Indonesia: Kemendikbudristek, 2021.

[3] H. Schomburg, Handbook for Graduate Tracer Studies. Turin, Italy: European Training Foundation, 2016.

[4] U. Teichler, "Does Higher Education Matter? Lessons from a Comparative Graduate Survey," European Journal of Education, vol. 42, no. 1, pp. 11–34, 2007.

[5] R. Chandra, S. Ruhama, and M. W. Sarjono, "Exploring Tracer Study Service in Career Center Web Site of Indonesia Higher Education," arXiv, arXiv:1304.5869, 2013.

[6] F. F. Abdulloh, M. Rahardi, A. Aminuddin, S. D. Anggita, and A. Y. A. Nugraha, "Observation of Imbalance Tracer Study Data for Graduates Employability Prediction in Indonesia," International Journal of Advanced Computer Science and Applications, vol. 13, no. 8, pp. 169–174, 2022, doi:10.14569/IJACSA.2022.0130820.

[7] A. Miranda and K. M. Lhaksamana, "Classification Analysis of Waiting Period for Telkom University Alumni to Get Jobs Using Decision Tree and Support Vector Machine," Building of Informatics, Technology and Science (BITS), vol. 4, no. 2, pp. 866–875, 2022, doi:10.47065/bits.v4i2.1963.

[8] M. F. Haikal and I. Palupi, "Predicting University Graduates Employability Using Support Vector Machine Classification," Building of Informatics, Technology and Science (BITS), vol. 6, no. 2, pp. 911–920, 2024, doi:10.47065/bits.v6i2.5655.

[9] S. R. Cholil, V. Vydia, Susanto, and Y. Cahyono, "Prediksi Lama Masa Tunggu Alumni Universitas Semarang dalam Mendapatkan Pekerjaan Menggunakan Algoritma K-Nearest Neighbor," Indonesian Journal on Computer and Information Technology, vol. 9, no. 2, pp. 246–255, 2024.

[10] D. A. Indah Cahya Dewi, N. K. Bagiastuti, and P. W. Sunu, "Development of a Tracer Study System Using Agile Approach for Meeting Higher Education Key Performance Indicators," in Proc. International Conference on Sustainable Applied Science (ICOSTAS-SAS), Atlantis Press, 2024.

[11] A. Salabi, M. A. M. Prasetyo, H. Halil, and M. Maulina, "Effectiveness of the Islamic Education Management Study Program Using Alumni Tracer Study Data," Jurnal Akuntabilitas Manajemen Pendidikan, vol. 12, no. 2, pp. 164–177, 2024.

[12] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic Minority Over-sampling Technique," Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002.

[13] H. He and E. A. Garcia, "Learning from Imbalanced Data," IEEE Transactions on Knowledge and Data Engineering, vol. 21, no. 9, pp. 1263–1284, 2009.

[14] H. Han, W. Y. Wang, and B. H. Mao, "Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning," in Advances in Intelligent Computing. Berlin, Germany: Springer, 2005, pp. 878–887.

[15] G. E. A. P. A. Batista, R. C. Prati, and M. C. Monard, "A Study of the Behavior of Several Methods for Balancing Machine Learning Training Data," ACM SIGKDD Explorations Newsletter, vol. 6, no. 1, pp. 20–29, 2004.

[16] L. Breiman, "Random Forests," Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.

[17] A. Liaw and M. Wiener, "Classification and Regression by randomForest," R News, vol. 2, no. 3, pp. 18–22, 2002.

[18] T. K. Ho, "Random Decision Forests," in Proc. 3rd International Conference on Document Analysis and Recognition, 1995, pp. 278–282.

[19] G. Biau and E. Scornet, "A Random Forest Guided Tour," TEST, vol. 25, no. 2, pp. 197–227, 2016.

[20] C. Strobl, A.-L. Boulesteix, A. Zeileis, and T. Hothorn, "Bias in Random Forest Variable Importance Measures: Illustrations, Sources and a Solution," BMC Bioinformatics, vol. 8, Art. no. 25, 2007.

[21] G. Chandrashekar and F. Sahin, "A Survey on Feature Selection Methods," Computers & Electrical Engineering, vol. 40, no. 1, pp. 16–28, 2014.

[22] P. Chapman et al., CRISP-DM 1.0: Step-by-Step Data Mining Guide. Chicago, IL, USA: SPSS Inc., 2000.

[23] C. Shearer, "The CRISP-DM Model: The New Blueprint for Data Mining," Journal of Data Warehousing, vol. 5, no. 4, pp. 13–22, 2000.

[24] R. Wirth and J. Hipp, "CRISP-DM: Towards a Standard Process Model for Data Mining," in Proc. Practical Applications of Knowledge Discovery and Data Mining, 2000, pp. 29–39.

[25] R. J. A. Little and D. B. Rubin, Statistical Analysis with Missing Data, 3rd ed. Hoboken, NJ: Wiley, 2019.

[26] S. Van Buuren, Flexible Imputation of Missing Data, 2nd ed. Boca Raton, FL: CRC Press, 2018.

[27] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning, 2nd ed. New York: Springer, 2009.

[28] M. Kuhn and K. Johnson, Applied Predictive Modeling. New York: Springer, 2013.

[29] J. R. Vergara and P. A. Estévez, "A Review of Feature Selection Methods Based on Mutual Information," Neural Computing and Applications, vol. 24, pp. 175–186, 2014.

[30] R. Kohavi, "A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection," in Proc. IJCAI, 1995, pp. 1137–1145.

[31] T. Fawcett, "An Introduction to ROC Analysis," Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, 2006.

[32] D. M. W. Powers, "Evaluation: From Precision, Recall and F-Measure to ROC, Informedness, Markedness and Correlation," Journal of Machine Learning Technologies, vol. 2, no. 1, pp. 37–63, 2011.

[33] G. C. Cawley and N. L. C. Talbot, "On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation," Journal of Machine Learning Research, vol. 11, pp. 2079–2107, 2010.

[34] C. Romero and S. Ventura, "Educational Data Mining and Learning Analytics: An Updated Survey," WIREs Data Mining and Knowledge Discovery, vol. 14, no. 2, 2024.

[35] Z. Ersozlu, S. Taheri, and I. Koch, "A Review of Machine Learning Methods Used for Educational Data," Education and Information Technologies, vol. 29, 2024.

[36] M. H. Alkhudhayr et al., "A Review of Educational Data Mining Trends," Procedia Computer Science, vol. 237, pp. 88–95, 2024.

[37] H. N. Mpia, L. W. Mburu, and S. N. Mwendia, "Applying Data Mining in Graduates' Employability: A Systematic Literature Review," International Journal of Engineering Pedagogy, vol. 13, no. 2, pp. 4–24, 2023.

[38] N. M. Almutairi et al., "Employability Prediction: A Survey of Current Approaches, Research Challenges and Applications," Journal of Ambient Intelligence and Humanized Computing, vol. 14, pp. 3479–3503, 2023.

[39] R. Scandurra et al., "Do Employability Programmes in Higher Education Improve Skills and Labour Market Outcomes?," Studies in Higher Education, vol. 49, no. 8, pp. 1381–1396, 2024.

[40] J. H. Guanin-Fajardo, J. Guaña-Moya, and J. Casillas, "Predicting Academic Success of College Students Using Machine Learning Techniques," Data, vol. 9, no. 4, Art. no. 60, 2024.

[41] A. H. Alqahtani et al., "Predictive Models for Educational Purposes: A Systematic Review," Big Data and Cognitive Computing, vol. 8, no. 12, Art. no. 187, 2024.

[42] D. Sartika et al., "The Impact of Internship Experience on the Employability of Vocational Students: A Bibliometric and Systematic Review," Cogent Education, vol. 11, 2024.

[43] M. S. N. Al-Din and H. A. Al Abdulqader, "Students' Academic Performance Prediction Using Educational Data Mining and Machine Learning: A Systematic Review," International Journal of Research and Innovation in Social Science, vol. 8, no. 8, 2024.

[44] V. A., S. B. E., and K. Sathish, "A Comprehensive Study of Employable and Technical Skill Development Programme in Higher Education Scenario: A Systematic Review," Educational Administration: Theory and Practice, vol. 30, no. 4, 2024.

Unduhan

Diterbitkan

2026-01-20

Terbitan

Bagian

Articles