EVALUASI EMPIRIS KINERJA DAN STABILITAS MODEL PADA KLASIFIKASI DATA TABULAR
DOI:
https://doi.org/10.47111/jti.v20i2.27818Keywords:
Benchmark Evaluation, Gradient Boosting, In-Context Learning, Tabular Classification, Tabular Data, Tabular Foundation ModelsAbstract
Data tabular banyak digunakan dalam berbagai aplikasi machine learning, dengan gradient boosting sebagai salah satu pendekatan yang umum digunakan untuk task klasifikasi. Perkembangan Tabular Foundation Models (FM) menawarkan alternatif melalui prapelatihan dan in-context learning. Namun, perbandingan FM dan gradient boosting dalam kondisi praktis tanpa hyperparameter tuning masih terbatas. Penelitian ini mengevaluasi tiga Tabular FM (TabPFN v2, TabICL, dan TabDPT) dan tiga metode gradient boosting (XGBoost, LightGBM, dan CatBoost) pada delapan dataset klasifikasi publik dengan karakteristik yang beragam, mencakup perbedaan ukuran data, jumlah fitur, jumlah kelas, dan distribusi kelas. Evaluasi dilakukan menggunakan konfigurasi bawaan melalui stratified five-fold cross-validation dan sepuluh random seed, dengan mempertimbangkan kinerja, ukuran dataset, ketidakseimbangan kelas, dan stabilitas. Hasil menunjukkan perbedaan kinerja yang signifikan (p < 0,001), dengan FM memberikan hasil terbaik pada tujuh dari delapan dataset. Keunggulan FM paling terlihat pada dataset kecil, tetapi keunggulan ini mengecil pada data yang tidak seimbang. CatBoost tetap kompetitif dan tidak berbeda signifikan dari ketiga FM, sedangkan stabilitas keenam model relatif serupa. Hasil ini menunjukkan bahwa FM merupakan pilihan yang kuat untuk dataset kecil tanpa tuning, namun demikian keunggulannya terhadap gradient boosting bergantung pada karakteristik data.
Downloads
References
[1] V. Borisov, T. Leemann, K. Seßler, and G. Kasneci, “Deep Neural Networks and Tabular Data: A Survey,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 6, pp. 7499–7519, 2024, doi: 10.1109/TNNLS.2022.3229161.
[2] N. Hollmann, S. Müller, K. Eggensperger, and F. Hutter, “TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second,” presented at the Proceedings of the 11th International Conference on Learning Representations, OpenReview.net, 2023. [Online]. Available: https://openreview.net/forum?id=cp5PvcI6w8_
[3] R. Shwartz-Ziv and A. Armon, “Tabular Data: Deep Learning Is Not All You Need,” Information Fusion, vol. 81, pp. 84–90, 2022, doi: 10.1016/j.inffus.2021.11.011.
[4] S. A. Fayaz, M. Zaman, S. Kaul, and M. A. Butt, “Is Deep Learning on Tabular Data Enough? An Assessment,” International Journal of Advanced Computer Science and Applications, vol. 13, no. 6, 2022, doi: 10.14569/IJACSA.2022.0130650.
[5] Y. Gorishniy, I. Rubachev, V. Khrulkov, and A. Babenko, “Revisiting Deep Learning Models for Tabular Data,” presented at the Advances in Neural Information Processing Systems, Curran Associates, 2021, pp. 18932–18943. [Online]. Available: https://proceedings.neurips.cc/paper/2021/hash/9d86d83f925f2149e9edb0ac3b49229c-Abstract.html
[6] A. Kadra, M. Lindauer, F. Hutter, and J. Grabocka, “Well-Tuned Simple Nets Excel on Tabular Datasets,” presented at the Advances in Neural Information Processing Systems, Curran Associates, 2021. [Online]. Available: https://proceedings.neurips.cc/paper/2021/hash/c902b497eb972281fb5b4e206db38ee6-Abstract.html
[7] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” presented at the Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA: ACM, 2016, pp. 785–794. doi: 10.1145/2939672.2939785.
[8] G. Ke et al., “LightGBM: A Highly Efficient Gradient Boosting Decision Tree,” presented at the Advances in Neural Information Processing Systems, Curran Associates, 2017. [Online]. Available: https://proceedings.neurips.cc/paper/2017/hash/6449f44a102fde848669bdd9eb6b76fa-Abstract.html
[9] L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, “CatBoost: Unbiased Boosting with Categorical Features,” presented at the Advances in Neural Information Processing Systems, Curran Associates, 2018.
[10] A. Y. Yildiz and A. Kalayci, “Gradient Boosting Decision Trees on Medical Diagnosis over Tabular Data,” presented at the 2025 IEEE Conference on AI and Data Analytics, IEEE, 2025.
[11] D. McElfresh, S. Khandagale, J. Valverde, and C. White, “When Do Neural Nets Outperform Boosted Trees on Tabular Data?,” presented at the Advances in Neural Information Processing Systems, Curran Associates, 2023.
[12] L. Grinsztajn, E. Oyallon, and G. Varoquaux, “Why Do Tree-Based Models Still Outperform Deep Learning on Tabular Data?,” presented at the Advances in Neural Information Processing Systems, Curran Associates, 2022, pp. 507–520. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/hash/0378c7692da36807bdec87ab043cdadc-Abstract-Datasets_and_Benchmarks.html
[13] G. Badaro, M. Saeed, and P. Papotti, “Transformers for Tabular Data Representation: A Survey of Models and Applications,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 227–249, 2023, doi: 10.1162/tacl_a_00544.
[14] S. Ö. Arik and T. Pfister, “TabNet: Attentive Interpretable Tabular Learning,” presented at the Proceedings of the AAAI Conference on Artificial Intelligence, AAAI Press, 2021, pp. 6679–6687. doi: 10.1609/aaai.v35i8.16826.
[15] N. Hollmann, S. Müller, L. Purucker, and F. Hutter, “Accurate Predictions on Small Data with a Tabular Foundation Model,” Nature, vol. 637, pp. 319–326, 2025, doi: 10.1038/s41586-024-08328-6.
[16] J. Qu, D. Holzmüller, G. Varoquaux, and M. Le Morvan, “TabICL: A Tabular Foundation Model for In-Context Learning on Large Data,” in Proceedings of the 42nd International Conference on Machine Learning, PMLR, 2025, pp. 50817–50847. [Online]. Available: https://proceedings.mlr.press/v267/qu25d.html
[17] J. Ma, “TabDPT: Scaling Tabular Foundation Models on Real Data.”
[18] J. Luo, Y. Quan, and S. Xu, “Robust-GBDT: Leveraging Robust Loss for Noisy and Imbalanced Classification with GBDT,” Knowledge and Information Systems, vol. 67, pp. 2481–2510, 2025, doi: 10.1007/s10115-024-02290-3.
[19] H.-J. Ye, S.-Y. Liu, H.-R. Cai, Q.-L. Zhou, and D.-C. Zhan, “A Closer Look at Deep Learning Methods on Tabular Datasets.”
[20] A. Shmuel, O. Glickman, and T. Lazebnik, “A Comprehensive Benchmark of Machine and Deep Learning Across Diverse Tabular Datasets,” arXiv preprint arXiv:2408.14817, 2024, doi: 10.48550/arXiv.2408.14817.
[21] B. Amirshahi and S. Lahmiri, “Bankruptcy Prediction Using Optimal Ensemble Models Under Balanced and Imbalanced Data,” Expert Systems, vol. 41, no. 5, p. e13444, 2024, doi: 10.1111/exsy.13444.
[22] Sarwo and Y. D. Prabowo, “Enhancing Classification Performance Through the Synergistic Use of XGBoost, TabPFN, and LGBM Models,” presented at the Proceedings of the 2023 15th International Congress on Advanced Applied Informatics Winter, IEEE, 2023. doi: 10.1109/IIAI-AAI-Winter60127.2023.
[23] R. Detrano, “International Application of a New Probability Algorithm for the Diagnosis of Coronary Artery Disease,” American Journal of Cardiology, vol. 64, no. 5, pp. 304–310, 1989, doi: 10.1016/0002-9149(89)90524-9.
[24] J. W. Smith, J. E. Everhart, W. C. Dickson, W. C. Knowler, and R. S. Johannes, “Using the ADAP Learning Algorithm to Forecast the Onset of Diabetes Mellitus,” presented at the Proceedings of the Annual Symposium on Computer Application in Medical Care, American Medical Informatics Association, 1988, pp. 261–265.
[25] P. Soundarapandian, L. Rubini, and P. Eswaran, “Chronic Kidney Disease Dataset,” UCI Machine Learning Repository. University of California, Irvine, 2015. [Online]. Available: https://archive.ics.uci.edu/ml/datasets/chronic_kidney_disease
[26] W. N. Street, W. H. Wolberg, and O. L. Mangasarian, “Nuclear Feature Extraction for Breast Tumor Diagnosis,” presented at the Proceedings of SPIE, SPIE, 1993, pp. 861–870. doi: 10.1117/12.148698.
[27] P. Cortez, A. Cerdeira, F. Almeida, T. Matos, and J. Reis, “Modeling Wine Preferences by Data Mining from Physicochemical Properties,” Decision Support Systems, vol. 47, no. 4, pp. 547–553, 2009, doi: 10.1016/j.dss.2009.05.016.
[28] M. Bohanec, “Car Evaluation Dataset,” UCI Machine Learning Repository. University of California, Irvine, 1997. [Online]. Available: https://archive.ics.uci.edu/ml/datasets/car+evaluation
[29] R. Kohavi, “Scaling Up the Accuracy of Naive-Bayes Classifiers: A Decision-Tree Hybrid,” presented at the Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining, AAAI Press, 1996, pp. 202–207.
[30] S. Moro, P. Cortez, and P. Rita, “A Data-Driven Approach to Predict the Success of Bank Telemarketing,” Decision Support Systems, vol. 62, pp. 22–31, 2014, doi: 10.1016/j.dss.2014.03.001.
[31] D.-H. Won, K.-S. Shin, and S. Youm, “Synthetic Data Augmentation for Imbalanced Tabular Data: A Comparative Study of Generation Methods,” Electronics, vol. 15, no. 4, p. 883, 2026, doi: 10.3390/electronics15040883.
[32] R. Kanász, P. Drotar, P. Gnip, and M. Zoricák, “Clash of Titans on Imbalanced Data: TabNet vs XGBoost,” presented at the Proceedings of the 2024 IEEE Conference on Artificial Intelligence, IEEE, 2024. doi: 10.1109/CAI59869.2024.
[33] R. A. Babu, V. Vishwa Priya, M. K. Mishra, and B. Swarna, “Transformer-Based Tabular Foundation Models: Outperforming Traditional Methods with TabPFN,” International Journal of Engineering, Science and Information Technology, vol. 5, no. 1, pp. 112–119, 2025.
[34] A. P. De Rosa, A. D’Ambrosio, C. Marzi, and F. Esposito, “XGBoost vs. TabPFN in Neuroimaging Machine Learning-Based Analysis,” presented at the Convegno Nazionale di Bioingegneria, 2023.
[35] Y. Zeng, T. Dinh, W. Kang, and A. C. Muller, “TABFLEX: Scaling Tabular Learning to Millions with Linear Attention,” presented at the Proceedings of Machine Learning Research, PMLR, 2025.





