Evaluating Short-Term Regime Prediction in the Korean Equity Market: A Large-Scale Study of Decision-Rule Dependence
Abstract
This study presents a large-scale benchmark and a decision-rule-aware evaluation protocol for short-term regime classification in the Korean stock market. Reliable evaluation in this setting is difficult because reported performance is sensitive to sample construction, look-ahead bias, class imbalance, and the decision rule used to convert model scores into predicted classes. The benchmark comprises 8,732,898 stock-day observations of 4,225 KOSPI and KOSDAQ common stocks from January 2012 to February 2026, with six three-class regime targets defined by ±5% and ±10% thresholds on 1-, 5-, and 10-trading-day forward returns. The protocol combines survivorship-bias-mitigating sampling, chronological splits with a 20-trading-day purge gap, training-only preprocessing, leakage checks, and full class-distribution reporting. Given severe class imbalance, models are evaluated using macro F1 and up/down directional F1 rather than accuracy alone. We report reference results for two tree-based baselines, LightGBM and XGBoost, and six tabular deep learning models. Under the default argmax rule, the deep learning models are more recall-oriented on extreme regimes, whereas LightGBM is more conservative. However, after validation-based threshold tuning, LightGBM’s directional F1 improves substantially, and the relative ranking narrows or reverses on several targets. These results show that apparent model superiority depends materially on the decision rule. The contribution of this study is therefore a reproducible benchmark and evaluation protocol, not a claim of universal model superiority or tradable profitability.
References
- Arik, S. Ö., & Pfister, T. (2021). TabNet: Attentive interpretable tabular learning. Proceedings of the AAAI Conference on Artificial Intelligence, 35(8), 6679–6687.
- Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794.
- Gorishniy, Y., Rubachev, I., Khrulkov, V., & Babenko, A. (2021). Revisiting deep learning models for tabular data. Advances in Neural Information Processing Systems, 34, 18932–18943.
- Gu, S., Kelly, B., & Xiu, D. (2020). Empirical asset pricing via machine learning. The Review of Financial Studies, 33(5), 2223–2273.
- Hamilton, J. D. (1989). A new approach to the economic analysis of nonstationary time series and the business cycle. Econometrica, 57(2), 357–384.
- He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263–1284.
- Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146–3154.
- Kim, S. (2023). A comparison of stock price prediction for global automotive companies using machine learning models [in Korean]. Journal of the Korean Data Analysis Society, 25(1), 249–263.
- López de Prado, M. (2018). Advances in financial machine learning. Wiley.
- Park, J., Jung, M., Kim, H., & Kim, S. (2024). An analysis of investment performance for portfolio selection models reflecting news sentiment analysis: Focusing on the Korean stock market [in Korean]. Journal of the Korean Operations Research and Management Science Society, 49(4), 57–72.
- Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427–437.
Keywords
Details
| Section | Articles |
| Issue | Vol. 2 No. 2 (2026): Volume 2 Issue 2 (Jun 2026) |
| Published | 2026-06-18 |
| Pages | 1-13 |
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

