Predicting Bond Defaults in China: A Comparative and Interpretable Machine Learning Approach Based on the IFind Database

Authors

  • Yangyi Guo School of Mathematics, College of Science and Engineering, University of Edinburgh, Mathematics and Business BSc (Hons), Edinburgh, Scotland, EH8 9YL, United Kingdom

Keywords:

Bond Default Prediction, Machine Learning, Xgboost, SHAP, KMV Model, Ifind Database, Chinese Bond Market

Abstract

Since 2014, the Chinese bond market has seen a large increase in default, whose shortcomings the conventional credit risk models reveal. The paper presents a systematic machine learning model to forecast bond defaults with a new dataset in the IFind database (2014-2024). We compare eight machine learning models, such as Logistic Regression, Random Forest, XGBoost, and Deep Neural Networks. The findings indicate that tree-based ensemble models, especially XGBoost (F1-score: 0.58, AUC: 0.92), are much better than the traditional linear models. Most importantly, we use SHAP (SHapley Additive exPlanations) to give model interpretability, which indicates that profitability (ROE), leverage (Debt-to-Asset Ratio), and liquidity (Quick Ratio) are the most significant predictors, whereas the predictive power of official credit ratings is surprisingly low. A suggested hybrid KMV-XGBoost model is similarly accurate but has theoretically based interpretability. This paper presents a testable, readable early warning mechanism of the Chinese bond market with direct risk management and regulatory policy implications.

Downloads

Published

2026-08-31

How to Cite

Guo, Y. (2026). Predicting Bond Defaults in China: A Comparative and Interpretable Machine Learning Approach Based on the IFind Database. CPS Digital Library - Series of Conferences, 57–65. Retrieved from https://seriesofconference.com/index.php/SCJ/article/view/401