信用风险分析中,能否针对各客户公司识别影响其风险的关键指标?
解决方案:为单个企业获取违约风险关联的关键指标
需求背景
我正在开展信用风险分析工作,目标是预测各企业对某虚构公司的违约风险。已通过RandomForestClassifier模型得到全局特征重要性,但希望针对每个新客户企业,找出与其违约风险(Risk_1)直接关联的关键指标——比如客户公司X违约风险为70%,关联指标为City、Age、Number_Employe;客户公司Y违约风险为80%,关联指标为City、Service、Average_Salary。
当前分析流程:使用含20个编码指标、各千余条违约(Classification=1)/未违约(Classification=0)标注数据训练模型,对无标注新企业进行风险概率预测,已得到各新企业的Risk_0和Risk_1值,现需明确每个企业对应的、与Risk_1最相关的指标。
当前代码实现
# X base composed of encoded indicators features = df_all_aux.columns.tolist() X = df_all_aux[features[:-1]] # all features except "Classification" # y base composed of the target: 1 if debt, 0 if no debt y = df_all_aux['Classification'] #Define the model rf_classifier = RandomForestClassifier(n_estimators=100, random_state=42) #Train the model using the training data rf_classifier.fit(X, y) #Predictions using the asset data y_pred = rf_classifier.predict_proba(df_new_companies) #Incorporating the data into the dataset df_new_companies['Risk_0'] = y_pred[:, 0] # Probability of being class 0 df_new_companies['Risk_1'] = y_pred[:, 1] # Probability of being class 1
数据结构示例
df_all_aux(训练数据集)
| City | Age | Number_Employe | Service | Average_Salary | Classification | ... |
|---|---|---|---|---|---|---|
| 1 | 100 | 20000 | 3 | 2000 | 1 | ... |
| 2 | 85 | 15000 | 1 | 5200 | 1 | ... |
| 1 | 103 | 20100 | 1 | 5200 | 1 | ... |
| 4 | 100 | 19800 | 2 | 5000 | 0 | ... |
| 1 | 101 | 30000 | 2 | 3500 | 0 | ... |
| 3 | 92 | 18900 | 3 | 5100 | 0 | ... |
df_new_companies(待预测数据集)
结构与df_all_aux一致,额外包含企业ID列。
具体实现方法
全局特征重要性反映的是特征在整个数据集的平均影响,要获取单个样本的关联指标,需要用模型可解释性工具,以下是两种直接可行的方案:
1. 使用SHAP值(推荐)
SHAP可以量化每个特征对单个样本预测结果的贡献值,能精准定位影响该企业Risk_1的核心指标。
代码实现
import shap # 初始化SHAP解释器(适配树模型) explainer = shap.TreeExplainer(rf_classifier) # 提取待预测数据的特征列(去掉无关列) X_new = df_new_companies.drop(['Risk_0', 'Risk_1', '企业ID'], axis=1) # 计算SHAP值 shap_values = explainer.shap_values(X_new) # 提取对应违约类别(Risk_1,即类别1)的SHAP值 shap_values_class1 = shap_values[1] feature_names = X_new.columns top_n = 3 # 自定义要提取的关键指标数量 # 为每个企业生成Top N关键指标 for idx in range(len(df_new_companies)): # 按特征贡献的绝对值排序 feature_contrib = sorted(zip(feature_names, shap_values_class1[idx]), key=lambda x: abs(x[1]), reverse=True) # 提取Top N指标名称 top_features = [item[0] for item in feature_contrib[:top_n]] # 存入数据集 df_new_companies.loc[idx, 'Top_Risk_Indicators'] = ', '.join(top_features)
2. 使用TreeInterpreter(轻量方案)
TreeInterpreter可解析随机森林中每棵树的决策路径,计算单个特征对预测概率的直接贡献,适合快速验证场景。
代码实现
from treeinterpreter import treeinterpreter as ti # 提取待预测数据的特征列 X_new = df_new_companies.drop(['Risk_0', 'Risk_1', '企业ID'], axis=1) # 获取预测结果、基准值和特征贡献 prediction, bias, contributions = ti.predict(rf_classifier, X_new) # 提取对应Risk_1的贡献值 contributions_class1 = contributions[:, :, 1] feature_names = X_new.columns top_n = 3 # 为每个企业生成Top N关键指标 for idx in range(len(df_new_companies)): feature_contrib = sorted(zip(feature_names, contributions_class1[idx]), key=lambda x: abs(x[1]), reverse=True) top_features = [item[0] for item in feature_contrib[:top_n]] df_new_companies.loc[idx, 'Top_Risk_Indicators'] = ', '.join(top_features)
内容的提问来源于stack exchange,提问作者Heloisa Ramos
相关产品推荐
相关产品推荐

