如何从Keras回归模型提取特征重要性?已遇两类问题待解
从Keras回归模型提取特征重要性/显著性图的解决方案
问题背景
用户构建了Keras回归模型后,尝试用eli5提取排列特征重要性时触发错误,自行编写的排列重要性函数输出所有特征的重要性数值完全相同,需要正确的特征重要性提取方法。
模型实现代码
import numpy as np import pandas as pd import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler import tensorflow as tf from tensorflow import keras from tensorflow.keras import layers import shap from sklearn.inspection import permutation_importance # Load data (modify as needed) s_data = np.array([ [1.0, 0.0, np.nan, 0.0], [2.0, 0.0, np.nan, 2.0], [0.0, 1.0, 2.0, 0.0] ]) # Load phenotype labels as a NumPy array (modify as needed) labels = np.array([ [7.], [9.], [2.] ]) s_data[np.isnan(s_data)] = -1 # Split data into training and testing sets X_train, X_test, y_train, y_test = train_test_split(s_data, labels, test_size=0.2, random_state=42) # Standardize the input features scaler = StandardScaler() X_train = scaler.fit_transform(X_train) X_test = scaler.transform(X_test) # Define a deep learning model model = keras.Sequential([ layers.Input(shape=(X_train.shape[1],)), layers.Dense(128, activation='relu'), layers.Dense(64, activation='relu'), layers.Dense(1, activation='linear') ]) # Compile the model model.compile(optimizer='adam', loss='mean_squared_error') # Train the model history = model.fit(X_train, y_train, epochs=50, batch_size=32, validation_split=0.2)
尝试eli5的错误代码及报错
import eli5 from eli5.sklearn import PermutationImportance perm = PermutationImportance(model, random_state=1).fit(X_train,y_train) eli5.show_weights(perm, feature_names = X_train.columns.tolist()) # Evaluate the model on the test set test_loss = model.evaluate(X_test, y_test) print(f'Test Loss: {test_loss}')
报错信息:
TypeError: If no scoring is specified, the estimator passed should have a 'score' method. The estimator <keras.engine.sequential.Sequential object at 0x000002B1ABA41448> does not.
自行编写的排列重要性函数及问题
n_permutations = 100 # Number of permutations to perform feature_importances = np.zeros(X_test.shape[1]) print (feature_importances.max()) for _ in range(n_permutations): shuffled_X_test = X_test.copy() np.random.shuffle(shuffled_X_test) # Permute the feature values shuffled_loss = model.evaluate(shuffled_X_test, y_test, verbose=0) importance = test_loss - shuffled_loss feature_importances += importance # Normalize feature importances feature_importances /= n_permutations # Print feature importances print("Feature Importances:") for i, importance in enumerate(feature_importances): print(f"SNP{i+1}: {importance:.4f}")
问题:输出的每个特征重要性数值完全相同。
问题根源分析
- eli5报错原因:Keras模型没有内置
score方法,eli5的PermutationImportance默认依赖该方法计算性能指标,需手动指定评分规则或包装模型适配sklearn接口。 - 自写函数结果相同的原因:
np.random.shuffle(shuffled_X_test)是对整个样本行打乱,而非单独打乱单个特征列,导致所有特征的随机化效果一致,最终重要性数值完全相同。
正确解决方案
方法1:修复排列特征重要性实现
核心逻辑是单独打乱每个特征列,计算该特征被打乱后的性能下降值:
# 先计算基准测试损失 test_loss = model.evaluate(X_test, y_test, verbose=0) n_permutations = 100 n_features = X_test.shape[1] feature_importances = np.zeros(n_features) for feature_idx in range(n_features): total_importance = 0.0 for _ in range(n_permutations): # 复制测试集,仅打乱当前特征列 shuffled_X = X_test.copy() np.random.shuffle(shuffled_X[:, feature_idx]) # 计算打乱后的损失 shuffled_loss = model.evaluate(shuffled_X, y_test, verbose=0) # 损失差值越大,特征对模型的影响越重要 importance = test_loss - shuffled_loss total_importance += importance # 取多次排列的平均值 feature_importances[feature_idx] = total_importance / n_permutations # 输出结果 print("Feature Importances:") for i, importance in enumerate(feature_importances): print(f"SNP{i+1}: {importance:.4f}")
方法2:使用sklearn的permutation_importance适配Keras模型
通过自定义评分函数让sklearn工具兼容Keras模型:
from sklearn.inspection import permutation_importance # 定义评分函数:返回负MSE(sklearn默认分数越高越好,MSE越小越好) def neg_mse_score(model, X, y): y_pred = model.predict(X, verbose=0) mse = np.mean((y - y_pred)**2) return -mse # 计算排列重要性 result = permutation_importance( model, X_test, y_test, scoring=neg_mse_score, n_repeats=100, random_state=42 ) # 输出结果 print("Feature Importances:") for i, importance in enumerate(result.importances_mean): print(f"SNP{i+1}: {importance:.4f}")
方法3:使用SHAP值解释特征重要性
SHAP是深度学习模型解释的主流工具,能更精准反映特征对预测结果的影响:
import shap # 创建解释器,用训练集子集作为背景数据(减少计算量) explainer = shap.KernelExplainer(model.predict, X_train[:10]) shap_values = explainer.shap_values(X_test) # 计算每个特征的平均绝对SHAP值作为重要性 feature_importances_shap = np.mean(np.abs(shap_values), axis=0) # 输出结果 print("SHAP Feature Importances:") for i, importance in enumerate(feature_importances_shap): print(f"SNP{i+1}: {importance:.4f}") # 可视化SHAP摘要图 shap.summary_plot(shap_values, X_test, feature_names=[f"SNP{i+1}" for i in range(n_features)])
方法4:生成显著性图(Saliency Map)
通过计算模型输出对输入特征的梯度,直观展示单个样本中各特征的影响程度:
import tensorflow as tf def compute_saliency_map(model, input_sample): input_tensor = tf.convert_to_tensor(input_sample.reshape(1, -1), dtype=tf.float32) with tf.GradientTape() as tape: tape.watch(input_tensor) output = model(input_tensor) # 计算输出对输入的梯度,取绝对值作为显著性值 gradients = tape.gradient(output, input_tensor) saliency = tf.abs(gradients).numpy().flatten() return saliency # 对测试集中的第一个样本计算显著性图 sample_idx = 0 saliency = compute_saliency_map(model, X_test[sample_idx]) # 可视化 plt.bar([f"SNP{i+1}" for i in range(n_features)], saliency) plt.title(f"Saliency Map for Sample {sample_idx+1}") plt.xlabel("Features") plt.ylabel("Saliency Value") plt.show()
内容的提问来源于stack exchange,提问作者honeymoon
相关产品推荐
相关产品推荐

