如何用Seaborn绘制基于目标值的特征组合相关性热力图?
使用Seaborn绘制特征组合与目标值关联的热力图
需求说明
我们需要的不是常规的特征间相关性热力图,而是识别特征组合(或单个特征状态)与target_value的关联关系,同时要能根据target_value的阈值筛选热力图轴上显示的特征,还要分析特定特征取值(比如feature_1=4)与target_value的相关性。
以下是原始数据和编码后的二值特征数据:
原始CSV数据
feature_1, feature_2, feature_3, feature_4, target_value 4, 8, 9, 8, 0.1 9, 7, 2, 0, 0.2 4, 4, 1, 4, 0.6 9, 7, 8, 4, 0.7 0, 9, 0, 7, 0.9
编码后的二值特征数据(基于阈值转0/1)
feature_1, feature_2, feature_3, feature_4, target_value 0, 1, 1, 1, 0.1 1, 1, 0, 0, 0.2 0, 0, 0, 0, 0.6 1, 1, 1, 0, 0.7 0, 1, 0, 1, 0.9
解决方案
1. 构建特征组合与target_value的关联热力图
要展示两个特征的组合对target_value的影响,我们可以计算特征交叉项与target_value的相关性,或者直接统计特征组合对应的target_value均值,再用Seaborn绘制热力图。
基于二值特征的实现
针对编码后的0/1特征,生成两两特征的交叉项(代表“两个特征同时为1”的组合),计算交叉项与target_value的相关性:
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt # 加载编码后的数据 encoded_df = pd.DataFrame({ 'feature_1': [0,1,0,1,0], 'feature_2': [1,1,0,1,1], 'feature_3': [1,0,0,1,0], 'feature_4': [1,0,0,0,1], 'target_value': [0.1,0.2,0.6,0.7,0.9] }) # 定义特征列表 features = ['feature_1', 'feature_2', 'feature_3', 'feature_4'] # 初始化相关性矩阵 corr_matrix = pd.DataFrame(index=features, columns=features, dtype=float) # 填充矩阵:计算每对特征交叉项与target的相关性 for f1 in features: for f2 in features: cross_term = encoded_df[f1] * encoded_df[f2] corr_matrix.loc[f1, f2] = cross_term.corr(encoded_df['target_value']) # 绘制热力图 plt.figure(figsize=(8,6)) sns.heatmap(corr_matrix, annot=True, cmap='coolwarm', vmin=-1, vmax=1) plt.title('特征组合与target_value的相关性热力图') plt.show()
分析特定特征取值与target_value的相关性
以feature_1=4为例,筛选对应样本后,计算该子集内特征与target_value的相关性:
# 加载原始数据 raw_df = pd.DataFrame({ 'feature_1': [4,9,4,9,0], 'feature_2': [8,7,4,7,9], 'feature_3': [9,2,1,8,0], 'feature_4': [8,0,4,4,7], 'target_value': [0.1,0.2,0.6,0.7,0.9] }) # 筛选feature_1=4的样本 f1_4_subset = raw_df[raw_df['feature_1'] == 4] # 计算该子集内各特征与target的相关性 corr_result = f1_4_subset[features + ['target_value']].corr()['target_value'].drop('target_value') print("feature_1=4时,各特征与target_value的相关性:\n", corr_result) # 若要可视化,可将结果转为矩阵形式绘制热力图
2. 根据target_value筛选热力图轴上的特征
要实现“按target阈值自定义X/Y轴特征”,只需先筛选样本,再手动指定热力图的index和columns为目标特征列表即可。
示例:target_value >=0.5时,X轴显示feature_1/2,Y轴显示feature_3/4
# 筛选target >=0.5的样本 high_target_subset = raw_df[raw_df['target_value'] >= 0.5] # 指定X/Y轴特征 x_features = ['feature_1', 'feature_2'] y_features = ['feature_3', 'feature_4'] # 构建关联矩阵(以特征交叉项与target的相关性为例) target_corr_matrix = pd.DataFrame(index=y_features, columns=x_features, dtype=float) for x in x_features: for y in y_features: cross_term = high_target_subset[x] * high_target_subset[y] target_corr_matrix.loc[y, x] = cross_term.corr(high_target_subset['target_value']) # 绘制热力图 plt.figure(figsize=(6,4)) sns.heatmap(target_corr_matrix, annot=True, cmap='Blues') plt.title('target_value >=0.5时,特征组合与target的相关性') plt.show() # target_value <0.5的场景只需替换筛选条件,重复上述步骤即可
补充提示
如果想更直观展示特征组合对target的影响,可以把相关性替换为特征组合对应的target均值(针对二值特征),代码示例:
# 替换相关性计算部分 target_mean = encoded_df[(encoded_df[f1]==1) & (encoded_df[f2]==1)]['target_value'].mean() # 若没有该组合的样本,可填充NaN或0 corr_matrix.loc[f1, f2] = target_mean if not pd.isna(target_mean) else 0
内容的提问来源于stack exchange,提问作者MonteCristo
相关产品推荐
相关产品推荐

