Pandas皮尔逊相关性表列排序:将因变量n_hhld_trip设为第一列
解决方案
你修改pivot_table参数不生效的原因是:pandas的pivot_table接收字典类型的aggfunc参数时,输出的列会默认按列名字母顺序排序,不会遵循你在字典中写的变量顺序。你可以通过以下两种方式实现需求:
方法1:直接调整相关表的列顺序
计算完皮尔逊相关系数后,直接对相关表的列重排序即可,代码修改如下:
zone_sum_mean_combo = pd.pivot_table( read_excel, index=['Zone'], aggfunc={'Household ID': np.mean, 'dwtype': np.mean, 'n_hhld_trip': np.sum, 'expf': np.mean, 'n_emp_ft': np.sum, 'n_emp_home': np.sum, 'n_emp_pt': np.sum, 'n_lic': np.sum, 'n_pers': np.sum, 'n_student': np.sum, 'n_veh': np.sum} ) index_reset = zone_sum_mean_combo.reset_index() print(index_reset) pearson_correlation = index_reset.corr(method='pearson') # 新增:把n_hhld_trip调整为第一列 pearson_correlation = pearson_correlation[['n_hhld_trip'] + [col for col in pearson_correlation.columns if col != 'n_hhld_trip']] # 如果需要同时把n_hhld_trip调整为第一行,可以再加下面这行 pearson_correlation = pearson_correlation.reindex(index=['n_hhld_trip'] + [i for i in pearson_correlation.index if i != 'n_hhld_trip']) print(pearson_correlation)
方法2:提前调整参与相关计算的数据集列顺序
如果Zone是区域编码、不需要参与相关性计算,可以提前指定参与计算的列顺序,后续生成的相关表会自动遵循该顺序:
zone_sum_mean_combo = pd.pivot_table( read_excel, index=['Zone'], aggfunc={'Household ID': np.mean, 'dwtype': np.mean, 'n_hhld_trip': np.sum, 'expf': np.mean, 'n_emp_ft': np.sum, 'n_emp_home': np.sum, 'n_emp_pt': np.sum, 'n_lic': np.sum, 'n_pers': np.sum, 'n_student': np.sum, 'n_veh': np.sum} ) index_reset = zone_sum_mean_combo.reset_index() print(index_reset) # 新增:指定参与相关计算的列顺序,把因变量放在第一个 feature_order = ['n_hhld_trip', 'Household ID', 'dwtype', 'expf', 'n_emp_ft', 'n_emp_home', 'n_emp_pt', 'n_lic', 'n_pers', 'n_student', 'n_veh'] pearson_correlation = index_reset[feature_order].corr(method='pearson') print(pearson_correlation)
内容的提问来源于stack exchange,提问作者rb1705
相关产品推荐
相关产品推荐

