如何显示交叉表中大于列均值的'Percentage of 1/Yes'行实际数值
问题描述
我有一个DataFrame df,想要显示Percentage of 1/Yes列中所有大于该列均值的行的实际数值。目前我的代码只输出布尔值(True/False),但我需要展示符合条件行的实际值,不显示不符合条件的行。
我的代码
trainData, validData = train_test_split(df, test_size=0.4, random_state=1) # RFM类别的响应率 # RFM:将R、F、M类别合并为一个类别 trainData['RFM'] = trainData['Mcode'].astype(str) + trainData['Rcode'].astype(str) + trainData['Fcode'].astype(str) rfm_crosstab = pd.crosstab(index = [trainData['RFM']], columns = trainData['Florence'], margins = True) rfm_crosstab['Percentage of 1/Yes'] = 100 * (rfm_crosstab[1] / rfm_crosstab['All']) # 显示百分比大于均值的行 rfm_crosstab['Percentage of 1/Yes'] > rfm_crosstab['Percentage of 1/Yes'].mean()
当前输出
RFM 111 False 121 True 131 False 141 False 211 False 212 True 221 True 222 True 231 False 232 False 241 False 242 False 311 False 312 False 313 True 321 True 322 True 323 True 331 False 332 False 333 False 341 True 342 False 343 False 411 True 412 False 413 False 421 False 422 True 423 True 431 False 432 False 433 False 441 False 442 False 443 False 511 True 512 False 513 True 521 True 522 False 523 True 531 False 532 False 533 True 541 False 542 False 543 False All False Name: Percentage of 1/Yes, dtype: bool
数据:df
Seq# ID# Gender M R F FirstPurch ChildBks YouthBks CookBks ... ItalCook ItalAtlas ItalArt Florence Related Purchase Mcode Rcode Fcode Yes_Florence No_Florence 0 1 25 1 297 14 2 22 0 1 1 ... 0 0 0 0 0 5 4 2 0 1 1 2 29 0 128 8 2 10 0 0 0 ... 0 0 0 0 0 4 3 2 0 1 2 3 46 1 138 22 7 56 2 1 2 ... 1 0 0 0 2 4 4 3 0 1 3 4 47 1 228 2 1 2 0 0 0 ... 0 0 0 0 0 5 1 1 0 1 4 5 51 1 257 10 1 10 0 0 0 ... 0 0 0 0 0 5 3 1 0 1
解决方案
你当前的代码仅生成了布尔索引,需要用该索引筛选原DataFrame的行,才能得到符合条件的实际数值。另外注意排除最后一行的All汇总行,避免它干扰均值计算和结果展示。
修改后的代码如下:
trainData, validData = train_test_split(df, test_size=0.4, random_state=1) # RFM类别的响应率 trainData['RFM'] = trainData['Mcode'].astype(str) + trainData['Rcode'].astype(str) + trainData['Fcode'].astype(str) rfm_crosstab = pd.crosstab(index = [trainData['RFM']], columns = trainData['Florence'], margins = True) rfm_crosstab['Percentage of 1/Yes'] = 100 * (rfm_crosstab[1] / rfm_crosstab['All']) # 计算均值时排除最后一行的All mean_value = rfm_crosstab['Percentage of 1/Yes'].iloc[:-1].mean() # 筛选符合条件的行,同时排除All行 filtered_rows = rfm_crosstab[rfm_crosstab['Percentage of 1/Yes'] > mean_value].drop('All') # 显示结果 print(filtered_rows['Percentage of 1/Yes'])
关键说明
- 排除汇总行计算均值:用
iloc[:-1]取除最后一行(All)外的所有行计算均值,避免汇总值拉低或抬高整体均值。 - 布尔索引筛选行:将布尔条件直接传入
rfm_crosstab[],即可筛选出符合条件的行。 - 移除汇总行:用
drop('All')去掉不需要的汇总行,确保结果只包含实际的RFM类别数据。
这样输出的就是符合条件的Percentage of 1/Yes列的实际数值,而非布尔值。
内容的提问来源于stack exchange,提问作者c200402
相关产品推荐
相关产品推荐

