You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何显示交叉表中大于列均值的'Percentage of 1/Yes'行实际数值

问题描述

我有一个DataFrame df,想要显示Percentage of 1/Yes列中所有大于该列均值的行的实际数值。目前我的代码只输出布尔值(True/False),但我需要展示符合条件行的实际值,不显示不符合条件的行。

我的代码

trainData, validData = train_test_split(df, test_size=0.4, random_state=1)

# RFM类别的响应率
# RFM:将R、F、M类别合并为一个类别
trainData['RFM'] = trainData['Mcode'].astype(str) + trainData['Rcode'].astype(str) + trainData['Fcode'].astype(str)

rfm_crosstab = pd.crosstab(index = [trainData['RFM']], columns = trainData['Florence'], margins = True)
rfm_crosstab['Percentage of 1/Yes'] = 100 * (rfm_crosstab[1] / rfm_crosstab['All'])

# 显示百分比大于均值的行
rfm_crosstab['Percentage of 1/Yes'] > rfm_crosstab['Percentage of 1/Yes'].mean()

当前输出

RFM
111    False
121     True
131    False
141    False
211    False
212     True
221     True
222     True
231    False
232    False
241    False
242    False
311    False
312    False
313     True
321     True
322     True
323     True
331    False
332    False
333    False
341     True
342    False
343    False
411     True
412    False
413    False
421    False
422     True
423     True
431    False
432    False
433    False
441    False
442    False
443    False
511     True
512    False
513     True
521     True
522    False
523     True
531    False
532    False
533     True
541    False
542    False
543    False
All    False
Name: Percentage of 1/Yes, dtype: bool

数据:df

Seq#    ID# Gender  M   R   F   FirstPurch  ChildBks    YouthBks    CookBks ... ItalCook    ItalAtlas   ItalArt Florence    Related Purchase    Mcode   Rcode   Fcode   Yes_Florence    No_Florence
0   1   25  1   297 14  2   22  0   1   1   ... 0   0   0   0   0   5   4   2   0   1
1   2   29  0   128 8   2   10  0   0   0   ... 0   0   0   0   0   4   3   2   0   1
2   3   46  1   138 22  7   56  2   1   2   ... 1   0   0   0   2   4   4   3   0   1
3   4   47  1   228 2   1   2   0   0   0   ... 0   0   0   0   0   5   1   1   0   1
4   5   51  1   257 10  1   10  0   0   0   ... 0   0   0   0   0   5   3   1   0   1
解决方案

你当前的代码仅生成了布尔索引,需要用该索引筛选原DataFrame的行,才能得到符合条件的实际数值。另外注意排除最后一行的All汇总行,避免它干扰均值计算和结果展示。

修改后的代码如下:

trainData, validData = train_test_split(df, test_size=0.4, random_state=1)

# RFM类别的响应率
trainData['RFM'] = trainData['Mcode'].astype(str) + trainData['Rcode'].astype(str) + trainData['Fcode'].astype(str)

rfm_crosstab = pd.crosstab(index = [trainData['RFM']], columns = trainData['Florence'], margins = True)
rfm_crosstab['Percentage of 1/Yes'] = 100 * (rfm_crosstab[1] / rfm_crosstab['All'])

# 计算均值时排除最后一行的All
mean_value = rfm_crosstab['Percentage of 1/Yes'].iloc[:-1].mean()

# 筛选符合条件的行,同时排除All行
filtered_rows = rfm_crosstab[rfm_crosstab['Percentage of 1/Yes'] > mean_value].drop('All')

# 显示结果
print(filtered_rows['Percentage of 1/Yes'])

关键说明

  • 排除汇总行计算均值:用iloc[:-1]取除最后一行(All)外的所有行计算均值,避免汇总值拉低或抬高整体均值。
  • 布尔索引筛选行:将布尔条件直接传入rfm_crosstab[],即可筛选出符合条件的行。
  • 移除汇总行:用drop('All')去掉不需要的汇总行,确保结果只包含实际的RFM类别数据。

这样输出的就是符合条件的Percentage of 1/Yes列的实际数值,而非布尔值。

内容的提问来源于stack exchange,提问作者c200402

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:30:33