如何将拆分扩展后的DataFrame列值转换为频次统计新列?
问题描述
我通过拆分某一列并扩展创建了一个新的DataFrame。现在我想要转换该DataFrame,为每个值创建新列,并仅展示该值的出现频次。
示例DataFrame
import pandas as pd import numpy as np df = pd.DataFrame({ 0: ['cake','fries', 'ketchup', 'potato', 'snack'], 1: ['fries', 'cake', 'potato', np.nan, 'snack'], 2: ['ketchup', 'cake', 'potatos', 'snack', np.nan], 3: ['potato', np.nan,'cake', 'ketchup',np.nan], 'index': ['james','samantha','ashley','tim', 'mo'] }) df.set_index('index')
预期输出
output = pd.DataFrame({ 'cake': [1, 2, 1, 0, 0], 'fries': [1, 1, 0, 0, 0], 'ketchup': [1, 0, 1, 1, 0], 'potatoes': [1, 0, 2, 1, 0], 'snack': [0, 0, 0, 1, 2], 'index': ['james', 'samantha', 'ashley', 'tim', 'mo'] }) output.set_index('index')
解决方案
可以通过Pandas的stack()、groupby()和unstack()方法实现需求,步骤如下:
- 将
index列设为数据框索引,方便按用户分组统计 - 用
stack()把每行多列数据转成长格式,自动忽略空值NaN - 按原索引分组,统计每个值的出现次数
- 用
unstack()转回宽格式,空值用0填充 - 若需要统一近似值(比如示例中的
potato和potatos),先做替换处理
完整代码
import pandas as pd import numpy as np # 初始化原数据框 df = pd.DataFrame({ 0: ['cake','fries', 'ketchup', 'potato', 'snack'], 1: ['fries', 'cake', 'potato', np.nan, 'snack'], 2: ['ketchup', 'cake', 'potatos', 'snack', np.nan], 3: ['potato', np.nan,'cake', 'ketchup',np.nan], 'index': ['james','samantha','ashley','tim', 'mo'] }).set_index('index') # 统一近似值,匹配预期输出中的potatoes stacked_data = df.stack().replace({'potato': 'potatoes', 'potatos': 'potatoes'}) # 按用户分组统计各值出现次数 counts = stacked_data.groupby(level=0).value_counts() # 转换为宽格式并填充空值为0 result = counts.unstack(fill_value=0) # 调整列顺序以匹配预期输出(可选) result = result[['cake', 'fries', 'ketchup', 'potatoes', 'snack']] print(result)
运行结果
cake fries ketchup potatoes snack index james 1 1 1 1 0 samantha 2 1 0 0 0 ashley 1 0 1 2 0 tim 0 0 1 1 1 mo 0 0 0 0 2
内容的提问来源于stack exchange,提问作者Nazar
相关产品推荐
相关产品推荐

