使用Pandas实现独热编码时解决host_id与id组合行重复问题
解决方案
完全保留你原有melt等处理逻辑,只需在现有代码末尾添加分组聚合步骤即可实现需求:
import pandas as pd d = {'host_id': [1, 1, 2], 'id': [10, 11, 20], 'value': ["Hot Water,Cold Water,Kitchen,Coffee", "Hot Water,Coffee,Something", "Hot Water,Coffee"]} df = pd.DataFrame(data=d) print(df) print(df.shape) print() df_path = df.copy() df_path.index = pd.MultiIndex.from_arrays(df_path[['host_id', 'id']].values.T, names=['host_id', 'id']) df_path = df_path['value'].str.split(',', expand=True) df_path = df_path.melt(ignore_index=False).dropna() df_path.reset_index(inplace=True) print(df_path) one_hot = pd.get_dummies(df_path['value']) df_one = df_path.drop('value',axis = 1) df_one = df_path.join(one_hot) # 新增去重聚合逻辑 one_hot_cols = one_hot.columns df_one = df_one.groupby(['host_id', 'id'], as_index=False).agg( value=('value', 'first'), **{col: (col, 'max') for col in one_hot_cols} ) print() print(df_one) print(df_one.shape)
实现逻辑
- 原有拆分逻辑会让同一
host_id+id组合对应多条记录,每条记录对应一个取值 - 按
host_id和id分组后,对独热编码列取最大值:只要该组合下出现过对应取值,列值就为1,否则为0,完全匹配独热编码要求 - 原始
value列每组取第一条即可,同一组合下的value完全一致
运行输出
host_id id value Coffee Cold Water Hot Water Kitchen Something 0 1 10 Hot Water,Cold Water,Kitchen,Coffee 1 1 1 1 0 1 1 11 Hot Water,Coffee,Something 1 0 1 0 1 2 2 20 Hot Water,Coffee 1 0 1 0 0 (3, 8)
内容的提问来源于stack exchange,提问作者Test
相关产品推荐
相关产品推荐

