如何生成独热编码列并将df1的months值映射填充到df2?
问题描述
现有两个DataFrame:
- df1(形状(52,2)),前5行数据如下:
id months 71911200001 22 71911200002 27 71911200004 30 71911200003 23 41911200003 35
- df2(形状(52,49)),已将
id设为行索引,前5行及列结构如下:
id M0 M1 M2 M3 M4.....M49 71911200001 0 0 0 0 0 0 71911200002 0 0 0 0 0 0 71911200004 0 0 0 0 0 0 71911200003 0 0 0 0 0 0 41911200003 0 0 0 0 0 0
需求:对每个id,当df2列名的数字部分与df1对应id的months值匹配时,将该列设为1,其余列保持0。两个DataFrame的id完全一致且无重复。
解决方案
方法一:循环赋值(直观易懂)
- 先将df1的
id设为行索引,与df2对齐:
df1 = df1.set_index('id')
- 生成每个id对应的目标列名:
target_cols = 'M' + df1['months'].astype(str)
- 遍历赋值:
for idx, col in target_cols.items(): df2.loc[idx, col] = 1
方法二:向量化操作(高效无循环)
适合数据量较大的场景,避免循环带来的性能损耗:
# 重置df1的索引,方便后续生成匹配掩码 df1_reset = df1.reset_index() # 生成布尔掩码:每一行对应df1中id的months匹配的列 mask = df2.columns == 'M' + df1_reset['months'].astype(str)[:, None] # 将掩码匹配的位置设为1,其余保持原0值 df2 = df2.where(~mask, 1)
内容的提问来源于stack exchange,提问作者R_Sengupta
相关产品推荐
相关产品推荐

