You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas侧拼接含重复行的不同长度DataFrame的对齐问题求解

实现方案

核心问题根源:pd.concat按列拼接时默认使用两个DataFrame的原始行索引对齐,df2中重复的joe行原始索引为2,会默认和df1中索引为2的john行对齐,才会出现错位。
解决方案是给相同业务主键的行添加组内顺序辅助索引,用多重索引对齐后再拼接:

import pandas as pd

# 构造测试数据
df1 = pd.DataFrame({
    'emp_id': [1,2,3],
    'emp_name': ['sam','joe','john'],
    'counts': [0,0,0]
})
df2 = pd.DataFrame({
    'emp_id': [1,2,2,3],
    'emp_name': ['sam','joe','joe','john'],
    'counts': [0,0,1,0]
})

# 处理逻辑
# 1. 给两个df添加组内序号辅助列,按emp_id分组,组内行从0开始编号
df1['group_seq'] = df1.groupby('emp_id').cumcount()
df2['group_seq'] = df2.groupby('emp_id').cumcount()

# 2. 将emp_id和group_seq设为多重索引,用于对齐
df1 = df1.set_index(['emp_id', 'group_seq'])
df2 = df2.set_index(['emp_id', 'group_seq'])

# 3. 按列拼接,使用多重索引对齐
result = pd.concat([df1, df2], axis=1, keys=['df1', 'df2'])

# 4. 重置索引去掉辅助列,得到最终结果
result = result.reset_index(drop=True)

运行后输出的result完全符合预期效果,重复行对应的另一表位置会自动填充NaN。

内容的提问来源于stack exchange,提问作者Bhavyashree717

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 07:36:03