You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python合并多层索引DataFrame的方法咨询

解决多级索引DataFrame合并问题

你的需求可以通过先拼接互补的df2和df3,再与df1按索引关联实现,完全避免重复列和NaN值:

步骤1:拼接df2与df3

df2和df3的列名一致,且索引分别对应index2='a'和index2='b'的分组,用concat纵向拼接即可得到包含全部索引组合的完整col3/col4数据:

import pandas as pd

# 初始化你的三个DataFrame
df1=pd.DataFrame({'col1': ['1a', '1b', '1c', '1d'], 'col2': ['2a','2b','2c','2d'],}, 
index=[[1,2,1,2],['a','a','b','b']])
df1.index.names = ['index1', 'index2']

df2=pd.DataFrame({'col3': ['3a','3b'], 'col4': ['4a','4b']}, index=[[1,2],['a','a']])
df2.index.names = ['index1', 'index2']

df3=pd.DataFrame({'col3': ['3c','3d'], 'col4': ['4c','4d']}, index=[[1,2],['b','b']])
df3.index.names = ['index1', 'index2']

# 拼接df2和df3,得到完整的col3/col4数据集
df23 = pd.concat([df2, df3])

步骤2:将df1与拼接后的df23按索引关联

由于df1和df23的多级索引完全匹配,直接用join按索引合并即可得到目标结果:

df5 = df1.join(df23)

验证结果

输出df5会和你期望的结构完全一致:

col1 col2 col3 col4
index1 index2                 
1      a      1a   2a  3a  4a
2      a      1b   2b  3b  4b
1      b      1c   2c  3c  4c
2      b      1d   2d  3d  4d

为什么之前的操作会出问题?

  • 若直接用df1.join(df2).join(df3),会因df2仅匹配index2='a'的行、df3仅匹配index2='b'的行,导致非匹配行出现NaN;
  • 若用merge多次合并,会因列名重复生成col3_x、col3_y这类重复列;
  • 先拼接df2和df3得到完整的补全数据,再与df1关联,就能从根源避免这两个问题。

内容的提问来源于stack exchange,提问作者Wick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 13:32:50