You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中按公共索引分组实现多级列拼接?

合并多索引DataFrame并生成多级列头

需求概述

将多个DataFrame合并为单个DataFrame,要求:

  • 结果包含多级列头(顶层为原DataFrame标识,底层为原列名)
  • 按多索引行进行分组合并,即使原DataFrame的索引顺序不一致,也需按索引匹配合并
  • 处理原DataFrame中的重复索引行,合并后每个索引组合仅保留一行

示例数据

DataFrame A

'a' 'b'  
'index4.1'  'index4.2'  'index4.3'  5   7
'index5.1'  'index5.2'  'index5.3'  3   8       
'index1.1'  'index1.2'  'index1.3'  9   8
'index2.1'  'index2.2'  'index2.3'  6   3
'index3.1'  'index3.2'  'index3.3'  8   4
'index4.1'  'index4.2'  'index4.3'  5   7

DataFrame B

'a' 'b' 
'index3.1'  'index3.2'  'index3.3'  1   5       
'index1.1'  'index1.2'  'index1.3'  6   2
'index2.1'  'index2.2'  'index2.3'  4   7
'index4.1'  'index4.2'  'index4.3'  4   7
'index5.1'  'index5.2'  'index5.3'  3   6 

DataFrame C

'b' 'c'        
'index1.1'  'index1.2'  'index1.3'  3   0
'index3.1'  'index3.2'  'index3.3'  1   5
'index4.1'  'index4.2'  'index4.3'  6   7
'index5.1'  'index5.2'  'index5.3'  3   1 
'index2.1'  'index2.2'  'index2.3'  2   1

期望合并结果

'A' 'A' 'B' 'B' 'C' 'C'
                                   'a' 'b' 'a' 'b' 'b' 'c'      
'index1.1'  'index1.2'  'index1.3'  9   8   6   2   3   0
'index2.1'  'index2.2'  'index2.3'  6   3   4   7   2   1
'index3.1'  'index3.2'  'index3.3'  8   4   1   5   1   5
'index4.1'  'index4.2'  'index4.3'  5   7   4   7   6   7
'index5.1'  'index5.2'  'index5.3'  3   8   3   6   3   1

解决方案代码

import pandas as pd

# 构造示例DataFrame A(含重复索引)
data_a = [
    ['index4.1', 'index4.2', 'index4.3', 5, 7],
    ['index5.1', 'index5.2', 'index5.3', 3, 8],
    ['index1.1', 'index1.2', 'index1.3', 9, 8],
    ['index2.1', 'index2.2', 'index2.3', 6, 3],
    ['index3.1', 'index3.2', 'index3.3', 8, 4],
    ['index4.1', 'index4.2', 'index4.3', 5, 7]
]
df_a = pd.DataFrame(data_a, columns=['idx1', 'idx2', 'idx3', 'a', 'b'])
df_a = df_a.set_index(['idx1', 'idx2', 'idx3']).drop_duplicates()  # 去重重复索引行

# 构造示例DataFrame B
data_b = [
    ['index3.1', 'index3.2', 'index3.3', 1, 5],
    ['index1.1', 'index1.2', 'index1.3', 6, 2],
    ['index2.1', 'index2.2', 'index2.3', 4, 7],
    ['index4.1', 'index4.2', 'index4.3', 4, 7],
    ['index5.1', 'index5.2', 'index5.3', 3, 6]
]
df_b = pd.DataFrame(data_b, columns=['idx1', 'idx2', 'idx3', 'a', 'b'])
df_b = df_b.set_index(['idx1', 'idx2', 'idx3'])

# 构造示例DataFrame C
data_c = [
    ['index1.1', 'index1.2', 'index1.3', 3, 0],
    ['index3.1', 'index3.2', 'index3.3', 1, 5],
    ['index4.1', 'index4.2', 'index4.3', 6, 7],
    ['index5.1', 'index5.2', 'index5.3', 3, 1],
    ['index2.1', 'index2.2', 'index2.3', 2, 1]
]
df_c = pd.DataFrame(data_c, columns=['idx1', 'idx2', 'idx3', 'b', 'c'])
df_c = df_c.set_index(['idx1', 'idx2', 'idx3'])

# 为每个DataFrame添加多级列头(顶层为原DataFrame标识)
df_a.columns = pd.MultiIndex.from_product([['A'], df_a.columns])
df_b.columns = pd.MultiIndex.from_product([['B'], df_b.columns])
df_c.columns = pd.MultiIndex.from_product([['C'], df_c.columns])

# 按列合并所有DataFrame,并对行索引排序
merged_df = pd.concat([df_a, df_b, df_c], axis=1).sort_index()

# 输出结果
print(merged_df)

关键步骤说明

  1. 去重处理:对存在重复索引的DataFrame(如示例中的A),使用drop_duplicates()确保每个索引组合唯一,避免合并后出现重复行。
  2. 构建多级列头:通过pd.MultiIndex.from_product为每个DataFrame的列添加顶层标签(A/B/C),实现多级列结构。
  3. 合并与排序:使用pd.concat按列合并所有DataFrame,自动按多索引匹配行;最后用sort_index()对行索引排序,得到有序的结果。

内容的提问来源于stack exchange,提问作者user17647940

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 04:50:23