Pandas多级索引DataFrame添加相邻空列时出现重复问题求助
解决DataFrame多级列添加子列的问题
问题分析
你当前的循环错误在于遍历了已经是MultiIndex类型的列(如('Pants', 'self')),添加新列时会把这个元组作为顶级列名的一部分,导致出现重复的顶级列(Pants出现两次,分别对应self和other),而不是在同一顶级列下扩展子列。
正确解决方案
方法1:遍历顶级列名添加子列
先将普通列转为带self的MultiIndex,再遍历顶级列名添加other子列,最后排序列确保结构一致:
import pandas as pd # 模拟你的初始DataFrame df_initial_unique = pd.DataFrame( {'Pants': ['yes', 'no', 'no'], 'Jacket': ['yes', 'no', 'no']}, index=['Denim', 'Cotton', 'Silk'] ) # 第一步:将列转为MultiIndex,添加self子列 df_initial_unique.columns = pd.MultiIndex.from_tuples( [(col, 'self') for col in df_initial_unique.columns] ) # 第二步:遍历顶级列名,添加other子列 top_cols = df_initial_unique.columns.get_level_values(0).unique() for col in top_cols: df_initial_unique[(col, 'other')] = 'new' # 第三步:按列索引排序,让同一顶级列的self/other相邻 df_initial_unique = df_initial_unique.sort_index(axis=1)
执行后即可得到目标结构:
Pants Jacket self other self other Denim yes new yes new Cotton no new no new Silk no new no new
方法2:用concat拼接两个子DataFrame
更简洁的方式是分别构建self和other对应的MultiIndex列DataFrame,再横向拼接:
import pandas as pd # 初始普通DataFrame df = pd.DataFrame( {'Pants': ['yes', 'no', 'no'], 'Jacket': ['yes', 'no', 'no']}, index=['Denim', 'Cotton', 'Silk'] ) # 构建带self子列的DF df_self = df.copy() df_self.columns = pd.MultiIndex.from_tuples([(col, 'self') for col in df_self.columns]) # 构建带other子列的DF,值全为'new' df_other = pd.DataFrame( 'new', index=df.index, columns=pd.MultiIndex.from_tuples([(col, 'other') for col in df.columns]) ) # 拼接并排序列 result = pd.concat([df_self, df_other], axis=1).sort_index(axis=1)
错误原因说明
你之前的循环for col in df_initial_unique.columns中,col是完整的MultiIndex元组(如('Pants', 'self')),此时df_initial_unique[(col,"other")] = "new"会创建一个三级列索引(('Pants', 'self'), 'other'),最终导致顶级列重复。正确的做法是遍历顶级列名(如'Pants'),而非已有的MultiIndex列元组。
内容的提问来源于stack exchange,提问作者Imakeweirdstuff
相关产品推荐
相关产品推荐

