多索引多列DataFrame中合并列:生成A+B_new与C_new列的方法
处理多索引多列DataFrame:合并A/B列并重命名C列
针对你拥有的多索引列DataFrame,要实现将每个日期分组下的A列与B列相加生成新列A+B_new,同时将C列重命名为C_new,可以按以下步骤操作:
1. 初始数据回顾
先确认你的初始DataFrame构造代码:
import pandas as pd data = {('2022/10', 'A'): {'ABC': 7, 'CDE': 4, 'FGH': 7}, ('2022/10', 'B'): {'ABC': 3, 'CDE': 3, 'FGH': 6}, ('2022/10', 'C'): {'ABC': 6, 'CDE': 4, 'FGH': 5}, ('2022/11', 'A'): {'ABC': 2, 'CDE': 5, 'FGH': 7}, ('2022/11', 'B'): {'ABC': 5, 'CDE': 8, 'FGH': 4}, ('2022/11', 'C'): {'ABC': 9, 'CDE': 3, 'FGH': 3}, ('2022/12', 'A'): {'ABC': 2, 'CDE': 4, 'FGH': 5}, ('2022/12', 'B'): {'ABC': 6, 'CDE': 7, 'FGH': 4}, ('2022/12', 'C'): {'ABC': 3, 'CDE': 8, 'FGH': 5 }} df = pd.DataFrame(data, index=['ABC','CDE','FGH'])
2. 解决方案代码
分步实现写法
# 按第二级列索引提取A、B列并相加 df_ab_sum = df.xs('A', level=1, axis=1) + df.xs('B', level=1, axis=1) # 为新列设置多索引(日期 + A+B_new) df_ab_sum.columns = pd.MultiIndex.from_tuples([(col, 'A+B_new') for col in df_ab_sum.columns]) # 提取原C列并重命名列名 df_c_renamed = df.xs('C', level=1, axis=1) df_c_renamed.columns = pd.MultiIndex.from_tuples([(col, 'C_new') for col in df_c_renamed.columns]) # 合并两个结果并按列索引排序 df_result = pd.concat([df_ab_sum, df_c_renamed], axis=1).sort_index(axis=1)
简洁链式写法
df_result = pd.concat( [ # 计算A+B并设置新列名 (df.xs('A', level=1, axis=1) + df.xs('B', level=1, axis=1)) .rename(columns=lambda x: (x, 'A+B_new')), # 提取C列并重命名 df.xs('C', level=1, axis=1) .rename(columns=lambda x: (x, 'C_new')) ], axis=1 ).sort_index(axis=1)
3. 代码说明
df.xs('A', level=1, axis=1):通过xs方法快速提取指定级别的列,level=1对应第二级列索引('A'/'B'/'C'),axis=1表示操作列维度。- 直接对提取后的A、B列执行加法运算,得到每个日期分组下的A+B结果,再通过
rename或构造多索引元组,为新列设置(日期, 'A+B_new')的双层索引。 - 提取原C列后,同样修改列索引为
(日期, 'C_new')。 - 用
pd.concat合并两个结果,sort_index(axis=1)保证列按日期顺序排列,与预期结构一致。
验证结果
运行代码后得到的df_result输出如下(注:你提供的期望结果中2022/11、2022/12的C_new值存在误差,以下为真实计算结果):
2022/10 2022/11 2022/12 A+B_new C_new A+B_new C_new A+B_new C_new ABC 10 6 7 9 8 3 CDE 7 4 13 3 11 8 FGH 13 5 11 3 9 5
内容的提问来源于stack exchange,提问作者John Doe
相关产品推荐
相关产品推荐

