如何一次性合并多个同结构Pandas DataFrame并添加来源标识列?
问题
我有多个列结构相同的Pandas DataFrame,这些DataFrame以字典形式存储(如df['A']、df['B']、df['F']至df['H']等),每个DataFrame对应不同的名称标识。现希望一次性将它们合并为一个DataFrame,同时新增一列来标记每条数据所属的原DataFrame名称。
示例输入
# df['A'] col1 col2 1 x 2 y # df['B'] col1 col2 1 x5 2 y2 # df['F'] col1 col2 13 x3 2 y3 ... # df['H'] col1 col2 13 x1 23 y2
期望输出
File col1 col2 A 1 x A 2 y B 1 x5 B 2 y2 F 13 x3 F 2 y3 ... ... ... H 13 x1 H 23 y2
解决方法
方式一:用pd.concat一键合并(推荐)
直接用Pandas自带的合并功能,自动添加来源标识,代码简洁高效:
import pandas as pd # 假设存储DataFrame的字典名为df,替换成你的实际字典名即可 combined_df = pd.concat(df.values(), keys=df.keys(), names=['File']) # 将File从索引转换为普通列,并清理多余的行索引 combined_df = combined_df.reset_index(level='File').reset_index(drop=True)
处理后得到的combined_df就完全符合需求,File列会清晰标记每条数据的来源。
方式二:循环手动添加列再合并
如果需要对来源标识做自定义处理(比如添加前缀、修改格式),可以用循环遍历字典,给每个DataFrame手动添加标识列后再合并:
import pandas as pd temp_dfs = [] for file_name, sub_df in df.items(): # 给当前DataFrame新增File列,值为对应的名称标识 sub_df['File'] = file_name temp_dfs.append(sub_df) # 合并所有临时DataFrame combined_df = pd.concat(temp_dfs, ignore_index=True) # 可选:调整列顺序,把File列放到最前面 combined_df = combined_df[['File', 'col1', 'col2']]
这种方式灵活性更高,能满足更多自定义需求。
内容的提问来源于stack exchange,提问作者navee pp
相关产品推荐
相关产品推荐

