You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何一次性合并多个同结构Pandas DataFrame并添加来源标识列?

问题

我有多个列结构相同的Pandas DataFrame,这些DataFrame以字典形式存储(如df['A']、df['B']、df['F']至df['H']等),每个DataFrame对应不同的名称标识。现希望一次性将它们合并为一个DataFrame,同时新增一列来标记每条数据所属的原DataFrame名称。

示例输入

# df['A']
col1  col2
 1     x
 2     y

# df['B']
col1  col2
 1     x5
 2     y2

# df['F']
col1  col2
 13    x3
 2     y3

...

# df['H']
col1  col2
 13    x1
 23    y2

期望输出

File  col1    col2 
 A     1       x
 A     2       y
 B     1       x5
 B     2       y2
 F     13      x3
 F     2       y3
 ...   ...     ...
 H     13      x1
 H     23      y2

解决方法

方式一:用pd.concat一键合并(推荐)

直接用Pandas自带的合并功能,自动添加来源标识,代码简洁高效:

import pandas as pd

# 假设存储DataFrame的字典名为df,替换成你的实际字典名即可
combined_df = pd.concat(df.values(), keys=df.keys(), names=['File'])
# 将File从索引转换为普通列,并清理多余的行索引
combined_df = combined_df.reset_index(level='File').reset_index(drop=True)

处理后得到的combined_df就完全符合需求,File列会清晰标记每条数据的来源。

方式二:循环手动添加列再合并

如果需要对来源标识做自定义处理(比如添加前缀、修改格式),可以用循环遍历字典,给每个DataFrame手动添加标识列后再合并:

import pandas as pd

temp_dfs = []
for file_name, sub_df in df.items():
    # 给当前DataFrame新增File列,值为对应的名称标识
    sub_df['File'] = file_name
    temp_dfs.append(sub_df)

# 合并所有临时DataFrame
combined_df = pd.concat(temp_dfs, ignore_index=True)
# 可选:调整列顺序,把File列放到最前面
combined_df = combined_df[['File', 'col1', 'col2']]

这种方式灵活性更高,能满足更多自定义需求。

内容的提问来源于stack exchange,提问作者navee pp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 17:17:25