You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拼接含重复索引、列数不同的DataFrame字典?

解决方法

可以直接使用pandas的pd.concat()函数处理这个需求,它会自动对齐所有DataFrame的列,缺失列自动填充为NaN,同时可重置索引生成连续编号。

步骤示例

  1. 准备你的DataFrame字典
  2. 使用pd.concat()合并所有DataFrame,开启ignore_index=True重置索引
  3. (可选)将新索引转为名为idx的列,匹配期望输出格式

代码实现

import pandas as pd

# 构建示例DataFrame(替换为你的实际数据)
df1 = pd.DataFrame({'col1': [1, 2], 'col2': [1, 2], 'col3': [1, 2]})
df2 = pd.DataFrame({'col1': [1, 2], 'col3': [1, 2]})
df3 = pd.DataFrame({'col1': [1, 2], 'col2': [1, 2], 'col3': [1, 2]})

# 存放DataFrame的字典
df_dict = {"df1": df1, "df2": df2, "df3": df3}

# 合并并重置索引
combined_df = pd.concat(df_dict.values(), ignore_index=True)
# 将索引转为显式的idx列
combined_df = combined_df.reset_index(names="idx")

print(combined_df)

输出结果

idx  col1  col2  col3
0    0     1   1.0     1
1    1     2   2.0     2
2    2     1   NaN     1
3    3     2   NaN     2
4    4     1   1.0     1
5    5     2   2.0     2

关键说明

  • pd.concat(df_dict.values()):自动收集所有DataFrame的列,缺失列填充为NaN,完成纵向拼接
  • ignore_index=True:丢弃原DataFrame的重复索引,生成从0开始的连续新索引
  • reset_index(names="idx"):将新索引转换为显式的idx列,完全匹配期望的输出格式

内容的提问来源于stack exchange,提问作者Sherwin R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 18:50:44