You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建包含原DataFrame及另一DataFrame非重复索引行的新DataFrame?

Pandas合并DataFrame:保留原数据+新增未重复索引行

原始数据

import pandas as pd

base_df = pd.DataFrame({'animal':['dog', 'cat'], 'noise':['bark', 'meow']}, index=[1,2])
new_info_df = pd.DataFrame({'animal':['wolf', 'sheep'], 'noise':['woof', 'baa']}, index=[1,3])

需求说明

需要生成新的DataFrame,包含base_df的所有数据,同时加入new_info_df中索引未在base_df出现的行,最终目标结果如下:

updated_df = pd.DataFrame({'animal':['dog', 'cat', 'sheep'], 'noise':['bark', 'meow', 'baa']}, index=[1,2,3])

问题分析

使用pd.concat([base_df, new_info_df], join='outer', axis=1)时,会因按列拼接导致animal和noise列重复(自动命名为animal_0、animal_1等),无法满足需求。

解决方案

方法1:使用combine_first(推荐)

combine_first会优先保留base_df的所有数据,自动用new_info_df填充base_df中缺失的索引行,完全匹配需求:

updated_df = base_df.combine_first(new_info_df)
# 可选:按索引排序,保证结果顺序和目标一致
updated_df = updated_df.sort_index()

方法2:筛选新增行后拼接

先从new_info_df中筛选出索引不在base_df里的行,再和base_df按行拼接:

# 筛选new_info_df中未在base_df出现的索引对应的行
new_rows = new_info_df[~new_info_df.index.isin(base_df.index)]
# 按行合并并排序索引
updated_df = pd.concat([base_df, new_rows]).sort_index()

内容的提问来源于stack exchange,提问作者SPAMINACAN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 15:12:34