如何创建包含原DataFrame及另一DataFrame非重复索引行的新DataFrame?
Pandas合并DataFrame:保留原数据+新增未重复索引行
原始数据
import pandas as pd base_df = pd.DataFrame({'animal':['dog', 'cat'], 'noise':['bark', 'meow']}, index=[1,2]) new_info_df = pd.DataFrame({'animal':['wolf', 'sheep'], 'noise':['woof', 'baa']}, index=[1,3])
需求说明
需要生成新的DataFrame,包含base_df的所有数据,同时加入new_info_df中索引未在base_df出现的行,最终目标结果如下:
updated_df = pd.DataFrame({'animal':['dog', 'cat', 'sheep'], 'noise':['bark', 'meow', 'baa']}, index=[1,2,3])
问题分析
使用pd.concat([base_df, new_info_df], join='outer', axis=1)时,会因按列拼接导致animal和noise列重复(自动命名为animal_0、animal_1等),无法满足需求。
解决方案
方法1:使用combine_first(推荐)
combine_first会优先保留base_df的所有数据,自动用new_info_df填充base_df中缺失的索引行,完全匹配需求:
updated_df = base_df.combine_first(new_info_df) # 可选:按索引排序,保证结果顺序和目标一致 updated_df = updated_df.sort_index()
方法2:筛选新增行后拼接
先从new_info_df中筛选出索引不在base_df里的行,再和base_df按行拼接:
# 筛选new_info_df中未在base_df出现的索引对应的行 new_rows = new_info_df[~new_info_df.index.isin(base_df.index)] # 按行合并并排序索引 updated_df = pd.concat([base_df, new_rows]).sort_index()
内容的提问来源于stack exchange,提问作者SPAMINACAN
相关产品推荐
相关产品推荐

