如何高效无循环地同时追加并覆盖Pandas DataFrame指定列?
高效实现DataFrame的覆盖更新与追加操作
问题场景
给定如下结构的Pandas DataFrame:
T x c1 c2 1 11-12 3 'yes' 'yes' 2 12-12 4 'no' 'yes' 3 13-12 4 'no' 'yes' 4 14-12 4 'yes' 'yes' 5 15-12 2 'no' 'no' 6 16-12 4 'yes' 'yes'
需要从索引4开始添加5条包含T和x的记录:前3条覆盖原有索引4-6的x值,后2条作为新索引7-8追加,最终保留原有c1、c2列的已有值,新增索引对应的c1、c2为NaN,结果如下:
T x c1 c2 1 11-12 3 'yes' 'yes' 2 12-12 4 'no' 'yes' 3 13-12 4 'no' 'yes' 4 14-12 3 'yes' 'yes' 5 15-12 3 'no' 'no' 6 16-12 4 'yes' 'yes' 7 17-12 4 nan nan 8 18-12 2 nan nan
简洁解决方案
可以通过reindex和combine_first组合实现,无需拆分数据分步骤操作:
- 先定义待添加的新数据:
import pandas as pd # 原DataFrame df = pd.DataFrame({ 'T': ['11-12', '12-12', '13-12', '14-12', '15-12', '16-12'], 'x': [3, 4, 4, 4, 2, 4], 'c1': ['yes', 'no', 'no', 'yes', 'no', 'yes'], 'c2': ['yes', 'yes', 'yes', 'yes', 'no', 'yes'] }, index=[1,2,3,4,5,6]) # 待添加的新数据 new_data = pd.DataFrame({ 'T': ['14-12', '15-12', '16-12', '17-12', '18-12'], 'x': [3, 3, 4, 4, 2] }, index=[4,5,6,7,8])
- 执行合并与更新:
# 扩展原DataFrame的索引,包含新数据的所有索引 result = df.reindex(df.index.union(new_data.index)) # 用新数据的x值覆盖对应索引,保留原有索引的x值 result['x'] = new_data['x'].combine_first(result['x'])
原理说明
reindex:将原DataFrame的索引扩展为原索引与新数据索引的并集,新增的索引(7-8)对应的列值默认填充为NaN,原有索引的c1、c2值保持不变。combine_first:优先使用新数据中的x值,当新数据中某索引无对应值时(即原索引1-3),保留原DataFrame的x值,完美实现“覆盖已有索引+追加新索引”的需求。
内容的提问来源于stack exchange,提问作者RiiNagaja
相关产品推荐
相关产品推荐

