Pandas中移除两DataFrame共享ISBN数据或更新价格折扣的方法
Pandas 处理两个DataFrame的ISBN匹配问题
解决方案1:生成不含双方共享ISBN的数据集
你之前的代码错误在于把DataFrame名称stop当成了ISBN值来比较,正确的做法是先提取两个数据集的共享ISBN,再过滤掉这些记录:
- 提取共享ISBN列表:
shared_isbns = db7['ISBN'].intersection(stop['ISBN'])
- 分别过滤出两个DataFrame中独有的记录,再合并成最终数据集:
# 从db7中排除共享ISBN db7_unique = db7[~db7['ISBN'].isin(shared_isbns)] # 从stop中排除共享ISBN stop_unique = stop[~stop['ISBN'].isin(shared_isbns)] # 合并两个独有数据集 combined_unique = pd.concat([db7_unique, stop_unique], ignore_index=True)
解决方案2:用stop的新值替换db7中的旧零售价和折扣
方法一:通过合并+填充实现(保留所有原列)
这种方法可以清晰处理匹配/未匹配的情况,且不会修改原DataFrame:
import pandas as pd # 只合并stop中需要的列,左连接保留db7所有行 merged = db7.merge( stop[['ISBN', 'retail', 'discount']], on='ISBN', how='left', suffixes=('_old', '') # 给db7的旧列加后缀区分 ) # 用stop的新值填充,未匹配到的保留db7原数据 merged['retail'] = merged['retail'].combine_first(merged['retail_old']) merged['discount'] = merged['discount'].combine_first(merged['discount_old']) # 移除旧列,得到最终更新后的db7 final_db7 = merged.drop(columns=['retail_old', 'discount_old'])
方法二:通过索引更新实现(更简洁)
如果只需要更新对应列,且可以接受修改原数据副本,可用索引匹配的方式:
# 将两个DataFrame都设置ISBN为索引 db7_copy = db7.set_index('ISBN') stop_update = stop.set_index('ISBN')[['retail', 'discount']] # 用stop的新值更新db7副本,仅更新匹配到的ISBN对应行 db7_copy.update(stop_update) # 恢复ISBN为普通列 final_db7 = db7_copy.reset_index()
内容的提问来源于stack exchange,提问作者jumpman23
相关产品推荐
相关产品推荐

