如何用Pandas覆盖DataFrame中已有ID对应的特征值?
问题:DataFrame重复添加相同ID行而非覆盖的解决方法
问题代码与现象
初始化代码
# Append columns to an empty DataFrame. self.df = pd.DataFrame(columns = ["ID", "Features"],index=['index1'])
逻辑代码
tracking_id = output[4] print(tracking_id in set(self.df['ID'])) if tracking_id in self.df['ID'] : df2 = pd.DataFrame([[tracking_id], features]) self.df.update(df2) print(self.df) else : self.df = self.df.append({'ID' : tracking_id, 'Features' : features}, ignore_index = True)
预期与实际偏差
预期逻辑:检查DataFrame中是否存在相同ID,存在则用新特征覆盖旧值,不存在则添加新行。
实际问题:相同ID的行被重复添加,偶尔能正常覆盖,输出示例:
ID Features 0 NaN NaN 1 1.0 [[1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 50.0... 2 4.0 [[0.0, 0.0, 0.0, 1.0, 89.0, 15.0, 0.0, 0.0, 10... 3 4.0 [[70.0, 41.0, 17.0, 41.0, 4.0, 0.0, 0.0, 0.0, ... 4 4.0 [[42.0, 18.0, 16.0, 14.0, 2.0, 0.0, 0.0, 0.0, ... 5 6.0 [[3.0, 0.0, 0.0, 0.0, 0.0, 0.0, 3.0, 8.0, 59.0... 6 6.0 [[0.0, 6.0, 7.0, 9.0, 12.0, 3.0, 0.0, 0.0, 51....
问题原因与修复方案
1. 核心问题点
- 初始化时指定
index=['index1'],导致DataFrame默认生成一行NaN数据,后续判断ID存在性时容易出现类型不匹配(比如tracking_id是整数,DataFrame中ID是浮点数),导致判断失效。 - 原更新逻辑中
df2的构造错误,生成的是两行数据而非一行两列,update方法无法正确匹配目标行。 append方法已被pandas弃用,且原逻辑中忽略索引的处理会加剧数据混乱。
2. 修复步骤
步骤1:正确初始化空DataFrame
去掉多余的index参数,避免默认生成NaN行:
self.df = pd.DataFrame(columns=["ID", "Features"])
步骤2:统一ID类型并准确判断存在性
确保tracking_id与DataFrame中ID的类型一致,用isin方法更可靠地判断存在性:
tracking_id = output[4] # 根据实际数据类型调整,比如转为整数或字符串 tracking_id = int(tracking_id) # 检查ID是否存在 id_exists = self.df['ID'].isin([tracking_id]).any()
步骤3:正确更新或添加数据
- 存在ID时,直接通过
loc定位到对应行更新特征值; - 不存在ID时,用
pd.concat替代append添加新行:
if id_exists: # 定位到对应ID的行,更新Features列 self.df.loc[self.df['ID'] == tracking_id, 'Features'] = features else: # 构造新行DataFrame并合并 new_row = pd.DataFrame([{'ID': tracking_id, 'Features': features}]) self.df = pd.concat([self.df, new_row], ignore_index=True)
3. 完整修复代码
# 初始化空DataFrame self.df = pd.DataFrame(columns=["ID", "Features"]) tracking_id = output[4] # 统一ID类型,根据实际情况调整(如字符串则用str(tracking_id)) tracking_id = int(tracking_id) id_exists = self.df['ID'].isin([tracking_id]).any() if id_exists: self.df.loc[self.df['ID'] == tracking_id, 'Features'] = features else: new_row = pd.DataFrame([{'ID': tracking_id, 'Features': features}]) self.df = pd.concat([self.df, new_row], ignore_index=True)
内容的提问来源于stack exchange,提问作者HusCet
相关产品推荐
相关产品推荐

