Pandas DataFrame重复数据处理异常求助(需保留for循环)
问题排查与修复
问题根源
你的代码中,dict_temp_df = df.to_dict("records")是在循环开始前生成的,它保存的是DataFrame的初始数据快照。当你在循环中通过df.at[i+1, ...]更新df的Value时,dict_temp_df里的数值并不会同步更新。这就导致后续循环中计算最大值时,使用的始终是初始的旧值,而非更新后的值。
举个具体的例子:
- 第一次循环(i=0):将行1的Value更新为
max(1.36,1.35)=1.36,但dict_temp_df[1]["Value"]仍然是初始的1.35。 - 第二次循环(i=1):计算
max(dict_temp_df[1]["Value"], dict_temp_df[2]["Value"])时,用的是1.35和1.35,结果还是1.35,最终行2的Value被设置为1.35,而非预期的1.36。
修复后的代码
直接从DataFrame中获取实时的数值,不再依赖静态的dict_temp_df:
df = df.sort_values(by=['ID', 'RecordingType', 'Date'], ascending=True).reset_index(drop=True) df["ToRemove"] = False outputcolnames = {'FEVR':'Value'} for i in range(df.shape[0]-1): curr_id = df.at[i, "ID"] next_id = df.at[i+1, "ID"] curr_recording_type = df.at[i, "RecordingType"] next_recording_type = df.at[i+1, "RecordingType"] curr_date = df.at[i, "Date"] next_date = df.at[i+1, "Date"] # 检查ID、RecordingType是否相同,且时间差小于60分钟 if curr_id == next_id and curr_recording_type == next_recording_type and abs((curr_date - next_date).total_seconds() / 60) < 60: df.at[i, "ToRemove"] = True if curr_recording_type == 'FEVR': # 直接从df中取当前和下一行的实时Value计算最大值 curr_val = df.at[i, outputcolnames[curr_recording_type]] next_val = df.at[i+1, outputcolnames[next_recording_type]] df.at[i+1, outputcolnames[next_recording_type]] = max(curr_val, next_val) else: curr_val = df.at[i, outputcolnames[curr_recording_type]] df.at[i+1, outputcolnames[next_recording_type]] += curr_val # 删除标记行 df = df[df["ToRemove"] == False].reset_index(drop=True)
修复说明
- 移除了静态的
dict_temp_df,改用df.at[...]直接获取和修改DataFrame的实时数据,确保每次计算都用最新的数值。 - 用
reset_index(drop=True)替代原有的reset_index().drop(columns=["index"]),代码更简洁。 - 拆分了条件判断中的变量,提升代码可读性,也方便后续扩展其他RecordingType的分支逻辑。
运行修复后的代码,会得到预期结果:
ID RecordingType Date Value 1 FEVR 2019-05-22 18:45:16 1.36
内容的提问来源于stack exchange,提问作者mariant
相关产品推荐
相关产品推荐

