You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中基于多列条件匹配更新旧DataFrame并新增行的问题

Pandas DataFrame 更新逻辑实现问题

需求说明

我有两个结构相似的DataFrame:old_df(包含date/time、Name、detect_ID、category、ID列)和new_df,需要按以下规则更新old_df:

  • 仅处理old_df中category为'B'的行;
  • 遍历new_df的每一行:
    • 若该行的ID在old_df中无匹配,将该行添加到old_df,新增Identify列并赋值为'new';
    • 若ID匹配,将detect_ID按逗号拆分转为整数逐一比对:
      1. 若detect_ID匹配,用new_df的该行替换old_df的对应行,Identify赋值为'updated';
      2. 若detect_ID不匹配,将该行添加到old_df,Identify赋值为'new';
  • old_df中满足以下条件的行保持不变,Identify赋值为'unchanged':
    • ID未在new_df中匹配;
    • ID匹配但detect_ID不匹配。

预期输出示例

>output
date/time    Name   detect_ID   category  ID   identify
13/1/2023    XXX    1           B        1400   updated  [Case A] 
14/1/2023    XXY    1           B        1402   updated  [Case A with multiple detect_ID]
14/1/2023    XXY    3           B        1402   updated
12/1/2023    XXY    7           B        1402   unchanged  [Step 3, Id matches but detect_id do not ]
14/1/2023    XXY    8           B        1402   new        [Case B]
12/1/2023    XXY    4           A        1403   unchanged   
12/1/2023    XXY    4           B        1407   unchanged [Step3 , id not found in new_df]

当前问题代码

当前代码会生成大量重复行,且未遍历足够的old_df行,无法达到预期效果:

old_df = pd.read_csv('old.csv')
new_df = pd.read_csv('new.csv')

# 生成旧数据集中唯一的(ID, 检测器ID)元组集合
unique_pairs = set()
for _, row in old_df.iterrows():
    detector_ids = [int(x) for x in str(row['Detect_ID']).split(',')]
    for detector_id in detector_ids:
        unique_pairs.add((row['ID'], detector_id))

# 遍历新数据集,检查每行的(ID, 检测器ID)是否在旧数据集的集合中
new_rows = []
updated_rows = []
for _, row in new_df.iterrows():
    detector_ids = [int(x) for x in str(row['Detect_ID']).split(',')]
    for detector_id in detector_ids:
        if (row['ID'], detector_id) in unique_pairs:
            old_row = old_df.loc[(old_df['ID'] == row['ID']) & (old_df['Detect_ID'].str.contains(str(detector_id)))]
            if not old_row.empty:
                old_row = old_row.iloc[0]
                old_row['Date/Time'] = row['date/time']
                old_df.loc[(old_df['ID'] == row['ID']) & (old_df['Detector_ID'].str.contains(str(detector_id))), 'date/time'] = old_row['date/time']
                updated_rows.append(old_row)
        else:
            row['Identify'] = 'new'
            new_rows.append(row)
            unique_pairs.add((row['ID'], detector_id))

# 将新增行追加到旧数据集,并将更新行也追加进去,最后写入文件
old_df = old_df.append(new_rows, ignore_index=True)
for row in updated_rows:
    row['Identify'] = 'updated'
old_df = old_df.append(updated_rows, ignore_index=True)
old_df.to_csv('updated.csv', index=False)

内容的提问来源于stack exchange,提问作者Jewel_R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 07:07:49