You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写Python函数循环处理DataFrame列表并移除合并重复行?

解决方案

你需要的是一个能迭代处理多个月度DataFrame、每次保留仅单侧存在行的函数,核心思路是每次合并后取对称差集(即排除两边都有的行),以下是简洁的实现:

函数代码

import pandas as pd

def merge_unique(main_df, monthly_dfs):
    # 复制主DataFrame,避免修改原始数据
    current_result = main_df.copy()
    # 遍历所有月度DataFrame
    for monthly_df in monthly_dfs:
        # 外连接合并并标记行来源
        merged = pd.merge(current_result, monthly_df, how='outer', indicator=True)
        # 筛选仅单侧存在的行,移除标记列
        current_result = merged[merged['_merge'] != 'both'].drop('_merge', axis=1)
    return current_result

用法说明

  1. 参数说明

    • main_df:初始的"Main" DataFrame
    • monthly_dfs:月度DataFrame的列表,最多可传入12个,直接添加到列表即可,无需修改函数
  2. 示例验证
    先构造你的示例数据:

    # 构造Main DataFrame
    main_df = pd.DataFrame({
        'Name': ['Bob', 'Dirk', 'Steve'],
        'Date': ['03/10/2022', '05/12/2022', '01/13/2022'],
        'Begin Time': ['11:04', '13:15', '11:11'],
        'End Time': ['14:10', '16:56', '13:13']
    })
    
    # 构造Other DataFrame(作为月度数据示例)
    other_df = pd.DataFrame({
        'Name': ['Rog', 'Dirk', 'Steve'],
        'Date': ['03/14/2022', '05/12/2022', '01/13/2022'],
        'Begin Time': ['11:44', '13:15', '11:11'],
        'End Time': ['14:30', '16:56', '13:13']
    })
    

    调用函数获取结果:

    # 传入月度数据列表,可添加更多月度DataFrame到列表中
    final_result = merge_unique(main_df, [other_df])
    print(final_result)
    

    输出结果与预期一致:

    Name        Date Begin Time End Time
    0    Bob  03/10/2022      11:04    14:10
    3    Rog  03/14/2022      11:44    14:30
    

扩展说明

如果不需要基于所有列判断重复行,而是仅基于特定列(比如Name和Date),可以在merge时指定on参数:

merged = pd.merge(current_result, monthly_df, how='outer', on=['Name', 'Date'], indicator=True)

这样只会根据指定列判断是否为重复行,其他列的差异不会影响结果。

内容的提问来源于stack exchange,提问作者brandooo23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 23:30:48