You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas两个DataFrame按名称、账号双条件匹配填充Category列

问题解决方案

报错原因

  1. 你遇到的TypeError: 'in ' requires string as left operand, not float是因为df_lookup中存在空值(NaN属于float类型),空值直接参与字符串in运算触发类型错误。
  2. 原有修改逻辑存在缺陷:新增的Account完全匹配条件需要同时获取df1每行的Account字段,仅传入Name字段无法完成账号校验,不需要额外构造Combined字段。

修复后代码

import pandas as pd

def func(row):
    global df_lookup
    # 统一转字符串避免类型不兼容问题
    current_name = str(row['Name'])
    current_account = row['Account']
    
    for _, lookup_row in df_lookup.iterrows():
        lookup_name = lookup_row['Name']
        lookup_account = lookup_row['Account']
        # 跳过空值行,避免in运算报错
        if pd.isna(lookup_name) or pd.isna(lookup_account):
            continue
        # 先校验账号完全匹配,再校验名称模糊匹配
        if lookup_account == current_account and str(lookup_name) in current_name:
            return lookup_row['Category']
    
    # 未匹配到则将当前行的名称、账号同步追加到lookup表
    new_row = pd.DataFrame([{
        "Name": current_name,
        "Account": current_account,
        "Category": "Needs Category"
    }])
    # 用concat替代已弃用的append方法,适配pandas全版本
    df_lookup = pd.concat([df_lookup, new_row], ignore_index=True)
    return "Needs Category"

# 逐行处理df1的全字段,生成Category列
df1['Category'] = df1.apply(func, axis=1)

逻辑说明

  • 调用apply时指定axis=1逐行读取df1的全部字段,同时获取每行的Name和Account值用于匹配
  • 匹配优先级:先校验账号完全相等,再校验lookup的Name值是否包含在df1当前行的Name中,两个条件同时满足才返回对应Category
  • 空值兼容处理:所有参与in运算的值统一转字符串,空值行直接跳过,彻底解决类型报错问题
  • 未匹配的条目会同时将Name和Account追加到lookup表,保证后续同账号同名称的条目可以命中匹配

如果数据量较大(万行以上),可以提前将lookup表按Account分组存储关键词映射,减少遍历次数提升运行效率。

内容的提问来源于stack exchange,提问作者user8200199

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 10:12:03