You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修正Pandas代码实现两列字符串的部分匹配判定?

修正后的代码实现

原代码存在的问题

  1. 未导入re模块,调用re.compile和re.search会抛出NameError。
  2. 匹配逻辑仅单向检查col2是否包含完整的col1字符串,无法覆盖两列存在共同子串的场景(比如第一行的数字部分匹配)。

方案1:通用双向部分匹配

适用于只要两列字符串存在有效重叠子串(或一方是另一方子串)即判定匹配的场景:

import pandas as pd
import re  # 补充导入re模块

data_in = {'col1': ['BANQ1049576495', 'HLCUSEL221162979', 'SEL1469779'],
           'col2': ['KNKX1049576495', 'SEL221162979', 'KROL1020107403']}
df_input = pd.DataFrame(data_in)

def compare_func(row):
    col1_val = row['col1']
    col2_val = row['col2']
    
    # 检查双向子串包含
    if col2_val in col1_val or col1_val in col2_val:
        return 'Match'
    
    # 提取数字部分检查是否完全匹配(覆盖第一行场景)
    num1 = ''.join(re.findall(r'\d+', col1_val))
    num2 = ''.join(re.findall(r'\d+', col2_val))
    if num1 == num2:
        return 'Match'
    
    return 'Mismatch'

df_input['col3'] = df_input.apply(compare_func, axis=1)
print(df_input)

方案2:针对数字部分的精准匹配

如果你的匹配逻辑仅关注字符串中的数字部分是否一致,可使用更简洁的实现:

import pandas as pd
import re

data_in = {'col1': ['BANQ1049576495', 'HLCUSEL221162979', 'SEL1469779'],
           'col2': ['KNKX1049576495', 'SEL221162979', 'KROL1020107403']}
df_input = pd.DataFrame(data_in)

def compare_func(row):
    # 提取并拼接字符串中的所有数字
    extract_num = lambda s: ''.join(re.findall(r'\d+', s))
    return 'Match' if extract_num(row['col1']) == extract_num(row['col2']) else 'Mismatch'

df_input['col3'] = df_input.apply(compare_func, axis=1)
print(df_input)

最终输出结果

两种方案运行后,df_input的col3列均符合期望:

col1               col2      col3
0  BANQ1049576495    KNKX1049576495     Match
1  HLCUSEL221162979    SEL221162979     Match
2       SEL1469779  KROL1020107403  Mismatch

内容的提问来源于stack exchange,提问作者Dew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 04:22:42