You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比同一DataFrame两列字符串并计算匹配率及匹配类型

Pandas实现DataFrame两列字符串逐行对比

实现效果

针对同一DataFrame的两列字符串逐行计算,新增两列输出结果:

  • 两字符串的匹配百分比
  • 匹配类型判定,共三类:完全匹配、部分匹配、完全不匹配
    实现效果参考:
    示例截图

实现代码

直接用Python标准库difflib即可完成相似度计算,无需额外安装第三方依赖:

import pandas as pd
from difflib import SequenceMatcher

def string_match_calc(str1, str2):
    # 空值兜底处理
    if pd.isna(str1) or pd.isna(str2):
        return 0.0, "完全不匹配"
    # 统一转字符串、去首尾空格后计算相似度
    s1 = str(str1).strip()
    s2 = str(str2).strip()
    sim_ratio = SequenceMatcher(None, s1, s2).ratio()
    match_pct = round(sim_ratio * 100, 2)
    # 匹配类型判定
    if match_pct == 100:
        match_type = "完全匹配"
    elif match_pct == 0:
        match_type = "完全不匹配"
    else:
        match_type = "部分匹配"
    return match_pct, match_type

# ------------------- 调用示例 -------------------
# 替换成你自己的DataFrame和对应列名即可
df = pd.DataFrame({
    "字符串列1": ["apple", "banana", "orange", "grape", "cherry"],
    "字符串列2": ["apple", "bananas", "pear", "grape", "berry"]
})

# 逐行计算生成结果列
df[["匹配百分比(%)", "匹配类型"]] = df.apply(
    lambda x: pd.Series(string_match_calc(x["字符串列1"], x["字符串列2"])),
    axis=1
)

自定义调整说明

  • 阈值调整:如果需要自定义完全不匹配的判定标准(比如相似度低于20%才算完全不匹配),直接修改判定分支的数值条件即可
  • 性能优化:如果是百万行以上的大表,可替换difflib为rapidfuzz库的fuzz.ratio接口,计算逻辑完全一致,速度可提升数十倍
  • 适配中文:上述逻辑对中文字符串同样生效,无需额外修改编码或匹配规则

运行结果示例

字符串列1字符串列2匹配百分比(%)匹配类型
appleapple100.0完全匹配
bananabananas92.31部分匹配
orangepear22.22部分匹配
grapegrape100.0完全匹配
cherryberry66.67部分匹配

内容的提问来源于stack exchange,提问作者Ravindra pol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 22:54:26