You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用SequenceMatcher筛选符合相似度比例的字符串及DataFrame示例

使用SequenceMatcher筛选高相似度字符串(匹配比例>0.8)

没问题,我来一步步教你怎么用Python的difflib.SequenceMatcher实现这个需求——把单个字符串和DataFrame某列的所有字符串比对,只留下相似度比例超过0.8的结果。直接上干货!

核心步骤说明

首先,我们需要用到两个库:difflib(提供SequenceMatcher)和pandas(处理DataFrame)。核心逻辑是:

  • 定义一个判断相似度的函数,计算两个字符串的匹配比例
  • 遍历DataFrame的目标列,用这个函数筛选符合阈值的行/字符串

完整代码示例

1. 导入依赖库

from difflib import SequenceMatcher
import pandas as pd

2. 定义相似度判断函数

这个函数会接收目标字符串和待比对字符串,返回是否满足相似度>0.8的条件:

def meets_similarity_threshold(target, compare_str):
    # 计算相似度比例:SequenceMatcher的ratio()返回0-1之间的数值,1表示完全匹配
    similarity = SequenceMatcher(None, target, compare_str).ratio()
    return similarity > 0.8

3. 创建测试数据并筛选

我们先构造一个测试用的DataFrame,然后执行筛选:

# 测试DataFrame,模拟你实际的数据
df = pd.DataFrame({
    'product_name': [
        "stainless steel water bottle",
        "steel water bottle",
        "glass coffee mug",
        "stainless steel bottle",
        "plastic water jug",
        "stainless steel water bottle with lid"
    ]
})

# 我们要比对的目标字符串
target_string = "stainless steel water bottle"

方式1:筛选出DataFrame中符合条件的行

用apply生成布尔索引,直接过滤DataFrame:

# 生成布尔索引,筛选符合条件的行
filtered_df = df[df['product_name'].apply(lambda x: meets_similarity_threshold(target_string, x))]

print("筛选后的结果:")
print(filtered_df)

运行结果:

筛选后的结果:
product_name
0 stainless steel water bottle
3 stainless steel bottle
5 stainless steel water bottle with lid

方式2:直接提取符合条件的字符串列表

如果只需要字符串结果,用列表推导式更高效:

matched_strings = [s for s in df['product_name'] if meets_similarity_threshold(target_string, s)]

print("符合条件的字符串列表:")
print(matched_strings)

运行结果:

符合条件的字符串列表:
['stainless steel water bottle', 'stainless steel bottle', 'stainless steel water bottle with lid']

可选优化:预处理字符串

默认的SequenceMatcher会严格匹配字符(包括大小写、空格)。如果需要忽略大小写或空格,可以在函数里加预处理步骤:

def meets_similarity_threshold_case_insensitive(target, compare_str):
    # 预处理:转小写+移除所有空格
    clean_target = target.lower().replace(" ", "")
    clean_compare = compare_str.lower().replace(" ", "")
    similarity = SequenceMatcher(None, clean_target, clean_compare).ratio()
    return similarity > 0.8

这样比如"Stainless Steel Bottle"和"stainlesssteelwaterbottle"也能被匹配到~

内容的提问来源于stack exchange,提问作者Rahul Syal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:02:42