如何用SequenceMatcher筛选符合相似度比例的字符串及DataFrame示例
没问题,我来一步步教你怎么用Python的difflib.SequenceMatcher实现这个需求——把单个字符串和DataFrame某列的所有字符串比对,只留下相似度比例超过0.8的结果。直接上干货!
核心步骤说明
首先,我们需要用到两个库:difflib(提供SequenceMatcher)和pandas(处理DataFrame)。核心逻辑是:
- 定义一个判断相似度的函数,计算两个字符串的匹配比例
- 遍历DataFrame的目标列,用这个函数筛选符合阈值的行/字符串
完整代码示例
1. 导入依赖库
from difflib import SequenceMatcher import pandas as pd
2. 定义相似度判断函数
这个函数会接收目标字符串和待比对字符串,返回是否满足相似度>0.8的条件:
def meets_similarity_threshold(target, compare_str): # 计算相似度比例:SequenceMatcher的ratio()返回0-1之间的数值,1表示完全匹配 similarity = SequenceMatcher(None, target, compare_str).ratio() return similarity > 0.8
3. 创建测试数据并筛选
我们先构造一个测试用的DataFrame,然后执行筛选:
# 测试DataFrame,模拟你实际的数据 df = pd.DataFrame({ 'product_name': [ "stainless steel water bottle", "steel water bottle", "glass coffee mug", "stainless steel bottle", "plastic water jug", "stainless steel water bottle with lid" ] }) # 我们要比对的目标字符串 target_string = "stainless steel water bottle"
方式1:筛选出DataFrame中符合条件的行
用apply生成布尔索引,直接过滤DataFrame:
# 生成布尔索引,筛选符合条件的行 filtered_df = df[df['product_name'].apply(lambda x: meets_similarity_threshold(target_string, x))] print("筛选后的结果:") print(filtered_df)
运行结果:
筛选后的结果:
product_name
0 stainless steel water bottle
3 stainless steel bottle
5 stainless steel water bottle with lid
方式2:直接提取符合条件的字符串列表
如果只需要字符串结果,用列表推导式更高效:
matched_strings = [s for s in df['product_name'] if meets_similarity_threshold(target_string, s)] print("符合条件的字符串列表:") print(matched_strings)
运行结果:
符合条件的字符串列表:
['stainless steel water bottle', 'stainless steel bottle', 'stainless steel water bottle with lid']
可选优化:预处理字符串
默认的SequenceMatcher会严格匹配字符(包括大小写、空格)。如果需要忽略大小写或空格,可以在函数里加预处理步骤:
def meets_similarity_threshold_case_insensitive(target, compare_str): # 预处理:转小写+移除所有空格 clean_target = target.lower().replace(" ", "") clean_compare = compare_str.lower().replace(" ", "") similarity = SequenceMatcher(None, clean_target, clean_compare).ratio() return similarity > 0.8
这样比如"Stainless Steel Bottle"和"stainlesssteelwaterbottle"也能被匹配到~
内容的提问来源于stack exchange,提问作者Rahul Syal

