如何更智能检测Spotify上相似名歌曲(含代码优化需求)
更智能的Spotify歌曲重复检测方案?
我现在用下面的代码检测Spotify上两首歌是否为同一首(比如存在拼写错误、现场版这类不同版本,名称相似的情况)。除了检查共同前缀或者用Levenshtein距离之外,有没有更智能的方案?当前代码偶尔会漏判重复项,还得手动删除。
def similarity_old(s1, s2): if len(s1) != len(s2): return False count = 0 for c1, c2 in zip(s1, s2): if c1 == c2: count += 1 similarity_percentage = (count / len(s1)) * 100 return similarity_percentage > 70 def similarity_old2(s1, s2): # implements levenshtein distance m = len(s1) n = len(s2) if abs(m - n) > max(m, n) * 0.9: # Allowing up to 90% length difference return False if s1 == s2: return True if m == 0 or n == 0: return False # Create a matrix to store the edit distances dp = [[0] * (n + 1) for _ in range(m + 1)] # Initialize the first row and column of the matrix for i in range(m + 1): dp[i][0] = i for j in range(n + 1): dp[0][j] = j # Compute the edit distances for i in range(1, m + 1): for j in range(1, n + 1): cost = 0 if s1[i - 1] == s2[j - 1] else 1 dp[i][j] = min(dp[i - 1][j] + 1, # Deletion dp[i][j - 1] + 1, # Insertion dp[i - 1][j - 1] + cost) # Substitution # Calculate the similarity percentage similarity_percentage = ((max(m, n) - dp[m][n]) / max(m, n)) * 100 return similarity_percentage > 70 def similarity(s1, s2): # optimized levenshtein distance len1 = len(s1) len2 = len(s2) if abs(len1 - len2) > max(len1, len2) * 0.9: # Allowing up to 90% length difference return False if s1 == s2: return True if len1 == 0 or len2 == 0: return False if len1 > len2: s1, s2 = s2, s1 len1, len2 = len2, len1 previous_row = list(range(len1 + 1)) for i, c2 in enumerate(s2): current_row = [i + 1] for j, c1 in enumerate(s1): insertions = previous_row[j + 1] + 1 deletions = current_row[j] + 1 substitutions = previous_row[j] + (c1 != c2) current_row.append(min(insertions, deletions, substitutions)) previous_row = current_row similarity_percentage = ((len1 - previous_row[-1]) / len1) * 100 return similarity_percentage > 70 def common_beginning(s1, s2): if len(s1) > 11 and s2.startswith(s1[:11]): return True else: return False def check_similarity(s, list1): # s = song name # list1 = list of song names # returns True if s is similar to something in list1, else False for item in list1: if common_beginning(s, item): return True #for item in list1: # if similarity(s, item): # return True return False
可行的优化方案
1. 先做针对性的文本预处理
这是提升匹配准确率的基础,能消除版本标识、格式差异带来的干扰:
- 统一大小写:将所有歌名转为小写,避免
Hello和hello被误判为不同 - 移除版本后缀/标记:正则匹配并移除
(Live)、[Acoustic Remix]、Radio Edit这类常见版本标识 - 清理冗余字符:去掉标点、特殊符号,合并多余空格,比如把
Don't Stop! Believin'转为dont stop believin - 过滤无意义冠词:移除
the、a、an这类不影响核心语义的词
2. 替换为更适配的字符串匹配算法
Levenshtein距离适合拼写错误,但对版本后缀这类大段附加内容的处理效果有限,可以试试这些:
- Jaccard相似度:基于n-gram(比如2-gram、3-gram)计算交集占并集的比例,能更好捕捉文本的整体相似性,适合处理有插入删除的场景
- Damerau-Levenshtein距离:在Levenshtein基础上支持相邻字符交换(比如
teh和the),更贴合实际拼写错误情况 - 部分匹配相似度:比如只匹配两个字符串中较短的那个的全部内容,适合处理原歌名和带长后缀的版本(比如
Hello和Hello (Live in Paris))
3. 结合Spotify元数据辅助判断
如果能获取歌曲的额外信息,匹配准确率会大幅提升:
- 歌手名称:同一首歌的不同版本通常由同一歌手演唱
- 时长:同一首歌的不同版本时长差异不会过大(比如控制在±10%范围内)
- 专辑关联:可以通过Spotify API查询歌曲的原版关联,直接判断是否为同一首歌的衍生版本
4. 使用成熟的模糊匹配库
不用自己造轮子,现成的库已经封装了优化后的算法:
fuzzywuzzy/rapidfuzz:提供多种相似度计算方法(比如partial_ratio处理长短文本、token_sort_ratio处理词序不同的情况),调用简单且准确率高- 示例代码(基于
rapidfuzz):
import re from rapidfuzz import fuzz def preprocess(name): # 预处理逻辑 name = name.lower() name = re.sub(r'\([^)]*\)|\[[^\]]*\]', '', name) name = re.sub(r'[^\w\s]', '', name) name = re.sub(r'\s+', ' ', name).strip() return name def is_duplicate(song_a, song_b): processed_a = preprocess(song_a) processed_b = preprocess(song_b) # 用partial_ratio处理长短文本,阈值可根据需求调整 return fuzz.partial_ratio(processed_a, processed_b) > 85
5. 进阶:语义匹配
如果处理大量数据或复杂改写场景,可以用预训练文本嵌入模型(比如Sentence-BERT),将歌名转为向量后计算余弦相似度,能识别语义相近但表述不同的情况。
内容的提问来源于stack exchange,提问作者xralf
相关产品推荐
相关产品推荐

