You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更智能检测Spotify上相似名歌曲(含代码优化需求)

更智能的Spotify歌曲重复检测方案?

我现在用下面的代码检测Spotify上两首歌是否为同一首(比如存在拼写错误、现场版这类不同版本,名称相似的情况)。除了检查共同前缀或者用Levenshtein距离之外,有没有更智能的方案?当前代码偶尔会漏判重复项,还得手动删除。

def similarity_old(s1, s2):
    if len(s1) != len(s2):
        return False

    count = 0
    for c1, c2 in zip(s1, s2):
        if c1 == c2:
            count += 1

    similarity_percentage = (count / len(s1)) * 100
    return similarity_percentage > 70


def similarity_old2(s1, s2):
    # implements levenshtein distance
    m = len(s1)
    n = len(s2)
    if abs(m - n) > max(m, n) * 0.9:  # Allowing up to 90% length difference
        return False

    if s1 == s2:
        return True

    if m == 0 or n == 0:
        return False

    # Create a matrix to store the edit distances
    dp = [[0] * (n + 1) for _ in range(m + 1)]

    # Initialize the first row and column of the matrix
    for i in range(m + 1):
        dp[i][0] = i
    for j in range(n + 1):
        dp[0][j] = j

    # Compute the edit distances
    for i in range(1, m + 1):
        for j in range(1, n + 1):
            cost = 0 if s1[i - 1] == s2[j - 1] else 1
            dp[i][j] = min(dp[i - 1][j] + 1,  # Deletion
                           dp[i][j - 1] + 1,  # Insertion
                           dp[i - 1][j - 1] + cost)  # Substitution

    # Calculate the similarity percentage
    similarity_percentage = ((max(m, n) - dp[m][n]) / max(m, n)) * 100

    return similarity_percentage > 70

def similarity(s1, s2):
    # optimized levenshtein distance
    len1 = len(s1)
    len2 = len(s2)
    if abs(len1 - len2) > max(len1, len2) * 0.9:  # Allowing up to 90% length difference
        return False

    if s1 == s2:
        return True

    if len1 == 0 or len2 == 0:
        return False

    if len1 > len2:
        s1, s2 = s2, s1
        len1, len2 = len2, len1

    previous_row = list(range(len1 + 1))

    for i, c2 in enumerate(s2):
        current_row = [i + 1]
        for j, c1 in enumerate(s1):
            insertions = previous_row[j + 1] + 1
            deletions = current_row[j] + 1
            substitutions = previous_row[j] + (c1 != c2)
            current_row.append(min(insertions, deletions, substitutions))
        previous_row = current_row

    similarity_percentage = ((len1 - previous_row[-1]) / len1) * 100

    return similarity_percentage > 70


def common_beginning(s1, s2):
    if len(s1) > 11 and s2.startswith(s1[:11]):
        return True
    else:
        return False


def check_similarity(s, list1):
    # s = song name
    # list1 = list of song names
    # returns True if s is similar to something in list1, else False

    for item in list1:
        if common_beginning(s, item):
            return True
    #for item in list1:
    #    if similarity(s, item):
    #        return True
    return False

可行的优化方案

1. 先做针对性的文本预处理

这是提升匹配准确率的基础,能消除版本标识、格式差异带来的干扰:

  • 统一大小写:将所有歌名转为小写,避免Hello和hello被误判为不同
  • 移除版本后缀/标记:正则匹配并移除(Live)、[Acoustic Remix]、Radio Edit这类常见版本标识
  • 清理冗余字符:去掉标点、特殊符号,合并多余空格,比如把Don't Stop! Believin'转为dont stop believin
  • 过滤无意义冠词:移除the、a、an这类不影响核心语义的词

2. 替换为更适配的字符串匹配算法

Levenshtein距离适合拼写错误,但对版本后缀这类大段附加内容的处理效果有限,可以试试这些:

  • Jaccard相似度:基于n-gram(比如2-gram、3-gram)计算交集占并集的比例,能更好捕捉文本的整体相似性,适合处理有插入删除的场景
  • Damerau-Levenshtein距离:在Levenshtein基础上支持相邻字符交换(比如teh和the),更贴合实际拼写错误情况
  • 部分匹配相似度:比如只匹配两个字符串中较短的那个的全部内容,适合处理原歌名和带长后缀的版本(比如Hello和Hello (Live in Paris))

3. 结合Spotify元数据辅助判断

如果能获取歌曲的额外信息,匹配准确率会大幅提升:

  • 歌手名称:同一首歌的不同版本通常由同一歌手演唱
  • 时长:同一首歌的不同版本时长差异不会过大(比如控制在±10%范围内)
  • 专辑关联:可以通过Spotify API查询歌曲的原版关联,直接判断是否为同一首歌的衍生版本

4. 使用成熟的模糊匹配库

不用自己造轮子,现成的库已经封装了优化后的算法:

  • fuzzywuzzy/rapidfuzz:提供多种相似度计算方法(比如partial_ratio处理长短文本、token_sort_ratio处理词序不同的情况),调用简单且准确率高
  • 示例代码(基于rapidfuzz):
import re
from rapidfuzz import fuzz

def preprocess(name):
    # 预处理逻辑
    name = name.lower()
    name = re.sub(r'\([^)]*\)|\[[^\]]*\]', '', name)
    name = re.sub(r'[^\w\s]', '', name)
    name = re.sub(r'\s+', ' ', name).strip()
    return name

def is_duplicate(song_a, song_b):
    processed_a = preprocess(song_a)
    processed_b = preprocess(song_b)
    # 用partial_ratio处理长短文本,阈值可根据需求调整
    return fuzz.partial_ratio(processed_a, processed_b) > 85

5. 进阶:语义匹配

如果处理大量数据或复杂改写场景,可以用预训练文本嵌入模型(比如Sentence-BERT),将歌名转为向量后计算余弦相似度,能识别语义相近但表述不同的情况。

内容的提问来源于stack exchange,提问作者xralf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 18:24:53