You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用DataSketch结合MFCC判断3段音频相似度的异常问题排查

问题描述

我用DataSketch库判断音频2、3是否和音频1相似,哪怕把阈值设为1(本该只匹配完全相同的音频),结果还是把和音频1差异极大的两段都列出来了。所有音频时长都是29秒,但内容完全不同。

错误代码
from datasketch import MinHash , MinHashLSH

x1 , Sr1 = librosa.load(r'path\f1.mp3')
mfcc1 = librosa.feature.mfcc(y=x1 , sr=Sr1)
mfcc1 = mfcc1.tobytes()

x2 , Sr2 = librosa.load(r'path\f2.mp3')
mfcc2 = librosa.feature.mfcc(y=x2 , sr=Sr2)
mfcc2 = mfcc2.tobytes()

x3 , Sr3 = librosa.load(r'path\f3.mp3')
mfcc3 = librosa.feature.mfcc(y=x3 , sr=Sr3)
mfcc3 = mfcc3.tobytes()

minhash1 = MinHash(num_perm=128 , hashfunc=hash)
minhash2 = MinHash(num_perm=128 , hashfunc=hash)
minhash3 = MinHash(num_perm=128 , hashfunc=hash)

for col1 in mfcc1:
    minhash1.update(col1)

for col2 in mfcc2:
    minhash2.update(col2)

for col3 in mfcc3:
    minhash3.update(col3)

lsh = MinHashLSH(threshold= 1 , num_perm=128)
lsh.insert("minhash2",minhash2)
lsh.insert("minhash3",minhash3)
result=lsh.query(minhash1)
print(result)
问题原因及修复方案

核心问题

当前处理逻辑完全错误:

  1. MFCC特征处理逻辑错误:直接将MFCC特征矩阵转成bytes后遍历单个字节更新MinHash,相当于把音频特征拆成了无意义的单个字节元素,完全丢失了MFCC的语义信息。MinHash需要的是有意义的集合元素,而非字节流。
  2. MinHash使用场景误解:MinHash用于计算集合相似度,必须先将MFCC特征转换为合理的集合元素,而非粗暴转成字节遍历。

修复步骤

方案1:将MFCC分帧特征作为集合元素

MFCC的每一列对应一帧音频特征,可将每帧MFCC向量转换为可哈希的元组,作为MinHash的更新元素:

from datasketch import MinHash, MinHashLSH
import librosa
import numpy as np

def get_mfcc_items(audio_path):
    y, sr = librosa.load(audio_path)
    mfcc = librosa.feature.mfcc(y=y, sr=sr)
    # 将每帧MFCC转为元组(可哈希类型),作为集合元素
    return [tuple(frame) for frame in mfcc.T]

# 获取各音频的MFCC集合元素
items1 = get_mfcc_items(r'path\f1.mp3')
items2 = get_mfcc_items(r'path\f2.mp3')
items3 = get_mfcc_items(r'path\f3.mp3')

# 初始化MinHash并更新
minhash1 = MinHash(num_perm=128)
for item in items1:
    minhash1.update(np.array(item).tobytes())  # 用数组字节更新,保证哈希一致性

minhash2 = MinHash(num_perm=128)
for item in items2:
    minhash2.update(np.array(item).tobytes())

minhash3 = MinHash(num_perm=128)
for item in items3:
    minhash3.update(np.array(item).tobytes())

# 构建LSH并查询
lsh = MinHashLSH(threshold=1.0, num_perm=128)
lsh.insert("audio2", minhash2)
lsh.insert("audio3", minhash3)
result = lsh.query(minhash1)
print(result)

方案2:特征量化(更适配音频相似度场景)

若追求更好效果,可先对MFCC特征进行量化(如聚类为离散符号),再用MinHash处理:

from datasketch import MinHash, MinHashLSH
import librosa
import numpy as np
from sklearn.cluster import KMeans

def quantize_mfcc(audio_path, n_clusters=100):
    y, sr = librosa.load(audio_path)
    mfcc = librosa.feature.mfcc(y=y, sr=sr).T
    # 用KMeans量化MFCC帧为离散标签
    kmeans = KMeans(n_clusters=n_clusters, random_state=42).fit(mfcc)
    return kmeans.labels_

# 获取量化后的特征标签(作为集合元素)
labels1 = quantize_mfcc(r'path\f1.mp3')
labels2 = quantize_mfcc(r'path\f2.mp3')
labels3 = quantize_mfcc(r'path\f3.mp3')

# 初始化MinHash
minhash1 = MinHash(num_perm=128)
for label in labels1:
    minhash1.update(str(label).encode('utf-8'))

minhash2 = MinHash(num_perm=128)
for label in labels2:
    minhash2.update(str(label).encode('utf-8'))

minhash3 = MinHash(num_perm=128)
for label in labels3:
    minhash3.update(str(label).encode('utf-8'))

# 查询
lsh = MinHashLSH(threshold=1.0, num_perm=128)
lsh.insert("audio2", minhash2)
lsh.insert("audio3", minhash3)
result = lsh.query(minhash1)
print(result)

额外说明

  • 阈值设为1.0时,仅当两个MinHash的所有置换哈希值完全相同时才会匹配,对应集合的Jaccard相似度为1(完全相同)。
  • 音频相似度判断也可直接计算MFCC的余弦相似度,或使用专门的音频指纹库(如dejavu),效果更稳定。

内容的提问来源于stack exchange,提问作者Faizan Ul Haq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 15:40:38