You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Librosa生成的音频相似度矩阵转为0-100匹配分数以识别面试代考

音频相似度计算与分数转换方案

问题背景

识别视频面试中的代考欺诈,通过对比同一受访者多次面试的音频判断是否为同一人发声。目前已用Librosa提取了chroma_cqt和MFCC特征,生成了余弦相似度矩阵,需将其转换为0-100的匹配分数,并设定阈值完成标签标注。

解决方案

1. 相似度矩阵转单一匹配分数

librosa.segment.cross_similarity返回的余弦相似度矩阵sim中,每个元素代表两段音频对应时间片段的余弦相似度(范围**[-1, 1]**)。需先将矩阵归一化到[0,1]区间,再转换为0-100的分数,同时从矩阵中提取代表整体匹配度的单一值:

import numpy as np
import librosa

hop_length = 1024
y_ref, sr1 = librosa.load(r"audio1.wav")
y_comp, sr2 = librosa.load(r"audio2.wav")
chroma_ref = librosa.feature.chroma_cqt(y=y_ref, sr=sr1, hop_length=hop_length)
chroma_comp = librosa.feature.chroma_cqt(y=y_comp, sr=sr2, hop_length=hop_length)

mfcc1 = librosa.feature.mfcc(y_ref, sr1, n_mfcc=13)
mfcc2 = librosa.feature.mfcc(y_comp, sr2, n_mfcc=13)

# 时间延迟嵌入优化特征矩阵
x_ref = librosa.feature.stack_memory(chroma_ref, n_steps=10, delay=3)
x_comp = librosa.feature.stack_memory(chroma_comp, n_steps=10, delay=3)

# 计算余弦相似度矩阵
sim = librosa.segment.cross_similarity(x_comp, x_ref, metric='cosine')

# 将余弦相似度从[-1,1]归一化到[0,1]
normalized_sim = (sim + 1) / 2

# 提取整体匹配分数(推荐用均值平衡全局匹配度,也可选用最大值侧重最相似片段)
match_score = np.mean(normalized_sim) * 100

print(f"音频匹配分数:{match_score:.2f}")

2. 阈值设定与标签标注

结合面试音频场景,需通过标注样本测试确定合适阈值:

  • 收集已知同一人的多组音频对,计算匹配分数后取这类分数的下限作为正样本阈值
  • 收集不同人的音频对,计算分数后取这类分数的上限作为负样本阈值
  • 标注逻辑示例:
    threshold = 70  # 需根据实际样本调整阈值
    label = "同一人" if match_score >= threshold else "代考欺诈"
    print(f"标注结果:{label}")
    

3. 优化建议

  • 融合MFCC特征:当前仅用chroma_cqt特征,可对MFCC做同样的相似度计算,通过加权平均融合两个特征的分数,提升判断准确性:
    # 处理MFCC特征并计算相似度
    mfcc_ref_stack = librosa.feature.stack_memory(mfcc1, n_steps=10, delay=3)
    mfcc_comp_stack = librosa.feature.stack_memory(mfcc2, n_steps=10, delay=3)
    sim_mfcc = librosa.segment.cross_similarity(mfcc_comp_stack, mfcc_ref_stack, metric='cosine')
    normalized_mfcc = (sim_mfcc + 1) / 2
    mfcc_score = np.mean(normalized_mfcc) * 100
    
    # 加权融合分数(示例:chroma占40%,MFCC占60%)
    final_score = 0.4 * match_score + 0.6 * mfcc_score
    
  • 对齐音频长度:若两段音频时长差异较大,先裁剪或对齐到相同长度,避免相似度矩阵出现极端值
  • 样本验证:用更多标注样本测试阈值,确保实际场景中的准确率

内容的提问来源于stack exchange,提问作者The6thSense

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 02:20:37