You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python开源工具定位语音替换音频的突变时间戳?

定位语音替换融合点的方法

场景术语说明

你遇到的场景属于音频篡改检测下的拼接边界定位,具体是语音片段替换后的拼接点检测,也可称为语音替换篡改边界检测。

方法一:传统信号处理方案(适配波形突变特征)

1. 改进的自适应阈值差分法

针对固定阈值误报多的问题,改用滑动窗口统计量生成自适应阈值,再对候选点去噪:

import numpy as np
import librosa

# 加载音频
y, sr = librosa.load("example.wav", sr=None)
# 计算一阶差分的绝对值
dy = np.abs(np.diff(y))

# 10ms滑动窗口计算自适应阈值(适配语音帧长度)
window_size = int(sr * 0.01)
thresholds = []
for i in range(len(dy) - window_size + 1):
    window = dy[i:i+window_size]
    # 用均值+3倍标准差作为当前窗口的突变阈值
    thresholds.append(np.mean(window) + 3 * np.std(window))
# 补全末尾窗口的阈值
thresholds += [thresholds[-1]] * (len(dy) - len(thresholds))

# 筛选超过阈值的候选突变点
candidates = np.where(dy > thresholds)[0]

# 合并连续候选点,取聚类中心作为拼接时间戳
if len(candidates) > 0:
    clusters = []
    current_cluster = [candidates[0]]
    for idx in candidates[1:]:
        # 时间差小于1个窗口的归为同一类
        if idx - current_cluster[-1] < window_size:
            current_cluster.append(idx)
        else:
            clusters.append(current_cluster)
            current_cluster = [idx]
    clusters.append(current_cluster)
    # 转换为秒级时间戳
    timestamps = [(np.mean(cluster) + 0.5) / sr for cluster in clusters]
    print("检测到的拼接点时间戳:", timestamps)

2. 多特征融合检测

结合HPSS分离结果,同时检测谐波、打击乐分量的突变,减少误报:

import numpy as np
import librosa

y, sr = librosa.load("example.wav", sr=None)
# 分离谐波与打击乐分量
y_harm, y_perc = librosa.effects.hpss(y)

# 分别计算三个分量的差分绝对值
dy_full = np.abs(np.diff(y))
dy_harm = np.abs(np.diff(y_harm))
dy_perc = np.abs(np.diff(y_perc))

# 计算各分量的自适应阈值
def get_adaptive_threshold(signal_diff, window_size):
    thresholds = []
    for i in range(len(signal_diff) - window_size + 1):
        window = signal_diff[i:i+window_size]
        thresholds.append(np.mean(window) + 3 * np.std(window))
    thresholds += [thresholds[-1]] * (len(signal_diff) - len(thresholds))
    return thresholds

window_size = int(sr * 0.01)
thresh_full = get_adaptive_threshold(dy_full, window_size)
thresh_harm = get_adaptive_threshold(dy_harm, window_size)
thresh_perc = get_adaptive_threshold(dy_perc, window_size)

# 仅保留三个分量都超过阈值的点
joint_candidates = np.where(
    (dy_full > thresh_full) & (dy_harm > thresh_harm) & (dy_perc > thresh_perc)
)[0]

# 聚类生成时间戳(同前一个方法的聚类逻辑)
if len(joint_candidates) > 0:
    clusters = []
    current_cluster = [joint_candidates[0]]
    for idx in joint_candidates[1:]:
        if idx - current_cluster[-1] < window_size:
            current_cluster.append(idx)
        else:
            clusters.append(current_cluster)
            current_cluster = [idx]
    clusters.append(current_cluster)
    timestamps = [(np.mean(cluster) + 0.5) / sr for cluster in clusters]
    print("多特征融合检测到的拼接点:", timestamps)

3. MFCC帧差异检测

利用语音的频谱特征差异定位拼接点,对音色相近的说话人也有效:

import numpy as np
import librosa

y, sr = librosa.load("example.wav", sr=None)
# 提取13维MFCC特征
mfcc = librosa.feature.mfcc(y=y, sr=sr, n_mfcc=13)
# 计算相邻帧的MFCC欧氏距离
mfcc_diff = np.linalg.norm(mfcc[:, 1:] - mfcc[:, :-1], axis=0)

# 自适应阈值筛选
mfcc_threshold = np.mean(mfcc_diff) + 2 * np.std(mfcc_diff)
candidate_frames = np.where(mfcc_diff > mfcc_threshold)[0]

# 转换为秒级时间戳(默认hop_length=512)
hop_length = 512
timestamps = (candidate_frames + 0.5) * hop_length / sr
print("MFCC检测到的拼接点:", timestamps)

方法二:机器学习/深度学习方案

如果传统方法仍有误报,可尝试以下方案:

  • 预训练模型复用:使用pyannote.audio的说话人分割模型,拼接点的特征变化与说话人切换模式相似,可直接用于边界检测
  • 自定义模型训练:用公开音频篡改数据集(如ASVspoof 2021)训练CNN/Transformer模型,学习拼接点的特征模式,实现精准检测

工具推荐

  • librosa:核心音频处理工具,支持特征提取、信号变换
  • pyannote.audio:提供预训练的音频分割、说话人识别模型
  • numpy:用于信号的数值计算与处理

内容的提问来源于stack exchange,提问作者Long

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:55:17