You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python检测音频重录片段并处理重复内容?

音频重复片段静音处理(保留原时长)

我有大量讲座音频,里面有很多重复的不完整表述,比如:

"this is the part"(随后重录)
"this is the part where"(随后重录)
"this is the part where we will explore the theory"

这类重复片段可以通过波形相似性识别,我需要用Python实现:保留最后一次重录的完整内容,将之前的重复片段静音,且不改变音频总时长。但当前的代码输出和原音频完全一致,没有达到预期效果。

原代码

import librosa
import numpy as np
import soundfile as sf
import os

# 加载音频文件
file_path = r'C:\test.wav'
y, sr = librosa.load(file_path, sr=None)

# 参数设置
chunk_duration = 1  # 秒
overlap = 0.75  # 75%重叠率,提升相似度匹配效果
chunk_length = int(chunk_duration * sr)
step_length = int(chunk_length * (1 - overlap))

# 存储每个片段最后一次出现的结束索引
last_occurrence_end = 0

# 存储最终音频片段的列表
final_audio = []

# 按片段遍历音频
i = 0
while i < len(y) - chunk_length:
    chunk = y[i:i + chunk_length]
    next_chunk_start = i + step_length
    
    # 如果是最后一个片段,直接保留
    if next_chunk_start + chunk_length > len(y):
        last_occurrence_end = len(y)
        break

    next_chunk = y[next_chunk_start:next_chunk_start + chunk_length]
    
    # 计算两个片段的欧氏距离
    distance = np.linalg.norm(chunk - next_chunk)
    
    # 设置相似度阈值
    if distance > 1000:  # 根据需要调整阈值
        if i > last_occurrence_end:
            final_audio.append(y[last_occurrence_end:i])
        last_occurrence_end = i + chunk_length
    
    i = next_chunk_start

# 循环结束后追加最后一段音频
if last_occurrence_end < len(y):
    final_audio.append(y[last_occurrence_end:])

# 检查是否有片段被添加到final_audio
if final_audio:
    # 拼接所有保留的片段,形成最终处理后的音频
    final_audio = np.concatenate(final_audio)
else:
    # 如果没有保留任何片段,返回原音频
    final_audio = y

# 定义带"_clean"后缀的新文件路径
new_file_path = os.path.splitext(file_path)[0] + '_clean.wav'

# 导出处理后的音频
sf.write(new_file_path, final_audio, sr)

print(f"处理后的音频已保存为: {new_file_path}")

原代码问题分析

  1. 逻辑偏离需求:原代码通过拼接不相似片段生成结果,直接改变了音频总时长,和「保留原时长、静音重复片段」的要求完全不符
  2. 相似性判断低效:仅对比相邻重叠片段的欧氏距离,无法识别跨区间的重复序列(比如多次重录的同一句)
  3. 阈值设置不合理:1000的欧氏距离阈值过大,几乎不会触发相似性判定,导致所有片段都被保留

修改后的代码

import librosa
import numpy as np
import soundfile as sf
import os

def detect_and_mute_duplicates(audio, sr, chunk_duration=0.5, similarity_threshold=0.8):
    chunk_length = int(chunk_duration * sr)
    step_length = int(chunk_length * 0.5)  # 50%重叠,平衡检测精度和速度
    
    # 生成所有滑动窗口的片段及对应索引
    chunks = []
    indices = []
    for i in range(0, len(audio) - chunk_length, step_length):
        chunks.append(audio[i:i+chunk_length])
        indices.append((i, i+chunk_length))
    
    # 对片段做归一化,避免音量差异影响相似性判断
    chunks_normalized = []
    for chunk in chunks:
        max_amp = np.max(np.abs(chunk))
        chunks_normalized.append(chunk / max_amp if max_amp != 0 else chunk)
    
    # 计算片段间的余弦相似度矩阵
    similarity_matrix = np.zeros((len(chunks), len(chunks)))
    for i in range(len(chunks)):
        for j in range(i+1, len(chunks)):
            norm_i = np.linalg.norm(chunks_normalized[i])
            norm_j = np.linalg.norm(chunks_normalized[j])
            if norm_i == 0 or norm_j == 0:
                sim = 0.0
            else:
                sim = np.dot(chunks_normalized[i], chunks_normalized[j]) / (norm_i * norm_j)
            similarity_matrix[i][j] = sim
            similarity_matrix[j][i] = sim
    
    # 标记需要静音的区间:仅保留每组相似序列的最后一段
    mute_mask = np.ones_like(audio, dtype=bool)  # True=保留,False=静音
    visited = set()
    
    for i in range(len(chunks)):
        if i in visited:
            continue
        # 找到所有高度相似的片段索引
        similar_indices = np.where(similarity_matrix[i] >= similarity_threshold)[0]
        if len(similar_indices) <= 1:
            continue
        
        # 按结束位置排序,取最后一个作为保留片段
        similar_intervals = [indices[idx] for idx in similar_indices]
        similar_intervals.sort(key=lambda x: x[1])
        keep_interval = similar_intervals[-1]
        
        # 标记除最后一段外的所有相似区间为静音
        for start, end in similar_intervals[:-1]:
            mute_mask[start:end] = False
        visited.update(similar_indices)
    
    # 应用静音掩码,生成处理后的音频
    processed_audio = audio.copy()
    processed_audio[~mute_mask] = 0.0
    return processed_audio

# 主流程
file_path = r'C:\test.wav'
y, sr = librosa.load(file_path, sr=None)

# 处理音频
clean_audio = detect_and_mute_duplicates(y, sr)

# 保存结果
new_file_path = os.path.splitext(file_path)[0] + '_clean.wav'
sf.write(new_file_path, clean_audio, sr)
print(f"处理后的音频已保存为: {new_file_path}")

关键修改说明

  • 核心逻辑调整:通过掩码数组标记静音区域,保证处理后音频总时长和原文件完全一致
  • 相似性判断优化:使用余弦相似度替代欧氏距离,同时对片段做归一化,更精准识别波形相似的重复内容
  • 重复序列处理:找到所有相似片段组,仅保留每组中最后出现的片段,其余全部静音
  • 参数可调:chunk_duration控制检测的片段长度(短重复用小值,长重复用大值),similarity_threshold控制相似判定的严格程度(值越高,仅匹配高度相似的片段)

内容的提问来源于stack exchange,提问作者Joan Venge

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 18:23:11