You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

音频语义分块技术问询:时间对齐与语义切片实现需求

音频语义分块解决方案与核心技术概念

一、需求实现方案

1. 转写文本与音频的时间对齐

Wav2Vec2-CTC默认仅输出转写文本,要实现时间对齐需从模型的帧级输出入手,具体步骤如下:

实现步骤:

  • 步骤1:建立帧与音频采样点的映射关系
    Wav2Vec2将输入音频编码为固定帧率的特征帧(16000Hz采样率下,每帧对应~20ms音频,对应320个采样点),先定义该映射常量:

    FRAME_STRIDE = 320  # 16000Hz采样率下,单帧对应的音频采样数
    
  • 步骤2:修改推理逻辑,获取token级时间戳
    对模型输出的帧级logits做CTC解码时,跟踪每个token对应的起始/结束帧,再转换为音频时间:

    from transformers import Wav2Vec2CTCTokenizer
    
    tokenizer = Wav2Vec2CTCTokenizer.from_pretrained("facebook/wav2vec2-base-960h")
    
    def transcribe_with_timestamps(input_file, chunk_duration=30):
        speech, sr = librosa.load(input_file, sr=16000, mono=True)
        chunk_samples = int(sr * chunk_duration)
        timestamped_transcript = []
        global_start_sample = 0
    
        for start in range(0, len(speech), chunk_samples):
            end = start + chunk_samples
            speech_chunk = speech[start:end]
            input_values = processor(speech_chunk, return_tensors="pt", sampling_rate=sr).input_values
    
            with torch.no_grad():
                logits = model(input_values).logits
    
            # 提取帧级预测ID
            frame_pred_ids = torch.argmax(logits, dim=-1)[0].numpy()
            # CTC规则去重:移除连续重复ID与空白ID
            filtered_ids = []
            prev_id = None
            blank_id = processor.tokenizer.pad_token_id
            for id in frame_pred_ids:
                if id != blank_id and id != prev_id:
                    filtered_ids.append(id)
                prev_id = id
    
            # 计算每个token对应的帧范围
            token_frame_ranges = []
            current_id = None
            start_frame = 0
            for idx, frame_id in enumerate(frame_pred_ids):
                if frame_id != blank_id:
                    if frame_id != current_id:
                        if current_id is not None:
                            token_frame_ranges.append((current_id, start_frame, idx-1))
                        current_id = frame_id
                        start_frame = idx
            # 处理最后一个token
            if current_id is not None:
                token_frame_ranges.append((current_id, start_frame, len(frame_pred_ids)-1))
    
            # 转换为全局音频时间(秒)
            for token_id, start_f, end_f in token_frame_ranges:
                token = tokenizer.convert_ids_to_tokens(token_id)
                start_sample = global_start_sample + start_f * FRAME_STRIDE
                end_sample = global_start_sample + (end_f + 1) * FRAME_STRIDE
                start_time = round(start_sample / sr, 3)
                end_time = round(end_sample / sr, 3)
                timestamped_transcript.append({
                    "text": token if token != "<s>" else "",
                    "start_time": start_time,
                    "end_time": end_time
                })
            
            global_start_sample += chunk_samples
    
        # 合并连续相同token的时间戳,生成完整句子级对齐结果
        merged = []
        for item in timestamped_transcript:
            if not merged:
                merged.append(item.copy())
            else:
                last = merged[-1]
                if last["text"] == item["text"]:
                    last["end_time"] = item["end_time"]
                else:
                    merged.append(item.copy())
        
        # 按标点分割为完整句子(可选)
        full_sentences = []
        current_sentence = {"text": "", "start_time": None, "end_time": None}
        for item in merged:
            if item["text"] in [".", "?", "!"]:
                current_sentence["text"] += item["text"]
                current_sentence["end_time"] = item["end_time"]
                full_sentences.append(current_sentence.copy())
                current_sentence = {"text": "", "start_time": None, "end_time": None}
            else:
                if current_sentence["start_time"] is None:
                    current_sentence["start_time"] = item["start_time"]
                current_sentence["text"] += item["text"] + " "
        if current_sentence["text"]:
            full_sentences.append(current_sentence)
        return full_sentences
    
  • 步骤3:验证对齐精度
    抽取10-30秒的音频片段,对比转写文本的时间戳与实际音频内容的匹配度,若存在偏差可微调FRAME_STRIDE(部分微调模型可能略有差异)。

2. 结合语义与语音活动的语义分块

目标是生成单段时长<15秒的音频-文本对,具体步骤如下:

实现步骤:

  • 步骤1:语音活动检测(VAD)提取有效语音段
    使用webrtcvad检测音频中的语音区间,过滤静音部分:

    import webrtcvad
    
    def vad_detect(input_file, sr=16000, aggressiveness=3):
        vad = webrtcvad.Vad(aggressiveness)
        speech, _ = librosa.load(input_file, sr=sr, mono=True)
        # 转换为webrtcvad所需的16位PCM格式
        pcm_data = (speech * 32767).astype('int16')
        frame_duration = 30  # 每帧时长(毫秒)
        frame_samples = int(sr * frame_duration / 1000)
        vad_segments = []
        start_sample = 0
        in_speech = False
    
        for i in range(0, len(pcm_data), frame_samples):
            frame = pcm_data[i:i+frame_samples]
            if len(frame) < frame_samples:
                break
            is_speech = vad.is_speech(frame.tobytes(), sr)
            if is_speech and not in_speech:
                in_speech = True
                start_sample = i
            elif not is_speech and in_speech:
                in_speech = False
                vad_segments.append((start_sample, i))
        if in_speech:
            vad_segments.append((start_sample, len(pcm_data)))
        # 转换为时间(秒)
        vad_time_segments = [(s/sr, e/sr) for s,e in vad_segments]
        return vad_time_segments
    
  • 步骤2:基于语义相似度的文本分割
    对带时间戳的转写文本,结合时长限制与语义相似度拆分语义单元:

    from sentence_transformers import SentenceTransformer, util
    
    def semantic_split(timestamped_sentences, max_duration=15):
        model = SentenceTransformer('all-MiniLM-L6-v2')
        sentences = [s["text"].strip() for s in timestamped_sentences]
        sentence_embeddings = model.encode(sentences)
        chunks = []
        current_chunk = {"text": "", "start_time": None, "end_time": None, "duration": 0}
    
        for idx, sent in enumerate(timestamped_sentences):
            sent_duration = sent["end_time"] - sent["start_time"]
            # 若当前chunk加该句子超过15秒,直接拆分
            if current_chunk["duration"] + sent_duration > max_duration:
                chunks.append(current_chunk.copy())
                current_chunk = {
                    "text": sent["text"], 
                    "start_time": sent["start_time"], 
                    "end_time": sent["end_time"], 
                    "duration": sent_duration
                }
            else:
                # 计算语义相似度,低于阈值则拆分
                if current_chunk["text"]:
                    last_emb = sentence_embeddings[idx-1]
                    curr_emb = sentence_embeddings[idx]
                    sim_score = util.cos_sim(last_emb, curr_emb).item()
                    if sim_score < 0.5:
                        chunks.append(current_chunk.copy())
                        current_chunk = {
                            "text": sent["text"], 
                            "start_time": sent["start_time"], 
                            "end_time": sent["end_time"], 
                            "duration": sent_duration
                        }
                    else:
                        current_chunk["text"] += " " + sent["text"]
                        current_chunk["end_time"] = sent["end_time"]
                        current_chunk["duration"] = current_chunk["end_time"] - current_chunk["start_time"]
                else:
                    current_chunk = {
                        "text": sent["text"], 
                        "start_time": sent["start_time"], 
                        "end_time": sent["end_time"], 
                        "duration": sent_duration
                    }
        if current_chunk["text"]:
            chunks.append(current_chunk)
        return chunks
    
  • 步骤3:生成最终音频-文本对
    结合VAD段与语义分块结果,确保每个块仅包含有效语音且时长<15秒:

    def generate_audio_text_pairs(input_file, timestamped_sentences):
        vad_segments = vad_detect(input_file)
        semantic_chunks = semantic_split(timestamped_sentences)
        final_pairs = []
        speech, sr = librosa.load(input_file, sr=16000, mono=True)
    
        for chunk in semantic_chunks:
            chunk_start = chunk["start_time"]
            chunk_end = chunk["end_time"]
            # 对齐到VAD段边界,过滤静音
            for vad_start, vad_end in vad_segments:
                if vad_start <= chunk_start <= vad_end:
                    chunk_start = vad_start
                if vad_start <= chunk_end <= vad_end:
                    chunk_end = vad_end
            # 对超15秒的块进一步拆分
            if chunk_end - chunk_start > 15:
                split_points = [chunk_start + i*15 for i in range(1, int((chunk_end - chunk_start)/15))]
                split_points.append(chunk_end)
                prev_split = chunk_start
                for split in split_points:
                    if split - prev_split > 0.5:  # 过滤过短片段
                        final_pairs.append({
                            "audio": speech[int(prev_split*sr):int(split*sr)],
                            "text": chunk["text"],
                            "start_time": prev_split,
                            "end_time": split
                        })
                    prev_split = split
            else:
                if chunk_end - chunk_start > 0.5:
                    final_pairs.append({
                        "audio": speech[int(chunk_start*sr):int(chunk_end*sr)],
                        "text": chunk["text"],
                        "start_time": chunk_start,
                        "end_time": chunk_end
                    })
        return final_pairs
    

二、核心技术概念

  • CTC对齐:Connectionist Temporal Classification,无需对齐标签的序列建模方法,通过帧级输出推导token与音频的时间映射,是语音文本对齐的核心。
  • 语音活动检测(VAD):识别音频中的语音/静音区间,过滤无效静音,减少后续处理的噪声干扰。
  • 语义文本分割:基于文本语义相似度或语法规则,将长文本拆分为语义独立的单元,确保分块后文本具备完整语义。
  • 语音特征提取:如Wav2Vec2的声学特征编码,将原始音频转换为高维特征,为识别、对齐提供基础。
  • 语义相似度计算:通过预训练语言模型生成文本嵌入,计算文本间的语义距离,判断是否属于同一语义块。

内容的提问来源于stack exchange,提问作者Kartick Narayan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 10:03:14