音频语义分块技术问询:时间对齐与语义切片实现需求
音频语义分块解决方案与核心技术概念
一、需求实现方案
1. 转写文本与音频的时间对齐
Wav2Vec2-CTC默认仅输出转写文本,要实现时间对齐需从模型的帧级输出入手,具体步骤如下:
实现步骤:
步骤1:建立帧与音频采样点的映射关系
Wav2Vec2将输入音频编码为固定帧率的特征帧(16000Hz采样率下,每帧对应~20ms音频,对应320个采样点),先定义该映射常量:FRAME_STRIDE = 320 # 16000Hz采样率下,单帧对应的音频采样数步骤2:修改推理逻辑,获取token级时间戳
对模型输出的帧级logits做CTC解码时,跟踪每个token对应的起始/结束帧,再转换为音频时间:from transformers import Wav2Vec2CTCTokenizer tokenizer = Wav2Vec2CTCTokenizer.from_pretrained("facebook/wav2vec2-base-960h") def transcribe_with_timestamps(input_file, chunk_duration=30): speech, sr = librosa.load(input_file, sr=16000, mono=True) chunk_samples = int(sr * chunk_duration) timestamped_transcript = [] global_start_sample = 0 for start in range(0, len(speech), chunk_samples): end = start + chunk_samples speech_chunk = speech[start:end] input_values = processor(speech_chunk, return_tensors="pt", sampling_rate=sr).input_values with torch.no_grad(): logits = model(input_values).logits # 提取帧级预测ID frame_pred_ids = torch.argmax(logits, dim=-1)[0].numpy() # CTC规则去重:移除连续重复ID与空白ID filtered_ids = [] prev_id = None blank_id = processor.tokenizer.pad_token_id for id in frame_pred_ids: if id != blank_id and id != prev_id: filtered_ids.append(id) prev_id = id # 计算每个token对应的帧范围 token_frame_ranges = [] current_id = None start_frame = 0 for idx, frame_id in enumerate(frame_pred_ids): if frame_id != blank_id: if frame_id != current_id: if current_id is not None: token_frame_ranges.append((current_id, start_frame, idx-1)) current_id = frame_id start_frame = idx # 处理最后一个token if current_id is not None: token_frame_ranges.append((current_id, start_frame, len(frame_pred_ids)-1)) # 转换为全局音频时间(秒) for token_id, start_f, end_f in token_frame_ranges: token = tokenizer.convert_ids_to_tokens(token_id) start_sample = global_start_sample + start_f * FRAME_STRIDE end_sample = global_start_sample + (end_f + 1) * FRAME_STRIDE start_time = round(start_sample / sr, 3) end_time = round(end_sample / sr, 3) timestamped_transcript.append({ "text": token if token != "<s>" else "", "start_time": start_time, "end_time": end_time }) global_start_sample += chunk_samples # 合并连续相同token的时间戳,生成完整句子级对齐结果 merged = [] for item in timestamped_transcript: if not merged: merged.append(item.copy()) else: last = merged[-1] if last["text"] == item["text"]: last["end_time"] = item["end_time"] else: merged.append(item.copy()) # 按标点分割为完整句子(可选) full_sentences = [] current_sentence = {"text": "", "start_time": None, "end_time": None} for item in merged: if item["text"] in [".", "?", "!"]: current_sentence["text"] += item["text"] current_sentence["end_time"] = item["end_time"] full_sentences.append(current_sentence.copy()) current_sentence = {"text": "", "start_time": None, "end_time": None} else: if current_sentence["start_time"] is None: current_sentence["start_time"] = item["start_time"] current_sentence["text"] += item["text"] + " " if current_sentence["text"]: full_sentences.append(current_sentence) return full_sentences步骤3:验证对齐精度
抽取10-30秒的音频片段,对比转写文本的时间戳与实际音频内容的匹配度,若存在偏差可微调FRAME_STRIDE(部分微调模型可能略有差异)。
2. 结合语义与语音活动的语义分块
目标是生成单段时长<15秒的音频-文本对,具体步骤如下:
实现步骤:
步骤1:语音活动检测(VAD)提取有效语音段
使用webrtcvad检测音频中的语音区间,过滤静音部分:import webrtcvad def vad_detect(input_file, sr=16000, aggressiveness=3): vad = webrtcvad.Vad(aggressiveness) speech, _ = librosa.load(input_file, sr=sr, mono=True) # 转换为webrtcvad所需的16位PCM格式 pcm_data = (speech * 32767).astype('int16') frame_duration = 30 # 每帧时长(毫秒) frame_samples = int(sr * frame_duration / 1000) vad_segments = [] start_sample = 0 in_speech = False for i in range(0, len(pcm_data), frame_samples): frame = pcm_data[i:i+frame_samples] if len(frame) < frame_samples: break is_speech = vad.is_speech(frame.tobytes(), sr) if is_speech and not in_speech: in_speech = True start_sample = i elif not is_speech and in_speech: in_speech = False vad_segments.append((start_sample, i)) if in_speech: vad_segments.append((start_sample, len(pcm_data))) # 转换为时间(秒) vad_time_segments = [(s/sr, e/sr) for s,e in vad_segments] return vad_time_segments步骤2:基于语义相似度的文本分割
对带时间戳的转写文本,结合时长限制与语义相似度拆分语义单元:from sentence_transformers import SentenceTransformer, util def semantic_split(timestamped_sentences, max_duration=15): model = SentenceTransformer('all-MiniLM-L6-v2') sentences = [s["text"].strip() for s in timestamped_sentences] sentence_embeddings = model.encode(sentences) chunks = [] current_chunk = {"text": "", "start_time": None, "end_time": None, "duration": 0} for idx, sent in enumerate(timestamped_sentences): sent_duration = sent["end_time"] - sent["start_time"] # 若当前chunk加该句子超过15秒,直接拆分 if current_chunk["duration"] + sent_duration > max_duration: chunks.append(current_chunk.copy()) current_chunk = { "text": sent["text"], "start_time": sent["start_time"], "end_time": sent["end_time"], "duration": sent_duration } else: # 计算语义相似度,低于阈值则拆分 if current_chunk["text"]: last_emb = sentence_embeddings[idx-1] curr_emb = sentence_embeddings[idx] sim_score = util.cos_sim(last_emb, curr_emb).item() if sim_score < 0.5: chunks.append(current_chunk.copy()) current_chunk = { "text": sent["text"], "start_time": sent["start_time"], "end_time": sent["end_time"], "duration": sent_duration } else: current_chunk["text"] += " " + sent["text"] current_chunk["end_time"] = sent["end_time"] current_chunk["duration"] = current_chunk["end_time"] - current_chunk["start_time"] else: current_chunk = { "text": sent["text"], "start_time": sent["start_time"], "end_time": sent["end_time"], "duration": sent_duration } if current_chunk["text"]: chunks.append(current_chunk) return chunks步骤3:生成最终音频-文本对
结合VAD段与语义分块结果,确保每个块仅包含有效语音且时长<15秒:def generate_audio_text_pairs(input_file, timestamped_sentences): vad_segments = vad_detect(input_file) semantic_chunks = semantic_split(timestamped_sentences) final_pairs = [] speech, sr = librosa.load(input_file, sr=16000, mono=True) for chunk in semantic_chunks: chunk_start = chunk["start_time"] chunk_end = chunk["end_time"] # 对齐到VAD段边界,过滤静音 for vad_start, vad_end in vad_segments: if vad_start <= chunk_start <= vad_end: chunk_start = vad_start if vad_start <= chunk_end <= vad_end: chunk_end = vad_end # 对超15秒的块进一步拆分 if chunk_end - chunk_start > 15: split_points = [chunk_start + i*15 for i in range(1, int((chunk_end - chunk_start)/15))] split_points.append(chunk_end) prev_split = chunk_start for split in split_points: if split - prev_split > 0.5: # 过滤过短片段 final_pairs.append({ "audio": speech[int(prev_split*sr):int(split*sr)], "text": chunk["text"], "start_time": prev_split, "end_time": split }) prev_split = split else: if chunk_end - chunk_start > 0.5: final_pairs.append({ "audio": speech[int(chunk_start*sr):int(chunk_end*sr)], "text": chunk["text"], "start_time": chunk_start, "end_time": chunk_end }) return final_pairs
二、核心技术概念
- CTC对齐:Connectionist Temporal Classification,无需对齐标签的序列建模方法,通过帧级输出推导token与音频的时间映射,是语音文本对齐的核心。
- 语音活动检测(VAD):识别音频中的语音/静音区间,过滤无效静音,减少后续处理的噪声干扰。
- 语义文本分割:基于文本语义相似度或语法规则,将长文本拆分为语义独立的单元,确保分块后文本具备完整语义。
- 语音特征提取:如Wav2Vec2的声学特征编码,将原始音频转换为高维特征,为识别、对齐提供基础。
- 语义相似度计算:通过预训练语言模型生成文本嵌入,计算文本间的语义距离,判断是否属于同一语义块。
内容的提问来源于stack exchange,提问作者Kartick Narayan
相关产品推荐
相关产品推荐

