使用OpenAI Whisper转录全音频时遇AssertionError错误咨询
解决Whisper转录完整音频时的AssertionError: incorrect audio shape问题
你遇到的问题根源是:Whisper的model.detect_language和whisper.decode方法仅支持处理30秒时长的单段音频(也就是pad_or_trim处理后固定形状的音频数据),直接传入完整长音频会因为形状不匹配触发断言错误。
正确解决方案:使用transcribe方法处理完整音频
不需要手动修改音频结构,直接用Whisper内置的transcribe方法即可,它会自动处理长音频的分段、语言检测、完整转录流程。修改后的脚本如下:
import whisper # 加载模型 model = whisper.load_model("large-v2") # 直接调用transcribe处理完整音频,自动处理分段和语言检测 result = model.transcribe("/content/file.mp3", fp16=False) # 打印检测到的语言和转录文本 print(f"Detected language: {result['language']}") print(result['text']) # 保存转录结果到文件 try: with open("output_of_file.txt", "w") as f: f.write(result['text']) print("Transcription saved to file.") except Exception as e: print(f"Error while saving transcription: {e}")
为什么原来的方法不行?
你原来的代码流程是针对单段30秒音频设计的:
pad_or_trim(audio)会把音频裁剪或补零到30秒,符合log_mel_spectrogram、detect_language、decode的输入要求- 去掉后,完整音频的长度远超过30秒,生成的梅尔频谱图形状不符合模型断言的要求,因此报错
若需手动处理长音频(不推荐)
如果一定要手动实现长音频处理,需要把音频分割成30秒的片段,逐个处理后拼接结果,示例代码如下(用Whisper自身的工具函数):
import whisper import numpy as np model = whisper.load_model("large-v2") audio = whisper.load_audio("/content/file.mp3") # 计算每个30秒片段的采样点数量(Whisper默认采样率16000,30秒=480000采样点) chunk_length = 16000 * 30 chunks = [audio[i:i+chunk_length] for i in range(0, len(audio), chunk_length)] # 处理每个片段 full_text = "" detected_lang = None for chunk in chunks: # 补零到30秒(如果最后一段不足) chunk = whisper.pad_or_trim(chunk) mel = whisper.log_mel_spectrogram(chunk).to(model.device) # 第一次检测语言 if not detected_lang: _, probs = model.detect_language(mel) detected_lang = max(probs, key=probs.get) print(f"Detected language: {detected_lang}") # 解码片段 options = whisper.DecodingOptions(fp16=False, language=detected_lang) result = whisper.decode(model, mel, options) full_text += result.text + " " # 打印并保存完整结果 print(full_text.strip()) try: with open("output_of_file.txt", "w") as f: f.write(full_text.strip()) print("Transcription saved to file.") except Exception as e: print(f"Error while saving transcription: {e}")
但强烈建议使用transcribe方法,它内部已经做了更优化的分段、上下文衔接处理,效果比手动拼接更好。
内容的提问来源于stack exchange,提问作者muratowski
相关产品推荐
相关产品推荐

