使用ivrit-ai/whisper-large-v2-tuned长音频转录遇多语言检测错误
解决Whisper模型多语言检测报错问题
你遇到的错误是因为模型在处理长音频分块时,检测到了多种语言,而当前版本不支持在单批次转录中处理多语言内容。解决方法有两种:
方法一:强制指定目标语言
直接在调用pipeline的pipe()方法时,添加language参数,指定音频对应的语言代码(比如希伯来语用'he',英语用'en'),这样模型会跳过语言检测,直接用指定语言转录。
修改后的核心代码片段:
prediction = pipe(audio, batch_size=8, return_timestamps=True, language='he')["chunks"]
完整修改后的代码:
import torch from transformers import pipeline device = "cuda:0" if torch.cuda.is_available() else "cpu" pipe = pipeline( "automatic-speech-recognition", model="ivrit-ai/whisper-large-v2-tuned", chunk_length_s=30, device=device, ) audio_file = './audio/sales_call.mp3' with open(audio_file, 'rb') as file: audio = file.read() # 添加language参数指定目标语言,示例为希伯来语(根据你的音频实际语言调整) prediction = pipe(audio, batch_size=8, return_timestamps=True, language='he')["chunks"] with open('transcription.txt', 'w', encoding='utf-8') as file: for item in prediction: file.write(f"{item['text']},{item['timestamp']}\n")
方法二:确保音频为单语言
如果你的音频确实包含多种语言,需要先将音频按语言拆分,再分别转录;如果是模型误检测,检查音频是否有噪音、多语言片段干扰,清理后再尝试。
内容的提问来源于stack exchange,提问作者Eyal Solomon
相关产品推荐
相关产品推荐

