You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ivrit-ai/whisper-large-v2-tuned长音频转录遇多语言检测错误

解决Whisper模型多语言检测报错问题

你遇到的错误是因为模型在处理长音频分块时,检测到了多种语言,而当前版本不支持在单批次转录中处理多语言内容。解决方法有两种:

方法一:强制指定目标语言

直接在调用pipeline的pipe()方法时,添加language参数,指定音频对应的语言代码(比如希伯来语用'he',英语用'en'),这样模型会跳过语言检测,直接用指定语言转录。

修改后的核心代码片段:

prediction = pipe(audio, batch_size=8, return_timestamps=True, language='he')["chunks"]

完整修改后的代码:

import torch
from transformers import pipeline

device = "cuda:0" if torch.cuda.is_available() else "cpu"

pipe = pipeline(
  "automatic-speech-recognition",
  model="ivrit-ai/whisper-large-v2-tuned",
  chunk_length_s=30,
  device=device,
)

audio_file = './audio/sales_call.mp3'  

with open(audio_file, 'rb') as file:
    audio = file.read()

# 添加language参数指定目标语言,示例为希伯来语(根据你的音频实际语言调整)
prediction = pipe(audio, batch_size=8, return_timestamps=True, language='he')["chunks"]

with open('transcription.txt', 'w', encoding='utf-8') as file:
    for item in prediction:
        file.write(f"{item['text']},{item['timestamp']}\n")

方法二:确保音频为单语言

如果你的音频确实包含多种语言,需要先将音频按语言拆分,再分别转录;如果是模型误检测,检查音频是否有噪音、多语言片段干扰,清理后再尝试。

内容的提问来源于stack exchange,提问作者Eyal Solomon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 18:15:58