使用OpenAI Whisper提取MP4音频文本遇TypeError问题求助
解决Whisper Python API解码时的TypeError问题
问题场景
尝试用OpenAI Whisper的Python包提取MP4音频文本,使用官方最简示例代码:
import whisper model = whisper.load_model("base") # load audio and pad/trim it to fit 30 seconds audio = whisper.load_audio("./portuguese.mp4") audio = whisper.pad_or_trim(audio) # make log-Mel spectrogram and move to the same device as the model mel = whisper.log_mel_spectrogram(audio).to(model.device) # detect the spoken language _, probs = model.detect_language(mel) print(f"Detected language: {max(probs, key=probs.get)}") # decode the audio options = whisper.DecodingOptions(fp16 = False) result = whisper.decode(model, mel, options) # print the recognized text print(result.text)
运行后抛出以下错误:
File ~/anaconda3/lib/python3.10/site-packages/spyder_kernels/py3compat.py:356 in compat_exec exec(code, globals, locals) File ~/Downloads/audio_to_text.py:18 result = whisper.decode(model, mel, options) File ~/anaconda3/lib/python3.10/site-packages/torch/utils/_contextlib.py:115 in decorate_context return func(*args, **kwargs) File ~/anaconda3/lib/python3.10/site-packages/whisper/decoding.py:811 in decode File ~/anaconda3/lib/python3.10/site-packages/torch/utils/_contextlib.py:115 in decorate_context return func(*args, **kwargs) File ~/anaconda3/lib/python3.10/site-packages/whisper/decoding.py:744 in run tokens: List[List[int]] = [t[i].tolist() for i, t in zip(selected, tokens)] File ~/anaconda3/lib/python3.10/site-packages/whisper/decoding.py:744 in <listcomp> tokens: List[List[int]] = [t[i].tolist() for i, t in zip(selected, tokens)] TypeError: only integer scalar arrays can be converted to a scalar index
但通过终端执行命令whisper portuguese.mp4 --language Portuguese可正常运行。
错误原因
官方示例的pad_or_trim将音频强制裁剪为30秒单段,而whisper.decode函数在处理单段短音频时,部分旧版本Whisper的内部索引逻辑存在bug;而终端命令行工具默认使用whisper.transcribe方法,该方法针对完整音频的分段处理逻辑更健壮,不会触发此错误。
解决方案
方案1:改用transcribe方法(推荐)
直接使用transcribe替代decode,这也是终端命令行工具的底层逻辑,无需手动裁剪音频,处理更稳定:
import whisper model = whisper.load_model("base") # 直接转录完整音频,指定语言可提升准确率 result = model.transcribe("./portuguese.mp4", language="Portuguese", fp16=False) print(f"Detected language: {result['language']}") print(result['text'])
方案2:修复decode的使用逻辑
如果坚持使用decode,需避免单段音频的索引问题,可调整音频处理方式,同时确保Whisper版本为最新:
- 升级Whisper到最新版本:
pip install --upgrade openai-whisper
- 修改代码,去掉
pad_or_trim,改用完整音频的Mel谱(注意:大音频可能需要更长处理时间):
import whisper model = whisper.load_model("base") # 加载完整音频,不裁剪 audio = whisper.load_audio("./portuguese.mp4") mel = whisper.log_mel_spectrogram(audio).to(model.device) # 检测语言 _, probs = model.detect_language(mel) detected_lang = max(probs, key=probs.get) print(f"Detected language: {detected_lang}") # 解码时指定语言 options = whisper.DecodingOptions(fp16=False, language=detected_lang) result = whisper.decode(model, mel, options) print(result.text)
内容的提问来源于stack exchange,提问作者donut
相关产品推荐
相关产品推荐

