You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用OpenAI Whisper提取MP4音频文本遇TypeError问题求助

解决Whisper Python API解码时的TypeError问题

问题场景

尝试用OpenAI Whisper的Python包提取MP4音频文本,使用官方最简示例代码:

import whisper

model = whisper.load_model("base")

# load audio and pad/trim it to fit 30 seconds
audio = whisper.load_audio("./portuguese.mp4")
audio = whisper.pad_or_trim(audio)

# make log-Mel spectrogram and move to the same device as the model
mel = whisper.log_mel_spectrogram(audio).to(model.device)

# detect the spoken language
_, probs = model.detect_language(mel)
print(f"Detected language: {max(probs, key=probs.get)}")

# decode the audio
options = whisper.DecodingOptions(fp16 = False)
result = whisper.decode(model, mel, options)

# print the recognized text
print(result.text)

运行后抛出以下错误:

File ~/anaconda3/lib/python3.10/site-packages/spyder_kernels/py3compat.py:356 in compat_exec
    exec(code, globals, locals)

  File ~/Downloads/audio_to_text.py:18
    result = whisper.decode(model, mel, options)

  File ~/anaconda3/lib/python3.10/site-packages/torch/utils/_contextlib.py:115 in decorate_context
    return func(*args, **kwargs)

  File ~/anaconda3/lib/python3.10/site-packages/whisper/decoding.py:811 in decode

  File ~/anaconda3/lib/python3.10/site-packages/torch/utils/_contextlib.py:115 in decorate_context
    return func(*args, **kwargs)

  File ~/anaconda3/lib/python3.10/site-packages/whisper/decoding.py:744 in run
    tokens: List[List[int]] = [t[i].tolist() for i, t in zip(selected, tokens)]

  File ~/anaconda3/lib/python3.10/site-packages/whisper/decoding.py:744 in <listcomp>
    tokens: List[List[int]] = [t[i].tolist() for i, t in zip(selected, tokens)]

TypeError: only integer scalar arrays can be converted to a scalar index

但通过终端执行命令whisper portuguese.mp4 --language Portuguese可正常运行。

错误原因

官方示例的pad_or_trim将音频强制裁剪为30秒单段,而whisper.decode函数在处理单段短音频时,部分旧版本Whisper的内部索引逻辑存在bug;而终端命令行工具默认使用whisper.transcribe方法,该方法针对完整音频的分段处理逻辑更健壮,不会触发此错误。

解决方案

方案1:改用transcribe方法(推荐)

直接使用transcribe替代decode,这也是终端命令行工具的底层逻辑,无需手动裁剪音频,处理更稳定:

import whisper

model = whisper.load_model("base")

# 直接转录完整音频,指定语言可提升准确率
result = model.transcribe("./portuguese.mp4", language="Portuguese", fp16=False)

print(f"Detected language: {result['language']}")
print(result['text'])

方案2:修复decode的使用逻辑

如果坚持使用decode,需避免单段音频的索引问题,可调整音频处理方式,同时确保Whisper版本为最新:

  1. 升级Whisper到最新版本:
pip install --upgrade openai-whisper
  1. 修改代码,去掉pad_or_trim,改用完整音频的Mel谱(注意:大音频可能需要更长处理时间):
import whisper

model = whisper.load_model("base")

# 加载完整音频,不裁剪
audio = whisper.load_audio("./portuguese.mp4")
mel = whisper.log_mel_spectrogram(audio).to(model.device)

# 检测语言
_, probs = model.detect_language(mel)
detected_lang = max(probs, key=probs.get)
print(f"Detected language: {detected_lang}")

# 解码时指定语言
options = whisper.DecodingOptions(fp16=False, language=detected_lang)
result = whisper.decode(model, mel, options)

print(result.text)

内容的提问来源于stack exchange,提问作者donut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 02:32:43