使用OpenAI Whisper转写音频触发ValueError问题求助
Whisper转写报错
ValueError: tuple.index(x): x not in tuple的解决办法 环境与问题
使用Python3.9,已安装ffmpeg及相关依赖,openai-whisper版本为20230308,可正常导入whisper库,但执行以下转写代码时出现错误:
audio = f'audio_dataset/testaudio.wav' model = whisper.load_model("base") result = model.transcribe(audio, fp16=False, language="en")
错误栈信息:
ValueError Traceback (most recent call last) Cell In [32], line 1 ----> 1 result = model.transcribe(audio, fp16=False, language="en") File ~/.local/lib/python3.9/site-packages/torch/autograd/grad_mode.py:27, in _DecoratorContextManager.__call__.<locals>.decorate_context(*args, **kwargs) 24 @functools.wraps(func) 25 def decorate_context(*args, **kwargs): 26 with self.clone(): ---> 27 return func(*args, **kwargs) File ~/.local/lib/python3.9/site-packages/whisper/decoding.py:811, in decode(model, mel, options, **kwargs) 808 if kwargs: 809 options = replace(options, **kwargs) ---> 811 result = DecodingTask(model, options).run(mel) 813 return result[0] if single else result File ~/.local/lib/python3.9/site-packages/whisper/decoding.py:522, in DecodingTask.__init__(self, model, options) 520 self.initial_tokens: Tuple[int] = self._get_initial_tokens() 521 self.sample_begin: int = len(self.initial_tokens) ---> 522 self.sot_index: int = self.initial_tokens.index(tokenizer.sot) 524 # inference: implements the forward pass through the decoder, including kv caching 525 self.inference = PyTorchInference(model, len(self.initial_tokens)) ValueError: tuple.index(x): x not in tuple
解决办法
该错误是20230308版本Whisper的已知bug:指定language参数后,初始token列表缺失sot(转录起始标记),导致代码无法找到其索引。可尝试以下方案:
升级Whisper版本(最推荐)
后续版本已修复此bug,执行升级命令即可:pip install --upgrade openai-whisper临时修改源码(不升级版本)
找到Python依赖目录下的whisper/decoding.py(路径示例:~/.local/lib/python3.9/site-packages/whisper/decoding.py),定位到DecodingTask类的_get_initial_tokens方法,确保返回的tokens元组中包含tokenizer.sot,比如在添加语言token前先插入sot。暂时移除language参数
不指定语言,让模型自动检测,验证是否能正常运行:result = model.transcribe(audio, fp16=False)
内容的提问来源于stack exchange,提问作者SRaj
相关产品推荐
相关产品推荐

