Whisperx Docker部署报错:TranscriptionOptions缺失3个必填参数
Whisperx Docker部署TranscriptionOptions参数错误问题解决
问题描述
在Mac mini M2的VSCode中使用Whisperx提取音频文本,不设置transcription_options时运行正常;但通过Docker(--platform=linux/arm64 python:3.10)部署时,触发TypeError,提示TranscriptionOptions.__new__()缺少max_new_tokens、clip_timestamps和hallucination_silence_threshold三个必填参数。已尝试手动添加这些参数,但问题仍未解决,且使用的是最新版Whisperx。
报错栈
2024-03-20 22:23:15 test-1 | Traceback (most recent call last): 2024-03-20 22:23:15 test-1 | import speech 2024-03-20 22:23:15 test-1 | File "/code/speech.py", line 274, in <module> 2024-03-20 22:23:15 test-1 | File "/code/speech.py", line 69, in get_speeach_to_text 2024-03-20 22:23:15 test-1 | model = whisperx.load_model("medium", device, compute_type=compute_type) 2024-03-20 22:23:15 test-1 | File "/usr/local/lib/python3.10/site-packages/whisperx/asr.py", line 332, in load_model 2024-03-20 22:23:15 test-1 | default_asr_options = faster_whisper.transcribe.TranscriptionOptions(**default_asr_options) 2024-03-20 22:23:15 test-1 | TypeError: TranscriptionOptions.__new__() missing 3 required positional arguments: 'max_new_tokens', 'clip_timestamps', and 'hallucination_silence_threshold'
修改后的代码
import whisperx import faster_whisper # 定义包含必填参数的TranscriptionOptions transcription_options = faster_whisper.transcribe.TranscriptionOptions( beam_size=4, best_of=1, patience=10, length_penalty=0.6, repetition_penalty=1.2, no_repeat_ngram_size=2, log_prob_threshold=-20, no_speech_threshold=0.5, compression_ratio_threshold=0.5, condition_on_previous_text=False, prompt_reset_on_temperature=True, temperatures=[0.7], initial_prompt="", prefix="", suppress_blank=False, suppress_tokens=False, without_timestamps=True, max_initial_timestamp=60, word_timestamps=False, prepend_punctuations="", append_punctuations="", max_new_tokens=50, clip_timestamps=60, hallucination_silence_threshold=0.5 ) device = "cpu" batch_size = 16 # GPU内存不足时可减小 compute_type = "int8" # 加载模型并转写音频 model = whisperx.load_model("medium", device, compute_type=compute_type) audio = whisperx.load_audio(source_audio_file) result = model.transcribe(audio, language='en', task='translate', batch_size=batch_size, transcription_options=transcription_options)
问题根源与解决方法
根源
报错发生在whisperx.load_model内部初始化默认配置的环节,核心原因是Docker环境中faster-whisper版本与本地Mac环境不一致:
- 本地Mac环境的
faster-whisper版本较低,max_new_tokens等三个参数为可选; - Docker中安装的
faster-whisper为最新版,将这三个参数设为必填,但当前使用的Whisperx版本未适配该变化,导致构建默认TranscriptionOptions时缺少参数。
解决步骤
- 锁定
faster-whisper兼容版本:在项目的requirements.txt中添加faster-whisper==0.10.0(该版本不要求这三个参数为必填,且与最新版Whisperx兼容)。 - 更新Docker构建流程:确保Docker镜像构建时安装指定版本的依赖,避免自动拉取最新版
faster-whisper。 - 恢复原有代码逻辑:使用兼容版本后,无需手动定义
transcription_options,本地正常运行的代码在Docker中也能直接生效。
验证方式:在Docker容器内执行pip show faster-whisper查看版本,若为0.11.x及以上则是问题版本,降级后重新运行即可。
内容的提问来源于stack exchange,提问作者AJ152
相关产品推荐
相关产品推荐

