能否将Distilled Whisper模型作为OpenAI Whisper的直接替代方案?
无法用Distil-Whisper模型直接替换OpenAI Whisper的问题解决
问题描述
现有基于OpenAI Whisper的视频转录流水线,尝试替换为更小更快的distil-small.en蒸馏模型时,调用whisper.load_model()报错模型未找到,相关信息如下:
原代码
def transcribe(self): file = "/path/to/video" model = whisper.load_model("small.en") # 正常运行 model = whisper.load_model("distil-small.en") # 运行报错 transcript = model.transcribe(word_timestamps=True, audio=file) print(transcript["text"])
错误信息
RuntimeError: Model distil-small.en not found; available models = ['tiny.en', 'tiny', 'base.en', 'base', 'small.en', 'small', 'medium.en', 'medium', 'large-v1', 'large-v2', 'large-v3', 'large']
Poetry依赖配置
[tool.poetry.dependencies] python = "^3.11" openai-whisper = "*" transformers = "*" # 用于加载蒸馏模型 accelerate = "*" # 蒸馏模型依赖 datasets = { version = "*", extras = ["audio"] } # 蒸馏模型依赖
答案
不能直接用whisper.load_model()加载Distil-Whisper模型,因为它不属于OpenAI官方Whisper的内置模型集合,而是Hugging Face社区推出的第三方蒸馏版本,必须通过transformers库加载和调用,无法直接替代原OpenAI Whisper的调用方式。
修改后的代码示例
from transformers import pipeline import torch def transcribe(self): file = "/path/to/video" # 加载Distil-Whisper蒸馏模型 transcriber = pipeline( "automatic-speech-recognition", model="distil-whisper/distil-small.en", torch_dtype=torch.float16 # GPU环境启用,CPU环境可移除 ) # 执行转录并获取词级时间戳 transcript_result = transcriber(file, return_timestamps="word") print(transcript_result["text"])
关键注意点
- 依赖验证:你配置的
transformers、accelerate和带audio扩展的datasets是正确的,确保已完成安装 - 格式适配:原OpenAI Whisper的
transcript结构和transformers返回的结构略有差异,若需和原流水线兼容,需手动转换结果格式 - 环境适配:CPU环境运行时,需移除
torch_dtype=torch.float16参数,避免数据类型错误
内容的提问来源于stack exchange,提问作者user2514157
相关产品推荐
相关产品推荐

