You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否将Distilled Whisper模型作为OpenAI Whisper的直接替代方案?

无法用Distil-Whisper模型直接替换OpenAI Whisper的问题解决

问题描述

现有基于OpenAI Whisper的视频转录流水线,尝试替换为更小更快的distil-small.en蒸馏模型时,调用whisper.load_model()报错模型未找到,相关信息如下:

原代码

def transcribe(self):
    file = "/path/to/video"

    model = whisper.load_model("small.en")          # 正常运行
    model = whisper.load_model("distil-small.en")   # 运行报错

    transcript = model.transcribe(word_timestamps=True, audio=file)
    print(transcript["text"])

错误信息

RuntimeError: Model distil-small.en not found; available models = ['tiny.en', 'tiny', 'base.en', 'base', 'small.en', 'small', 'medium.en', 'medium', 'large-v1', 'large-v2', 'large-v3', 'large']

Poetry依赖配置

[tool.poetry.dependencies]
python = "^3.11"
openai-whisper = "*"
transformers  = "*" # 用于加载蒸馏模型
accelerate  = "*" # 蒸馏模型依赖
datasets = { version = "*", extras = ["audio"] } # 蒸馏模型依赖

答案

不能直接用whisper.load_model()加载Distil-Whisper模型,因为它不属于OpenAI官方Whisper的内置模型集合,而是Hugging Face社区推出的第三方蒸馏版本,必须通过transformers库加载和调用,无法直接替代原OpenAI Whisper的调用方式。

修改后的代码示例

from transformers import pipeline
import torch

def transcribe(self):
    file = "/path/to/video"
    
    # 加载Distil-Whisper蒸馏模型
    transcriber = pipeline(
        "automatic-speech-recognition",
        model="distil-whisper/distil-small.en",
        torch_dtype=torch.float16  # GPU环境启用,CPU环境可移除
    )
    
    # 执行转录并获取词级时间戳
    transcript_result = transcriber(file, return_timestamps="word")
    print(transcript_result["text"])

关键注意点

  1. 依赖验证:你配置的transformers、accelerate和带audio扩展的datasets是正确的,确保已完成安装
  2. 格式适配:原OpenAI Whisper的transcript结构和transformers返回的结构略有差异,若需和原流水线兼容,需手动转换结果格式
  3. 环境适配:CPU环境运行时,需移除torch_dtype=torch.float16参数,避免数据类型错误

内容的提问来源于stack exchange,提问作者user2514157

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:22:21