You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用OpenAI Whisper做STT遇'slow_conv2d_cpu' Half类型错误,如何解决?

解决Whisper运行时slow_conv2d_cpu not implemented for 'Half'错误

错误原因

这个错误的核心是模型使用了半精度(FP16)数据类型,但你的CPU不支持对应的卷积运算。Whisper默认会根据设备自动选择精度,若检测到GPU会用半精度提速,但CPU通常只能兼容全精度(FP32),此时半精度的卷积操作就会触发未实现的报错。

解决方案

1. 强制用全精度加载模型(最直接有效)

修改模型加载代码,明确指定设备为CPU并使用全精度(FP32):

import torch
import whisper

# 强制CPU设备+全精度加载模型
model = whisper.load_model("base", device="cpu", dtype=torch.float32)

这样模型会全程用FP32运行,避开CPU不支持的半精度操作。

2. 升级PyTorch版本

部分旧版PyTorch对CPU半精度的支持存在缺陷,尝试升级到最新稳定版:

pip install --upgrade torch

注:如果是较老的CPU,升级后仍可能不支持半精度,优先使用第一个方案。

3. 确保张量与模型精度一致

如果需要保留原有加载逻辑,可将mel张量强制转换为全精度:

mel = whisper.log_mel_spectrogram(audio).to(model.device, dtype=torch.float32)

注:此方法需配合模型精度调整使用,单独使用可能仍有问题。

原始错误信息

Traceback (most recent call last):
  File "/Users/reallymemorable/git/fp-stt/2-stt.py", line 20, in <module>
    result = whisper.decode(model, mel, options)
  File "/opt/homebrew/lib/python3.10/site-packages/torch/autograd/grad_mode.py", line 27, in decorate_context
    return func(*args, **kwargs)
  File "/opt/homebrew/lib/python3.10/site-packages/whisper/decoding.py", line 705, in decode
    result = DecodingTask(model, options).run(mel)
  File "/opt/homebrew/lib/python3.10/site-packages/torch/autograd/grad_mode.py", line 27, in decorate_context
    return func(*args, **kwargs)
  File "/opt/homebrew/lib/python3.10/site-packages/whisper/decoding.py", line 621, in run
    audio_features: Tensor = self._get_audio_features(mel)  # encoder forward pass
  File "/opt/homebrew/lib/python3.10/site-packages/whisper/decoding.py", line 565, in _get_audio_features
    audio_features = self.model.encoder(mel)
  File "/opt/homebrew/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1190, in _call_impl
    return forward_call(*input, **kwargs)
  File "/opt/homebrew/lib/python3.10/site-packages/whisper/model.py", line 148, in forward
    x = F.gelu(self.conv1(x))
  File "/opt/homebrew/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1190, in _call_impl
    return forward_call(*input, **kwargs)
  File "/opt/homebrew/lib/python3.10/site-packages/torch/nn/modules/conv.py", line 313, in forward
    return self._conv_forward(input, self.weight, self.bias)
  File "/opt/homebrew/lib/python3.10/site-packages/whisper/model.py", line 43, in _conv_forward
    return super()._conv_forward(
  File "/opt/homebrew/lib/python3.10/site-packages/torch/nn/modules/conv.py", line 309, in _conv_forward
    return F.conv1d(input, weight, bias, self.stride,
RuntimeError: "slow_conv2d_cpu" not implemented for 'Half'

原始代码

import whisper

model = whisper.load_model("base")

# load audio and pad/trim it to fit 30 seconds
audio = whisper.load_audio("speech-to-text-sample.wav")
audio = whisper.pad_or_trim(audio)

# make log-Mel spectrogram and move to the same device as the model
mel = whisper.log_mel_spectrogram(audio).to(model.device)

# detect the spoken language
_, probs = model.detect_language(mel)
print(f"Detected language: {max(probs, key=probs.get)}")

# decode the audio
options = whisper.DecodingOptions()
result = whisper.decode(model, mel, options)

# print the recognized text
print(result.text)

内容的提问来源于stack exchange,提问作者reallymemorable

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:40:32