You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CPU模式运行CLIP Interrogator遇slow_conv2d_cpu未实现Half类型错误求解决

问题描述

我尝试在CPU模式下运行CLIP Interrogator,目的是为其他操作保留显存——CUDA模式下主程序的扩散任务会占用过多显存,导致内存不足错误。但运行时出现报错:"slow_conv2d_cpu" not implemented for 'Half',报错来自/usr/local/lib/python3.7/dist-packages/torch/nn/modules/conv.py的_conv_forward函数(代码片段如下):

def _conv_forward(self, input: Tensor, weight: Tensor, bias: Optional[Tensor]):
    if self.padding_mode != 'zeros':
        return F.conv2d(F.pad(input, self._reversed_padding_repeated_twice, mode=self.padding_mode),
                        weight, bias, self.stride,
                        _pair(0), self.dilation, self.groups)
    return F.conv2d(input, weight, bias, self.stride,
                    self.padding, self.dilation, self.groups) #<-- Error on this line

想知道能否强制使用全精度浮点数,或者有没有其他解决办法?

解决办法

1. 强制切换为全精度(FP32)运行

CPU不支持半精度(FP16)的卷积操作,而CLIP Interrogator默认可能启用半精度,因此需要显式指定使用全精度:

  • 初始化Interrogator时,在Config中添加model_kwargs={"torch_dtype": torch.float32}并指定device="cpu":
from clip_interrogator import Config, Interrogator

ci = Interrogator(Config(
    device="cpu",
    model_kwargs={"torch_dtype": torch.float32}
))
  • 如果使用便捷调用函数clip_interrogator,直接传入对应参数:
from clip_interrogator import clip_interrogator

result = clip_interrogator(image, device="cpu", torch_dtype=torch.float32)

2. 其他优化方案

  • 选用轻量CLIP模型:若全精度CPU运行速度过慢,可切换为更小的CLIP模型,比如clip_model_name="openai/clip-vit-base-patch32",降低计算负载。
  • 分批次处理任务:处理大量图片时,分批次加载和处理,避免CPU内存过载。
  • 手动转换张量精度:检查所有输入张量和模型参数,若存在半精度张量,用tensor.float()转换为全精度后再传入模型。

内容的提问来源于stack exchange,提问作者WASasquatch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 18:40:31