You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在仅CPU设备加载量化LLM?报错问题求助

解决CPU环境加载量化StableLM-2-Zephyr-1.6B模型时的NameError: torch未定义问题

问题本质

报错源于transformers与bitsandbytes的集成代码未正确识别torch,同时伴随的警告说明当前bitsandbytes版本不兼容4/8位量化序列化,且你使用的纯CPU环境不适用常规GPU版bitsandbytes。

分步解决方案

1. 显式导入torch

在加载模型的代码开头必须提前导入torch,避免transformers内部调用时找不到依赖:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

# 后续加载模型代码
model = AutoModelForCausalLM.from_pretrained(
    "你的本地量化模型路径",
    device_map="cpu"
)

2. 安装CPU兼容版bitsandbytes

常规pip install bitsandbytes默认是GPU版本,CPU环境需要安装专门分支:

pip uninstall -y bitsandbytes
pip install bitsandbytes-cpu

3. 强制升级到兼容版本组合

确保transformers、torch、bitsandbytes版本匹配,执行以下命令统一升级:

pip install --upgrade transformers>=4.35.0 torch>=2.0.0 bitsandbytes-cpu>=0.41.1

4. 加载时明确指定CPU设备

纯CPU环境下必须显式设置device_map="cpu",避免transformers尝试寻找GPU设备:

model = AutoModelForCausalLM.from_pretrained(
    "./quantized_stablelm",  # 替换为你的模型本地路径
    device_map="cpu",
    trust_remote_code=True  # StableLM系列模型需要此参数
)

额外排查项

  • 确认环境一致性:检查执行pip install的Python环境和运行代码的环境是同一个(比如虚拟环境是否激活)
  • 验证torch可用性:运行python -c "import torch; print(torch.__version__)",确认torch能正常导入且版本符合要求
  • 重新量化模型:若上述步骤无效,可能是量化过程配置错误,重新量化时需指定device_map="cpu"确保生成CPU兼容的量化权重

内容的提问来源于stack exchange,提问作者James Arnold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 06:26:22