You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅在CPU环境下运行DeepSeek-V3模型推理?

问题

通过SSH连接无GPU但具备多CPU核心的远程机器,尝试运行DeepSeek-V3模型推理,两种方法均失败:

  • 使用DeepSeek-Infer Demo方法,执行命令:
    generate.py --ckpt-path /path/to/DeepSeek-V3-Demo --config configs/config_671B.json --interactive --temperature 0.7 --max-new-tokens 200
    
    报错:
    RuntimeError: Found no NVIDIA driver on your system. Please check that you have an NVIDIA GPU and installed a driver from http://www.nvidia.com/Download/index.aspx
    
  • 使用Hugging-Face Transformer库(v4.51.3),脚本如下:
    # `run_deepseek_v1.py`
    from transformers import AutoModelForCausalLM, AutoTokenizer
    import torch
    torch.manual_seed(30)
    
    tokenizer = AutoTokenizer.from_pretrained("path/to/local/deepseek-v3")
    
    chat = [
      {"role": "user", "content": "Hello, how are you?"},
      {"role": "assistant", "content": "I'm doing great. How can I help you today?"},
      {"role": "user", "content": "I'd like to show off how chat templating works!"},
    ]
    
    
    model = AutoModelForCausalLM.from_pretrained("path/to/local/deepseek-v3", device_map="auto", torch_dtype=torch.bfloat16)
    inputs = tokenizer.apply_chat_template(chat, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device)
    import time
    start = time.time()
    outputs = model.generate(inputs, max_new_tokens=50)
    print(tokenizer.batch_decode(outputs))
    print(time.time()-start)
    
    运行时报错:
    transformers/quantizers/quantizer_finegrained_fp8.py, line 51, in validate_environment raise RuntimeError("No GPU found. A GPU is needed for FP8 quantization.")
    
    修改device_map="auto"为device_map="cpu"后仍报错。

咨询:是否有方法仅在CPU环境下运行DeepSeek-V3推理?理想情况下可使用上述方法之一或其他可行方法。

解决方案

针对Hugging Face Transformers脚本的调整

1. 禁用FP8量化并指定CPU设备

原脚本报错核心是模型默认启用FP8量化(仅支持GPU),需显式禁用量化、强制使用CPU,并切换为CPU兼容的torch.float32数据类型(bfloat16在多数CPU上支持不佳)。修改后的脚本:

# `run_deepseek_v1.py`
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
torch.manual_seed(30)

tokenizer = AutoTokenizer.from_pretrained("path/to/local/deepseek-v3")

chat = [
  {"role": "user", "content": "Hello, how are you?"},
  {"role": "assistant", "content": "I'm doing great. How can I help you today?"},
  {"role": "user", "content": "I'd like to show off how chat templating works!"},
]

# 禁用量化,指定CPU设备,使用float32数据类型
model = AutoModelForCausalLM.from_pretrained(
    "path/to/local/deepseek-v3",
    device_map="cpu",
    torch_dtype=torch.float32,
    load_in_8bit=False,
    load_in_4bit=False,
    quantization_config=None
)
inputs = tokenizer.apply_chat_template(chat, tokenize=True, add_generation_prompt=True, return_tensors="pt").to("cpu")
import time
start = time.time()
outputs = model.generate(inputs, max_new_tokens=50)
print(tokenizer.batch_decode(outputs))
print(time.time()-start)

2. CPU量化优化(可选,减少内存占用)

若机器内存有限,可使用CPU兼容的4-bit量化(需提前安装bitsandbytes的CPU兼容版本):

model = AutoModelForCausalLM.from_pretrained(
    "path/to/local/deepseek-v3",
    device_map="cpu",
    torch_dtype=torch.float32,
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float32
)

针对DeepSeek-Infer Demo的调整

DeepSeek-Infer Demo默认依赖GPU,需修改配置和代码:

  • 修改configs/config_671B.json:设置device: "cpu",将quantization字段改为"none",移除所有GPU专属配置
  • 修改generate.py:替换所有torch.cuda相关调用为CPU版本,强制模型在CPU上初始化和运行

替代方案:使用vLLM CPU推理模式

vLLM针对CPU推理做了多核心优化,适合大模型运行:

  1. 安装CPU版本的vLLM
  2. 交互式推理示例:
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("path/to/local/deepseek-v3")
chat = [
  {"role": "user", "content": "Hello, how are you?"},
  {"role": "assistant", "content": "I'm doing great. How can I help you today?"},
  {"role": "user", "content": "I'd like to show off how chat templating works!"},
]

sampling_params = SamplingParams(max_tokens=50, temperature=0.7)
llm = LLM(model="path/to/local/deepseek-v3", device="cpu")
prompt = tokenizer.apply_chat_template(chat, add_generation_prompt=True)
outputs = llm.generate([prompt], sampling_params)

for output in outputs:
    print(output.outputs[0].text)

内容的提问来源于stack exchange,提问作者The_Average_Engineer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 09:22:12