RTX 4070 Super加载DeepSeek-7B-Chat模型CUDA OOM问题求助
问题分析与解决方案
问题背景
- 硬件:NVIDIA RTX 4070 Super(12GB VRAM)
- 目标:加载
deepseek-ai/deepseek-llm-7b-chat模型,开启4bit量化,限制GPU显存使用不超过4GB - 异常情况:加载时触发报错,
nvidia-smi显示GPU显存占用超10GB;设置device_map="cpu"可正常加载
报错信息
Error loading quantized model: Some modules are dispatched on the CPU or the disk. Make sure you have enough GPU RAM to fit the quantized model. If you want to dispatch the model on the CPU or the disk while keeping these modules in 32-bit, you need to set `llm_int8_enable_fp32_cpu_offload=True` and pass a custom `device_map` to `from_pretrained`. Check https://huggingface.co/docs/transformers/main/en/main_classes/quantization#offload-between-cpu-and-gpu for more details.
核心原因
- 手动推断的
device_map未适配4bit量化模型:infer_auto_device_map基于空权重模型计算,未考虑4bit量化后的模型尺寸,导致生成的设备分配逻辑仍按FP16/FP32模型分配,实际加载时GPU显存占用远超限制。 - 缺少量化模型加载的关键优化参数:未启用
low_cpu_mem_usage=True,也未让框架自动处理跨设备offload逻辑。
修复步骤
修改后的代码片段
model_name: str = "deepseek-ai/deepseek-llm-7b-chat" quantization_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4" ) # 直接加载模型,让框架自动处理设备分配 model = AutoModelForCausalLM.from_pretrained( model_name, quantization_config=quantization_config, trust_remote_code=True, device_map="auto", # 自动分配设备 max_memory={0: "4GB", "cpu": "32GB"}, # 限制GPU显存 low_cpu_mem_usage=True, # 优化内存加载 torch_dtype=torch.float16 # 指定张量类型,减少内存占用 )
关键调整说明
- 移除手动推断
device_map的逻辑:改用device_map="auto",让Transformers框架根据量化后的模型尺寸和max_memory限制自动分配设备,适配4bit量化场景。 - 添加
low_cpu_mem_usage=True:减少CPU内存占用,优化模型加载时的内存管理流程。 - 指定
torch_dtype=torch.float16:统一张量类型,避免不必要的类型转换导致的显存浪费。
验证
运行修改后的代码后,通过nvidia-smi查看GPU显存占用,应控制在4GB左右,模型可正常加载并运行。
内容的提问来源于stack exchange,提问作者user938363
相关产品推荐
相关产品推荐

