使用HuggingFace加载DeepSeek-V3时遇fp8量化类型错误求解决
解决HuggingFace调用DeepSeek-V3时的fp8量化类型错误
问题场景
调用DeepSeek-V3作为嵌入模型时,执行以下代码:
model_kwargs = {"trust_remote_code": True} embedding_model = HuggingFaceEmbeddings( model_kwargs=model_kwargs, model_name="deepseek-ai/DeepSeek-V3")
触发错误:
ValueError:未知量化类型,得到fp8 - 支持的类型包括:['awq', 'bitsandbytes_4bit', 'bitsandbytes_8bit', 'gptq', 'aqlm', 'quanto', 'eetq', 'hqq', 'compressed-tensors', 'fbgemm_fp8', 'torchao']
解决方案
方案1:指定支持的fp8量化类型
错误提示中列出的fbgemm_fp8是官方支持的fp8量化实现,修改model_kwargs添加量化配置指定该类型:
model_kwargs = { "trust_remote_code": True, "quantization_config": {"quantization_method": "fbgemm_fp8"} } embedding_model = HuggingFaceEmbeddings( model_kwargs=model_kwargs, model_name="deepseek-ai/DeepSeek-V3")
方案2:禁用量化,使用全精度模型
如果不需要量化加速,直接在model_kwargs中禁用量化参数:
model_kwargs = { "trust_remote_code": True, "load_in_8bit": False, "load_in_4bit": False, "device_map": "auto" } embedding_model = HuggingFaceEmbeddings( model_kwargs=model_kwargs, model_name="deepseek-ai/DeepSeek-V3")
方案3:更新依赖包到兼容版本
确保transformers、accelerate等核心库版本足够新,以支持所有列出的量化类型:
pip install --upgrade transformers accelerate
内容的提问来源于stack exchange,提问作者John
相关产品推荐
相关产品推荐

