使用PyTorch加载madlad400-3b-mt时device_map="xpu"显存不足问题
在Intel XPU上运行google/madlad400-3b-mt模型的问题与解决方案
问题详情
- 参考google/madlad400-3b-mt模型的使用示例
- 设置
device_map="AUTO"时模型在XPU上无法运行,但在CPU上可正常运行 - 询问缺少哪些配置,如何让模型在XPU上正常运行
运行代码
import torch from transformers import T5ForConditionalGeneration, T5Tokenizer print(torch.xpu.memory_allocated()) print(torch.xpu.device_count()) print(torch.xpu.get_device_name(0)) model_name = 'google/madlad400-3b-mt' model = T5ForConditionalGeneration.from_pretrained(model_name, device_map="xpu", torch_dtype=torch.float16) tokenizer = T5Tokenizer.from_pretrained(model_name) text = "<2pt> I love pizza!" input_ids = tokenizer(text, return_tensors="pt").input_ids.to(device_name) outputs = model.generate(input_ids=input_ids) tokenizer.decode(outputs[0], skip_special_tokens=True) # Eu adoro pizza! print(tokenizer.decode(outputs[0], skip_special_tokens=True)) print(outputs[0])
报错信息
0 1 Intel(R) Arc(TM) A770 Graphics Traceback (most recent call last): File "C:\Users\zhang\source\repos\intel_pytorch2.6.10\translater.py", line 10, in model = T5ForConditionalGeneration.from_pretrained(model_name,device_map="xpu", torch_dtype=torch.float16) File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 279, in _wrapper return func(*args, **kwargs) File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 4399, in from_pretrained ) = cls._load_pretrained_model( ~~~~~~~~~~~~~~~~~~~~~~~~~~^ model, ^^^^^^ ...<13 lines>... weights_only=weights_only, ^^^^^^^^^^^^^^^^^^^^^^^^^^ ) ^ File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 4793, in _load_pretrained_model caching_allocator_warmup(model_to_load, expanded_device_map, factor=2 if hf_quantizer is None else 4) ~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 5803, in caching_allocator_warmup _ = torch.empty(byte_count // factor, dtype=torch.float16, device=device, requires_grad=False) torch.OutOfMemoryError: XPU out of memory. Tried to allocate 7.46 GiB. GPU 0 has a total capacity of 15.56 GiB. Of the allocated memory 0 bytes is allocated by PyTorch, and 0 bytes is reserved by PyTorch but unallocated. Please use `empty_cache` to release all unoccupied cached memory.
解决方案
1. 关闭内存预分配并启用低内存模式
transformers的缓存分配器预热会预分配2倍于模型大小的内存,这是导致OOM的主要原因。加载模型时添加low_cpu_mem_usage=True,并配合device_map="auto"实现智能内存分配:
model = T5ForConditionalGeneration.from_pretrained( model_name, device_map="auto", torch_dtype=torch.float16, low_cpu_mem_usage=True, offload_folder="./offload" # 可选,内存不足时自动卸载部分参数到磁盘 )
2. 修正设备变量错误
代码中device_name未定义,需明确指定XPU设备:
device = "xpu:0" if torch.xpu.is_available() else "cpu" input_ids = tokenizer(text, return_tensors="pt").input_ids.to(device)
3. 手动清理XPU缓存
在加载模型前执行缓存清理,释放闲置内存:
torch.xpu.empty_cache()
4. 启用模型量化(内存紧张时可选)
使用4位/8位量化大幅降低内存占用,需安装bitsandbytes库:
from transformers import BitsAndBytesConfig bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.float16 ) model = T5ForConditionalGeneration.from_pretrained( model_name, device_map="auto", quantization_config=bnb_config, low_cpu_mem_usage=True )
内容的提问来源于stack exchange,提问作者Jzhang
相关产品推荐
相关产品推荐

