You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyTorch加载madlad400-3b-mt时device_map="xpu"显存不足问题

在Intel XPU上运行google/madlad400-3b-mt模型的问题与解决方案

问题详情

  • 参考google/madlad400-3b-mt模型的使用示例
  • 设置device_map="AUTO"时模型在XPU上无法运行,但在CPU上可正常运行
  • 询问缺少哪些配置,如何让模型在XPU上正常运行

运行代码

import torch
from transformers import T5ForConditionalGeneration, T5Tokenizer

print(torch.xpu.memory_allocated())
print(torch.xpu.device_count())
print(torch.xpu.get_device_name(0))

model_name = 'google/madlad400-3b-mt'
model = T5ForConditionalGeneration.from_pretrained(model_name, device_map="xpu", torch_dtype=torch.float16)
 
tokenizer = T5Tokenizer.from_pretrained(model_name)

text = "<2pt> I love pizza!"
input_ids = tokenizer(text, return_tensors="pt").input_ids.to(device_name)
outputs = model.generate(input_ids=input_ids)

tokenizer.decode(outputs[0], skip_special_tokens=True)

# Eu adoro pizza!
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
print(outputs[0])

报错信息

0
1
Intel(R) Arc(TM) A770 Graphics
Traceback (most recent call last):
File "C:\Users\zhang\source\repos\intel_pytorch2.6.10\translater.py", line 10, in 
model = T5ForConditionalGeneration.from_pretrained(model_name,device_map="xpu", torch_dtype=torch.float16)
File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 279, in _wrapper
return func(*args, **kwargs)
File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 4399, in from_pretrained
) = cls._load_pretrained_model(
~~~~~~~~~~~~~~~~~~~~~~~~~~^
model,
^^^^^^
...<13 lines>...
weights_only=weights_only,
^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 4793, in _load_pretrained_model
caching_allocator_warmup(model_to_load, expanded_device_map, factor=2 if hf_quantizer is None else 4)
~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\zhang\anaconda3\envs\Pytorch-ipx\Lib\site-packages\transformers\modeling_utils.py", line 5803, in caching_allocator_warmup
_ = torch.empty(byte_count // factor, dtype=torch.float16, device=device, requires_grad=False)
torch.OutOfMemoryError: XPU out of memory. Tried to allocate 7.46 GiB. GPU 0 has a total capacity of 15.56 GiB. Of the allocated memory 0 bytes is allocated by PyTorch, and 0 bytes is reserved by PyTorch but unallocated. Please use `empty_cache` to release all unoccupied cached memory.

解决方案

1. 关闭内存预分配并启用低内存模式

transformers的缓存分配器预热会预分配2倍于模型大小的内存,这是导致OOM的主要原因。加载模型时添加low_cpu_mem_usage=True,并配合device_map="auto"实现智能内存分配:

model = T5ForConditionalGeneration.from_pretrained(
    model_name,
    device_map="auto",
    torch_dtype=torch.float16,
    low_cpu_mem_usage=True,
    offload_folder="./offload"  # 可选,内存不足时自动卸载部分参数到磁盘
)

2. 修正设备变量错误

代码中device_name未定义,需明确指定XPU设备:

device = "xpu:0" if torch.xpu.is_available() else "cpu"
input_ids = tokenizer(text, return_tensors="pt").input_ids.to(device)

3. 手动清理XPU缓存

在加载模型前执行缓存清理,释放闲置内存:

torch.xpu.empty_cache()

4. 启用模型量化(内存紧张时可选)

使用4位/8位量化大幅降低内存占用,需安装bitsandbytes库:

from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16
)

model = T5ForConditionalGeneration.from_pretrained(
    model_name,
    device_map="auto",
    quantization_config=bnb_config,
    low_cpu_mem_usage=True
)

内容的提问来源于stack exchange,提问作者Jzhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 02:42:18