You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:HuggingFace Accelerate的init_empty_weights在Diffusers中是否正确实现?

讨论

测试发现,HuggingFace Accelerate的init_empty_weights()可将所有文本编码器正确加载到PyTorch meta设备,加载时无明显内存或磁盘占用。

但在测试的两款HuggingFace Diffusers模型(Flux和Stable Diffusion XL)中,该方法表现异常:模型被加载到CPU/CUDA设备,Windows 11性能管理器显示内存占用增加。

与文本编码器相比,init_empty_weights()在Diffusers中的实现是否存在差异、错误或未实现?

代码示例

init_empty_weights()在文本编码器中正常工作

with init_empty_weights():
    text_encoder_2 = T5EncoderModel.from_pretrained(
        "black-forest-labs/FLUX.1-dev",
        subfolder="text_encoder_2",
        torch_dtype=torch.float32
    )

text_encoder_2.device

Jupyter Notebook返回结果:

device(type='meta')

符合预期,模型仅加载到meta设备,无RAM/VRAM占用增加。

init_empty_weights()在Diffusers中似乎无法正常工作

init_empty_weights()在Flux中似乎无法正常工作

with init_empty_weights():
    transformer = FluxTransformer2DModel.from_pretrained(
        "black-forest-labs/FLUX.1-dev",
        subfolder="transformer",
        torch_dtype=torch.bfloat16
    )

transformer.device

Jupyter Notebook返回结果:

device(type='cpu')

与预期不符,模型加载到CPU,RAM占用相应增加。

init_empty_weights()在SDXL中似乎无法正常工作

with init_empty_weights():
    pipeline = StableDiffusionXLPipeline.from_pretrained(
        "stabilityai/stable-diffusion-xl-base-1.0", 
        torch_dtype=torch.float16, 
        variant="fp16", 
        use_safetensors=True
    )

pipeline.unet.device

Jupyter Notebook返回结果:

device(type='cpu')

与预期不符,模型加载到CPU,RAM占用相应增加。

背景说明

提出该问题是希望通过初始化空权重模型,传入HuggingFace Accelerate的infer_auto_device_map(),让其自动推断模型层的加载设备。仅为获取形状加载完整模型速度慢,目前有重启内核复用设备映射的繁琐临时解决办法。

内容的提问来源于stack exchange,提问作者Matthew Ross

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 13:24:59