技术问询:HuggingFace Accelerate的init_empty_weights在Diffusers中是否正确实现?
讨论
测试发现,HuggingFace Accelerate的init_empty_weights()可将所有文本编码器正确加载到PyTorch meta设备,加载时无明显内存或磁盘占用。
但在测试的两款HuggingFace Diffusers模型(Flux和Stable Diffusion XL)中,该方法表现异常:模型被加载到CPU/CUDA设备,Windows 11性能管理器显示内存占用增加。
与文本编码器相比,init_empty_weights()在Diffusers中的实现是否存在差异、错误或未实现?
代码示例
init_empty_weights()在文本编码器中正常工作
with init_empty_weights(): text_encoder_2 = T5EncoderModel.from_pretrained( "black-forest-labs/FLUX.1-dev", subfolder="text_encoder_2", torch_dtype=torch.float32 ) text_encoder_2.device
Jupyter Notebook返回结果:
device(type='meta')
符合预期,模型仅加载到meta设备,无RAM/VRAM占用增加。
init_empty_weights()在Diffusers中似乎无法正常工作
init_empty_weights()在Flux中似乎无法正常工作
with init_empty_weights(): transformer = FluxTransformer2DModel.from_pretrained( "black-forest-labs/FLUX.1-dev", subfolder="transformer", torch_dtype=torch.bfloat16 ) transformer.device
Jupyter Notebook返回结果:
device(type='cpu')
与预期不符,模型加载到CPU,RAM占用相应增加。
init_empty_weights()在SDXL中似乎无法正常工作
with init_empty_weights(): pipeline = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", torch_dtype=torch.float16, variant="fp16", use_safetensors=True ) pipeline.unet.device
Jupyter Notebook返回结果:
device(type='cpu')
与预期不符,模型加载到CPU,RAM占用相应增加。
背景说明
提出该问题是希望通过初始化空权重模型,传入HuggingFace Accelerate的infer_auto_device_map(),让其自动推断模型层的加载设备。仅为获取形状加载完整模型速度慢,目前有重启内核复用设备映射的繁琐临时解决办法。
内容的提问来源于stack exchange,提问作者Matthew Ross
相关产品推荐
相关产品推荐

