You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Windows10下Docker环境PyTorch模型转GPU速度慢6倍求助

Docker中Stable Diffusion模型迁移到CUDA速度过慢问题

问题背景

在Windows10系统上使用pytorch/pytorch:latest Docker镜像运行DiffusionPipeline,遇到generator.to("cuda")命令耗时3-4分钟,是主机环境(不到30秒)的6倍。需要确认该差异是否合理,以及是否存在解决方案。

Docker环境运行代码

from diffusers import DiffusionPipeline
import transformers
import torch
import os
from global_config import global_config

from torch.profiler import profile, record_function, ProfilerActivity

# 尝试优化性能
torch.set_num_threads(1)

print("Torch version:", torch.__version__)

print("Is CUDA enabled?", torch.cuda.is_available())

print(f"torch version: {torch.version.cuda}")

generator = DiffusionPipeline.from_pretrained(
    "runwayml/stable-diffusion-v1-5", cache_dir=global_config.model_dir()
)

with profile(activities=[ProfilerActivity.CPU], record_shapes=True) as prof:
    with record_function("model_inference"):
        generator.to("cuda")

t = torch.cuda.get_device_properties(0).total_memory
r = torch.cuda.memory_reserved(0)
a = torch.cuda.memory_allocated(0)
f = r-a  # free inside reserved
print(f'total (MB)    : {t/1e6}')
print(f'free (MB)     : {f/1e6}')
print(f'allocated (MB) : {a/1e6}')
print(f'reserved (MB) : {r/1e6}')


print(prof.key_averages().table(sort_by="cpu_time_total", row_limit=10))

Docker环境运行输出

(base) root@22010904699b:/workspace# /opt/conda/bin/python /workspace/src/main.py
Torch version: 2.0.1
Is CUDA enabled? True
torch version: 11.7
device name: NVIDIA GeForce GTX 1080
`text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["id2label"]` will be overriden.
`text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["bos_token_id"]` will be overriden.
`text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["eos_token_id"]` will be overriden.
WARNING:2023-07-25 04:05:31 718:718 init.cpp:146] function cbapi->getCuptiStatus() failed with error CUPTI_ERROR_NOT_INITIALIZED (15)
WARNING:2023-07-25 04:05:31 718:718 init.cpp:147] CUPTI initialization failed - CUDA profiler activities will be missing
INFO:2023-07-25 04:05:31 718:718 init.cpp:149] If you see CUPTI_ERROR_INSUFFICIENT_PRIVILEGES, refer to https://developer.nvidia.com/nvidia-development-tools-solutions-err-nvgpuctrperm-cupti
STAGE:2023-07-25 04:05:31 718:718 ActivityProfilerController.cpp:311] Completed Stage: Warm Up
STAGE:2023-07-25 04:11:16 718:718 ActivityProfilerController.cpp:317] Completed Stage: Collection
STAGE:2023-07-25 04:11:16 718:718 ActivityProfilerController.cpp:321] Completed Stage: Post Processing
100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 50/50 [00:21<00:00,  2.30it/s]
-------------------------------------------  ------------  ------------  ------------  ------------  ------------  ------------  
                                       Name    Self CPU %      Self CPU   CPU total %     CPU total  CPU time avg    # of Calls  
-------------------------------------------  ------------  ------------  ------------  ------------  ------------  ------------  
                            model_inference         0.26%     886.103ms       100.00%      344.795s      344.795s             1  
                                   aten::to         0.01%      18.660ms        99.74%      343.899s     225.065ms          1528  
                             aten::_to_copy         0.01%      33.231ms        99.74%      343.889s     225.058ms          1528  
                                aten::copy_        94.98%      327.473s        94.98%      327.473s     214.315ms          1528  
                        aten::empty_strided         4.75%       16.383s         4.75%       16.383s      10.722ms          1528  
    aten::_has_compatible_shallow_copy_type         0.00%     897.000us         0.00%     897.000us       0.294us          3052  
-------------------------------------------  ------------  ------------  ------------  ------------  ------------  ------------  
Self CPU time total: 344.795s

补充内存使用信息

Docker内内存使用

total (MB)    : 8589.672448
free (MB)     : 88.072192
allocated (MB) : 5507.129344
reserved (MB) : 5595.201536

主机环境内存使用

total (MB)    : 8589.672448
free (MB)     : 129.753088
allocated (MB) : 5499.002880
reserved (MB) : 5628.755968

注:free = reserved - allocated(原注释笔误已修正)

通过nvidia-smi -l 5观察到:两者最大内存使用量均约为6903MiB / 8192MiB;主机环境内存快速达到峰值,Docker环境内存分配逐步上升,耗时数分钟才到峰值。


问题分析与解决方案

1. 差异是否合理?

这种6倍的耗时差异不合理,Docker容器内的CUDA操作性能不应出现如此大幅的下降,核心原因通常和内存分配策略、容器资源限制或CUDA环境配置有关。

2. 常见原因及解决方法

(1)调整CUDA内存分配策略

主机环境可能默认使用了更高效的内存预分配策略,而Docker容器内Torch采用逐步分配模式。可以强制开启内存预分配:

# 在初始化Torch后添加
torch.cuda.empty_cache()
torch.cuda.memory._set_allocator_settings("max_split_size_mb:1024")
# 或者直接预分配固定内存
torch.cuda.memory_reserve(6*1024*1024*1024)  # 预分配6GB

(2)放宽容器CPU资源限制

代码中设置了torch.set_num_threads(1),但Docker容器默认CPU核心/线程限制可能过低,导致占95%耗时的aten::copy_内存拷贝操作效率低下。启动容器时指定更多CPU资源:

docker run --gpus all --cpus 4 --memory 16G [其他参数] pytorch/pytorch:latest

(3)优化WSL2后端性能

如果使用WSL2运行Docker,WSL2与Windows主机的交互存在额外开销:

  • 将模型缓存目录挂载到WSL2内部文件系统,而非Windows主机目录
  • 升级WSL2内核到最新版本
  • 在.wslconfig中设置memory=16GB调整WSL2内存上限

(4)检查CUDA版本与驱动兼容性

确保主机CUDA驱动版本支持容器内的CUDA 11.7版本,驱动版本过低会导致内存操作效率下降。

(5)禁用CUDA Profiler

输出中显示CUPTI初始化失败,可能带来额外性能开销,可在代码开头禁用:

import os
os.environ["CUDA_PROFILE_DISABLE"] = "1"

3. 验证方法

每次修改配置后,重新运行generator.to("cuda")并计时,同时用nvidia-smi观察内存分配速度,确认是否有提升。

内容的提问来源于stack exchange,提问作者Don Kong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 07:22:02