Windows10下Docker环境PyTorch模型转GPU速度慢6倍求助
问题背景
在Windows10系统上使用pytorch/pytorch:latest Docker镜像运行DiffusionPipeline,遇到generator.to("cuda")命令耗时3-4分钟,是主机环境(不到30秒)的6倍。需要确认该差异是否合理,以及是否存在解决方案。
Docker环境运行代码
from diffusers import DiffusionPipeline import transformers import torch import os from global_config import global_config from torch.profiler import profile, record_function, ProfilerActivity # 尝试优化性能 torch.set_num_threads(1) print("Torch version:", torch.__version__) print("Is CUDA enabled?", torch.cuda.is_available()) print(f"torch version: {torch.version.cuda}") generator = DiffusionPipeline.from_pretrained( "runwayml/stable-diffusion-v1-5", cache_dir=global_config.model_dir() ) with profile(activities=[ProfilerActivity.CPU], record_shapes=True) as prof: with record_function("model_inference"): generator.to("cuda") t = torch.cuda.get_device_properties(0).total_memory r = torch.cuda.memory_reserved(0) a = torch.cuda.memory_allocated(0) f = r-a # free inside reserved print(f'total (MB) : {t/1e6}') print(f'free (MB) : {f/1e6}') print(f'allocated (MB) : {a/1e6}') print(f'reserved (MB) : {r/1e6}') print(prof.key_averages().table(sort_by="cpu_time_total", row_limit=10))
Docker环境运行输出
(base) root@22010904699b:/workspace# /opt/conda/bin/python /workspace/src/main.py Torch version: 2.0.1 Is CUDA enabled? True torch version: 11.7 device name: NVIDIA GeForce GTX 1080 `text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["id2label"]` will be overriden. `text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["bos_token_id"]` will be overriden. `text_config_dict` is provided which will be used to initialize `CLIPTextConfig`. The value `text_config["eos_token_id"]` will be overriden. WARNING:2023-07-25 04:05:31 718:718 init.cpp:146] function cbapi->getCuptiStatus() failed with error CUPTI_ERROR_NOT_INITIALIZED (15) WARNING:2023-07-25 04:05:31 718:718 init.cpp:147] CUPTI initialization failed - CUDA profiler activities will be missing INFO:2023-07-25 04:05:31 718:718 init.cpp:149] If you see CUPTI_ERROR_INSUFFICIENT_PRIVILEGES, refer to https://developer.nvidia.com/nvidia-development-tools-solutions-err-nvgpuctrperm-cupti STAGE:2023-07-25 04:05:31 718:718 ActivityProfilerController.cpp:311] Completed Stage: Warm Up STAGE:2023-07-25 04:11:16 718:718 ActivityProfilerController.cpp:317] Completed Stage: Collection STAGE:2023-07-25 04:11:16 718:718 ActivityProfilerController.cpp:321] Completed Stage: Post Processing 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 50/50 [00:21<00:00, 2.30it/s] ------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ Name Self CPU % Self CPU CPU total % CPU total CPU time avg # of Calls ------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ model_inference 0.26% 886.103ms 100.00% 344.795s 344.795s 1 aten::to 0.01% 18.660ms 99.74% 343.899s 225.065ms 1528 aten::_to_copy 0.01% 33.231ms 99.74% 343.889s 225.058ms 1528 aten::copy_ 94.98% 327.473s 94.98% 327.473s 214.315ms 1528 aten::empty_strided 4.75% 16.383s 4.75% 16.383s 10.722ms 1528 aten::_has_compatible_shallow_copy_type 0.00% 897.000us 0.00% 897.000us 0.294us 3052 ------------------------------------------- ------------ ------------ ------------ ------------ ------------ ------------ Self CPU time total: 344.795s
补充内存使用信息
Docker内内存使用
total (MB) : 8589.672448 free (MB) : 88.072192 allocated (MB) : 5507.129344 reserved (MB) : 5595.201536
主机环境内存使用
total (MB) : 8589.672448 free (MB) : 129.753088 allocated (MB) : 5499.002880 reserved (MB) : 5628.755968
注:free = reserved - allocated(原注释笔误已修正)
通过nvidia-smi -l 5观察到:两者最大内存使用量均约为6903MiB / 8192MiB;主机环境内存快速达到峰值,Docker环境内存分配逐步上升,耗时数分钟才到峰值。
问题分析与解决方案
1. 差异是否合理?
这种6倍的耗时差异不合理,Docker容器内的CUDA操作性能不应出现如此大幅的下降,核心原因通常和内存分配策略、容器资源限制或CUDA环境配置有关。
2. 常见原因及解决方法
(1)调整CUDA内存分配策略
主机环境可能默认使用了更高效的内存预分配策略,而Docker容器内Torch采用逐步分配模式。可以强制开启内存预分配:
# 在初始化Torch后添加 torch.cuda.empty_cache() torch.cuda.memory._set_allocator_settings("max_split_size_mb:1024") # 或者直接预分配固定内存 torch.cuda.memory_reserve(6*1024*1024*1024) # 预分配6GB
(2)放宽容器CPU资源限制
代码中设置了torch.set_num_threads(1),但Docker容器默认CPU核心/线程限制可能过低,导致占95%耗时的aten::copy_内存拷贝操作效率低下。启动容器时指定更多CPU资源:
docker run --gpus all --cpus 4 --memory 16G [其他参数] pytorch/pytorch:latest
(3)优化WSL2后端性能
如果使用WSL2运行Docker,WSL2与Windows主机的交互存在额外开销:
- 将模型缓存目录挂载到WSL2内部文件系统,而非Windows主机目录
- 升级WSL2内核到最新版本
- 在
.wslconfig中设置memory=16GB调整WSL2内存上限
(4)检查CUDA版本与驱动兼容性
确保主机CUDA驱动版本支持容器内的CUDA 11.7版本,驱动版本过低会导致内存操作效率下降。
(5)禁用CUDA Profiler
输出中显示CUPTI初始化失败,可能带来额外性能开销,可在代码开头禁用:
import os os.environ["CUDA_PROFILE_DISABLE"] = "1"
3. 验证方法
每次修改配置后,重新运行generator.to("cuda")并计时,同时用nvidia-smi观察内存分配速度,确认是否有提升。
内容的提问来源于stack exchange,提问作者Don Kong

