Windows Docker容器中NVIDIA-SMI无法运行的问题求助
Let's break down your problem and walk through actionable fixes step by step.
Your Setup & Core Problem
You're trying to get GPU access working in a Windows Docker container, but when you run nvidia-smi.exe inside the container, you hit this error:
NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running. This can also be happening if non-NVIDIA GPU is running as primary display, and NVIDIA GPU is in WDDM mode.
Here's what you've already confirmed:
- Host has 2 GTX 1080 Ti GPUs, no physical monitor, using TeamViewer (one GPU acts as Display 1 per dxdiag), no integrated CPU GPU.
- Your Dockerfile uses a vanilla Windows 1903 base image:
FROM mcr.microsoft.com/windows:1903 CMD [ "ping", "-t", "localhost" ]
- Build/run commands include manually mounting the NVSMI folder:
docker build -t debug_image . docker run -d --gpus all --mount src="C:\Program Files\NVIDIA Corporation\NVSMI",target="C:\Program Files\NVIDIA Corporation\NVSMI",type=bind debug_image docker exec -it CONTAINER_ID powershell
- Updated GPU drivers from v431 to v441, confirmed GTX series can't use TCC mode (stuck on WDDM).
Why This Is Happening
Let's unpack the root causes:
- WDDM Mode + Remote Display Lock: WDDM mode ties your GPU to the Windows display stack. Even without a physical monitor, TeamViewer is using one GPU to render the remote display, which locks that GPU's core resources—Docker can't fully access it while it's serving the display session.
- Manual NVSMI Bind Isn't Enough: Mounting just the NVSMI folder doesn't give the container access to the full NVIDIA driver runtime components. The NVIDIA Container Toolkit handles proper driver injection, which a simple folder bind can't replicate.
- Potential OS Version Mismatch: Your container uses Windows 1903—make sure your host is running the exact same Windows build. Docker for Windows requires host/container OS version parity for GPU passthrough to work reliably.
Step-by-Step Solutions
1. Use NVIDIA's Pre-Configured Container Base Images
Ditch the vanilla Windows image for NVIDIA's official base images, which include all necessary runtime components for GPU access. Replace your Dockerfile with something like this (adjust the CUDA version to match your driver):
FROM mcr.microsoft.com/windows/servercore:1903 # Install CUDA runtime (includes nvidia-smi and required driver components) RUN powershell -Command \ Invoke-WebRequest -Uri https://developer.download.nvidia.com/compute/cuda/11.8.0/local_installers/cuda_11.8.0_windows.exe -OutFile cuda_install.exe; \ Start-Process .\cuda_install.exe -ArgumentList "-s" -Wait; \ Remove-Item cuda_install.exe CMD ["ping", "-t", "localhost"]
This ensures the container has all the right bits to communicate with the host's GPU driver.
2. Install & Use the NVIDIA Container Toolkit for Windows
Stop manually mounting NVSMI—install the NVIDIA Container Toolkit, which automatically injects driver components into your container. Once installed, simplify your run command to:
docker run -d --gpus all debug_image
The toolkit handles all underlying driver communication, so you don't need to mess with folder binds.
3. Isolate One GPU from Remote Display
Since TeamViewer is locking one GPU, configure it to use only one card, leaving the other free for Docker:
- Open NVIDIA Control Panel on your host.
- Navigate to 3D Settings > Manage 3D Settings > Program Settings.
- Add
TeamViewer.exeto the program list, then set its preferred GPU to one of your GTX 1080 Ti cards. - Restart TeamViewer and Docker to apply the changes. Now the other GPU should be available for container use.
4. Verify Host & Container OS Version Match
Run winver on your host to check the exact build number. Make sure your Docker base image uses the same major version (e.g., 1903). Mismatched versions often cause silent driver communication failures.
5. Test Without Remote Display (If Possible)
Temporarily disconnect TeamViewer and run nvidia-smi inside the container. If it works, you know the remote display session was the bottleneck. For a permanent fix, you could set up a headless display driver (using tools like a virtual display adapter) for the compute GPU to avoid tying up resources with remote sessions.
Final Thoughts
Windows Docker GPU passthrough is trickier than Linux, especially with WDDM-mode GPUs. The key fixes are using NVIDIA's official container images, leveraging the Container Toolkit, and ensuring your display session doesn't lock all your GPU resources.
内容的提问来源于stack exchange,提问作者Btc Sources

