AWS DLAMI使用Keras multi_gpu_model时GPU无法识别问题求助
Hey there, let's break down what's happening here and how to fix it:
Why You're Seeing This Error
The g2.8xlarge EC2 instance uses GRID K520 GPUs, which have a CUDA compute capability of 3.0. Starting from TensorFlow 1.10 and later, official builds dropped support for GPUs with compute capability below 3.5. Since Keras' multi_gpu_model relies on TensorFlow's GPU handling logic, this is why your GPU is getting ignored. Reinstalling a newer TensorFlow version won't help here—it's intentional behavior in those recent releases.
Possible Solutions
Downgrade TensorFlow to a version that supports compute capability 3.0
You'll need to install a TensorFlow version prior to 1.10. The last official release that supports 3.0 is TensorFlow 1.9. Make sure to pair it with a compatible Keras version (e.g., Keras 2.2.4 works well with TensorFlow 1.9). Run one of these commands depending on your package manager:# Using pip pip install tensorflow-gpu==1.9.0 keras==2.2.4 # Using conda (for the Deep Learning AMI's conda environment) conda install tensorflow-gpu=1.9.0 keras=2.2.4After downgrading, restart your Jupyter Notebook session and try using
multi_gpu_modelagain.Skip multi_gpu_model and use single GPU training
If downgrading isn't ideal for your project, you can train your model on a single GPU instead. The GRID K520 has 4GB of VRAM, so as long as your LSTM model fits within that memory limit, this will work. Just remove themulti_gpu_modelwrapper from your code and compile/train the base model directly.Upgrade your EC2 instance type
For long-term projects, consider switching to an EC2 instance with GPUs that have a CUDA compute capability of 3.5 or higher. Suitable options include:p2.xlarge/p2.8xlarge/p2.16xlarge(uses NVIDIA K80, compute capability 3.7)g3.xlarge/g3.8xlarge(uses NVIDIA M60, compute capability 5.2)p3.2xlarge/p3.8xlarge/p3.16xlarge(uses NVIDIA V100, compute capability 7.0)
These instances will work seamlessly with modern TensorFlow/Keras versions andmulti_gpu_model.
内容的提问来源于stack exchange,提问作者Joe Urc

