在GCE VM Ubuntu16.04安装TensorFlow-GPU报错libcublas.so.9.0缺失求解决
Hey there, let's break down this issue clearly and walk through the fixes step by step.
The ImportError: libcublas.so.9.0: cannot open shared object file: No such file or directory error boils down to your system being unable to locate the CUDA 9.0 library file libcublas.so.9.0. Here are the key reasons:
- Version Mismatch: You installed cuDNN v6.0, which is designed to work with CUDA 8.0, not CUDA 9.0. If the TensorFlow GPU version you installed requires CUDA 9.0, it will fail to find the missing library from your CUDA 8.0 setup.
- Incomplete CUDA Installation: If you attempted to install CUDA 9.0 but the process was interrupted or incomplete, the required library files would be missing.
- Failed Environment Variable Configuration: Even if you have the correct CUDA version installed, if the environment variables aren't set properly (or haven't taken effect), your system won't know where to look for the libraries.
Option 1: Match TensorFlow, CUDA, and cuDNN Versions (Recommended)
Since you already have cuDNN v6.0 installed, we'll align everything to work with CUDA 8.0, which is the compatible pair for this cuDNN version.
Step-by-Step Commands:
- Uninstall your current mismatched TensorFlow GPU version:
pip uninstall tensorflow-gpu -y - Install a TensorFlow version that supports CUDA 8.0 + cuDNN 6.0 (e.g., TensorFlow 1.4.0):
pip install tensorflow-gpu==1.4.0 - Verify and fix your environment variables for CUDA 8.0:
Open your bash profile for editing:
Add these lines at the end of the file:nano ~/.bashrc
Save and exit, then apply the changes immediately:export PATH=/usr/local/cuda-8.0/bin${PATH:+:${PATH}} export LD_LIBRARY_PATH=/usr/local/cuda-8.0/lib64${LD_LIBRARY_PATH:+:${LD_LIBRARY_PATH}}source ~/.bashrc - Confirm CUDA 8.0 is properly recognized:
You should see output indicating CUDA version 8.0.x.nvcc --version
Option 2: Install CUDA 9.0 + Compatible cuDNN
If you want to use a TensorFlow version that requires CUDA 9.0, you'll need to update your CUDA and cuDNN to matching versions (CUDA 9.0 pairs with cuDNN 7.0).
Step-by-Step Commands:
- Uninstall your existing CUDA setup first:
sudo apt-get purge nvidia-cuda* -y sudo rm -rf /usr/local/cuda* - Download and install CUDA 9.0:
During installation, choose NOT to install the NVIDIA driver—GCE already has a pre-configured driver for the Tesla K80.wget https://developer.nvidia.com/compute/cuda/9.0/Prod/local_installers/cuda_9.0.176_384.81_linux-run chmod +x cuda_9.0.176_384.81_linux-run sudo ./cuda_9.0.176_384.81_linux-run --override - Install cuDNN 7.0 for CUDA 9.0:
After downloading the cuDNN 7.0 tar package (requires an NVIDIA account), extract and copy the files:tar -xzvf cudnn-9.0-linux-x64-v7.tgz sudo cp cuda/include/cudnn*.h /usr/local/cuda-9.0/include sudo cp cuda/lib64/libcudnn* /usr/local/cuda-9.0/lib64 sudo chmod a+r /usr/local/cuda-9.0/include/cudnn*.h /usr/local/cuda-9.0/lib64/libcudnn* - Configure environment variables for CUDA 9.0:
Edit your bash profile:
Add these lines:nano ~/.bashrc
Apply the changes:export PATH=/usr/local/cuda-9.0/bin${PATH:+:${PATH}} export LD_LIBRARY_PATH=/usr/local/cuda-9.0/lib64${LD_LIBRARY_PATH:+:${LD_LIBRARY_PATH}}source ~/.bashrc - Install a TensorFlow version compatible with CUDA 9.0 + cuDNN 7.0 (e.g., TensorFlow 1.8.0):
pip install tensorflow-gpu==1.8.0
Skip all manual setup by using GCE's pre-built Deep Learning VM Images:
- When launching your GCE instance, search the Marketplace for "Deep Learning VM Image"
- Select the Ubuntu 16.04 variant, choose the Tesla K80 GPU type, and set your disk size to 25GB
- These images come pre-configured with matching versions of CUDA, cuDNN, TensorFlow, and all dependencies—you can start using TensorFlow GPU immediately without any manual setup.
内容的提问来源于stack exchange,提问作者Bob Hopez

