You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu AWS环境下Anaconda安装TensorFlow-GPU导入失败求助

Fixing "libcuda.so.1: cannot open shared object file" Error with TensorFlow-GPU on AWS Ubuntu

Hey there, let's work through this issue together! The error you're seeing means TensorFlow's GPU version can't locate the system-level CUDA driver files—and this is super common with AWS EC2 instances for a couple of key reasons. Here's how to fix it:

First: Check Your AWS Instance Type

The most likely culprit is that you're using an EC2 instance without a GPU (like t2, m5, c5 series). TensorFlow-GPU requires a NVIDIA GPU to run.

  • If you're on a non-GPU instance: Either switch to a GPU-enabled instance (g4dn, p3, p4, etc.) or install the CPU-only version of TensorFlow instead:
    conda uninstall tensorflow-gpu
    conda install tensorflow
    

If You're Using a GPU Instance: Install NVIDIA Drivers & Verify Compatibility

GPU instances on AWS don't come with NVIDIA drivers pre-installed by default—you have to set them up manually. Here's the step-by-step:

  1. Update your system packages

    sudo apt update && sudo apt upgrade -y
    
  2. Install the correct NVIDIA driver
    Use Ubuntu's package manager to install a stable driver version compatible with your GPU. For most modern AWS GPUs (like T4, A10G), driver 525+ works well:

    sudo apt install nvidia-driver-525 -y
    

    After installation, reboot your instance to apply changes:

    sudo reboot
    
  3. Verify the driver installation
    Once your instance restarts, run this command to confirm the GPU and driver are recognized:

    nvidia-smi
    

    You should see a table showing your GPU model, driver version, and supported CUDA version.

  4. Ensure TensorFlow-GPU matches your system's CUDA version
    Conda's TensorFlow-GPU package includes its own CUDA runtime, but it needs to be compatible with the driver's supported CUDA version. For example, if nvidia-smi shows CUDA 12.0, install a TensorFlow version that supports it:

    conda install tensorflow-gpu==2.13.0 -y
    
  5. Fix missing libcuda.so.1 symlink (if needed)
    If the driver is installed but you still get the error, the soft link for libcuda.so.1 might be missing. First find the actual driver file:

    sudo find / -name libcuda.so.*
    

    You'll get a path like /usr/lib/x86_64-linux-gnu/libcuda.so.525.147.05. Create a symlink to libcuda.so.1:

    sudo ln -s /usr/lib/x86_64-linux-gnu/libcuda.so.525.147.05 /usr/lib/x86_64-linux-gnu/libcuda.so.1
    

    Then update the dynamic link cache:

    sudo ldconfig
    

Test Your Setup

Finally, activate your conda environment and test TensorFlow:

conda activate tensorflow-gpu
python -c "import tensorflow as tf; print(tf.config.list_physical_devices('GPU'))"

If you see output listing your GPU device, you're all set!

内容的提问来源于stack exchange,提问作者ZHANG Juenjie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:46:15