You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow指定GPU0运行失效求助:Ubuntu环境GPU2独占问题

Troubleshooting: Forcing TensorFlow 1.9 to Run on GPU 0 (Instead of GPU 2)

Hey there, let's tackle this frustrating GPU assignment issue you're dealing with. It's super annoying when the standard tricks don't stick, especially after trying every environment variable tweak you can think of. Let's break down targeted fixes tailored to your Ubuntu 16.04 + TensorFlow 1.9 setup:

1. Ensure Environment Variables Are Set Before Importing TensorFlow

TensorFlow detects and initializes GPU devices the moment you import it—so if you set CUDA_VISIBLE_DEVICES after importing tensorflow, it won't have any effect. Fix this by moving the environment variable setup to the absolute top of your script:

import os
# Set this FIRST, before any TensorFlow imports
os.environ['CUDA_DEVICE_ORDER'] = 'PCI_BUS_ID'  # Ensures TF uses same numbering as nvidia-smi
os.environ['CUDA_VISIBLE_DEVICES'] = '0'

# Now import TensorFlow
import tensorflow as tf

The CUDA_DEVICE_ORDER line is critical here—it forces TensorFlow to number GPUs the same way nvidia-smi does. Without this, TF might use a different internal ordering, meaning "GPU 0" in your code is actually GPU 2 in nvidia-smi.

2. Verify No System-Level Environment Variables Are Overriding Your Settings

Check if your system has a persistent CUDA_VISIBLE_DEVICES set that's overriding your script/PyCharm config:

  • Open a terminal and run: echo $CUDA_VISIBLE_DEVICES
  • If it returns 2, you'll need to remove this from your shell config files (like ~/.bashrc, ~/.profile, or /etc/profile).
  • Restart your terminal and PyCharm after making changes to ensure the new environment takes effect.

3. Manually Specify the GPU in TensorFlow Code

Skip environment variables entirely and force TensorFlow to use GPU 0 directly with device contexts and session configs:

import tensorflow as tf

# Configure session to only use GPU 0
config = tf.ConfigProto(
    allow_soft_placement=False,  # Prevents TF from falling back to CPU if GPU 0 is unavailable
    log_device_placement=True,   # Prints which device each operation uses (great for debugging)
    gpu_options=tf.GPUOptions(
        visible_device_list='0',
        allow_growth=True  # Prevents TF from allocating all GPU memory at once
    )
)

# Wrap your model/training code in a GPU 0 device context
with tf.device('/device:GPU:0'):
    # Example operations to test
    x = tf.constant([1.0, 2.0, 3.0], dtype=tf.float32)
    y = tf.multiply(x, 2.0)

# Start the session and verify device usage
with tf.Session(config=config) as sess:
    print("Output:", sess.run(y))
    print("Active GPU:", tf.test.gpu_device_name())

The log_device_placement=True flag will print detailed logs showing exactly which device each tensor/operation is assigned to—this lets you confirm if GPU 0 is actually being used.

4. Validate GPU Numbering Matches Between TensorFlow and nvidia-smi

Sometimes TF's internal GPU numbering doesn't match nvidia-smi. To confirm:

  1. Run nvidia-smi -L in a terminal to get the UUID and numbering of each GPU:
    GPU 0: NVIDIA GeForce RTX 2080 Ti (UUID: GPU-xxxxxx-xxxx-xxxx-xxxx-xxxxxx)
    GPU 1: NVIDIA GeForce RTX 2080 Ti (UUID: GPU-yyyyyy-yyyy-yyyy-yyyy-yyyyyy)
    ...
    
  2. Run this code to check TF's detected devices:
    from tensorflow.python.client import device_lib
    devices = device_lib.list_local_devices()
    for dev in devices:
        if dev.device_type == 'GPU':
            print(f"TF GPU {dev.name.split(':')[-1]}: {dev.physical_device_desc}")
    

Compare the UUIDs in TF's output to nvidia-smi—if TF's "GPU 0" corresponds to nvidia-smi's GPU 2, you'll need to adjust CUDA_VISIBLE_DEVICES to the number that maps to your desired GPU (or use the CUDA_DEVICE_ORDER fix from step 1 to align the numbering).

5. Double-Check PyCharm Run Configuration

Make sure your PyCharm environment variable settings are applied to the correct run configuration:

  1. Go to Run > Edit Configurations
  2. Select the configuration you're using to run your script
  3. Under Environment variables, add CUDA_DEVICE_ORDER=PCI_BUS_ID and CUDA_VISIBLE_DEVICES=0
  4. Uncheck "Include system environment variables" temporarily to rule out overrides from the system
  5. Click Apply and re-run your script

Final Checks

  • Restart your machine after making any system-level environment variable changes—sometimes old settings linger until a reboot.
  • Ensure no other processes are hogging GPU 0 (run nvidia-smi to check for active processes).

内容的提问来源于stack exchange,提问作者Muhammad Ahsan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:07:13