TensorFlow指定GPU0运行失效求助:Ubuntu环境GPU2独占问题
Hey there, let's tackle this frustrating GPU assignment issue you're dealing with. It's super annoying when the standard tricks don't stick, especially after trying every environment variable tweak you can think of. Let's break down targeted fixes tailored to your Ubuntu 16.04 + TensorFlow 1.9 setup:
1. Ensure Environment Variables Are Set Before Importing TensorFlow
TensorFlow detects and initializes GPU devices the moment you import it—so if you set CUDA_VISIBLE_DEVICES after importing tensorflow, it won't have any effect. Fix this by moving the environment variable setup to the absolute top of your script:
import os # Set this FIRST, before any TensorFlow imports os.environ['CUDA_DEVICE_ORDER'] = 'PCI_BUS_ID' # Ensures TF uses same numbering as nvidia-smi os.environ['CUDA_VISIBLE_DEVICES'] = '0' # Now import TensorFlow import tensorflow as tf
The CUDA_DEVICE_ORDER line is critical here—it forces TensorFlow to number GPUs the same way nvidia-smi does. Without this, TF might use a different internal ordering, meaning "GPU 0" in your code is actually GPU 2 in nvidia-smi.
2. Verify No System-Level Environment Variables Are Overriding Your Settings
Check if your system has a persistent CUDA_VISIBLE_DEVICES set that's overriding your script/PyCharm config:
- Open a terminal and run:
echo $CUDA_VISIBLE_DEVICES - If it returns
2, you'll need to remove this from your shell config files (like~/.bashrc,~/.profile, or/etc/profile). - Restart your terminal and PyCharm after making changes to ensure the new environment takes effect.
3. Manually Specify the GPU in TensorFlow Code
Skip environment variables entirely and force TensorFlow to use GPU 0 directly with device contexts and session configs:
import tensorflow as tf # Configure session to only use GPU 0 config = tf.ConfigProto( allow_soft_placement=False, # Prevents TF from falling back to CPU if GPU 0 is unavailable log_device_placement=True, # Prints which device each operation uses (great for debugging) gpu_options=tf.GPUOptions( visible_device_list='0', allow_growth=True # Prevents TF from allocating all GPU memory at once ) ) # Wrap your model/training code in a GPU 0 device context with tf.device('/device:GPU:0'): # Example operations to test x = tf.constant([1.0, 2.0, 3.0], dtype=tf.float32) y = tf.multiply(x, 2.0) # Start the session and verify device usage with tf.Session(config=config) as sess: print("Output:", sess.run(y)) print("Active GPU:", tf.test.gpu_device_name())
The log_device_placement=True flag will print detailed logs showing exactly which device each tensor/operation is assigned to—this lets you confirm if GPU 0 is actually being used.
4. Validate GPU Numbering Matches Between TensorFlow and nvidia-smi
Sometimes TF's internal GPU numbering doesn't match nvidia-smi. To confirm:
- Run
nvidia-smi -Lin a terminal to get the UUID and numbering of each GPU:GPU 0: NVIDIA GeForce RTX 2080 Ti (UUID: GPU-xxxxxx-xxxx-xxxx-xxxx-xxxxxx) GPU 1: NVIDIA GeForce RTX 2080 Ti (UUID: GPU-yyyyyy-yyyy-yyyy-yyyy-yyyyyy) ... - Run this code to check TF's detected devices:
from tensorflow.python.client import device_lib devices = device_lib.list_local_devices() for dev in devices: if dev.device_type == 'GPU': print(f"TF GPU {dev.name.split(':')[-1]}: {dev.physical_device_desc}")
Compare the UUIDs in TF's output to nvidia-smi—if TF's "GPU 0" corresponds to nvidia-smi's GPU 2, you'll need to adjust CUDA_VISIBLE_DEVICES to the number that maps to your desired GPU (or use the CUDA_DEVICE_ORDER fix from step 1 to align the numbering).
5. Double-Check PyCharm Run Configuration
Make sure your PyCharm environment variable settings are applied to the correct run configuration:
- Go to Run > Edit Configurations
- Select the configuration you're using to run your script
- Under Environment variables, add
CUDA_DEVICE_ORDER=PCI_BUS_IDandCUDA_VISIBLE_DEVICES=0 - Uncheck "Include system environment variables" temporarily to rule out overrides from the system
- Click Apply and re-run your script
Final Checks
- Restart your machine after making any system-level environment variable changes—sometimes old settings linger until a reboot.
- Ensure no other processes are hogging GPU 0 (run
nvidia-smito check for active processes).
内容的提问来源于stack exchange,提问作者Muhammad Ahsan

