VS Code中Jupyter Notebook训练命令异常:GPU未调用且快速终止问题排查求助
Hey there, let's break down your problem to figure out if it's a VS Code Jupyter quirk or a GPU configuration glitch. Here's what you can check step by step:
1. Verify TensorFlow can access your GPU first
This is the most critical check—if TensorFlow isn't recognizing your GPU, the training will default to CPU (which explains the fast finish and no GPU activity). Run this snippet in your VS Code Notebook:
import tensorflow as tf # List available GPUs print("Available GPUs:", tf.config.list_physical_devices('GPU')) # Check if TensorFlow was built with CUDA support print("Built with CUDA:", tf.test.is_built_with_cuda())
- If the output shows no GPUs or
Built with CUDAisFalse, your TensorFlow installation isn't properly linked to your GPU setup. Even though your CUDA/cuDNN versions match TF 2.6.2, you might have installed the CPU-only version of TensorFlow, or your system environment variables (likeLD_LIBRARY_PATH) aren't pointing to CUDA's libraries. - If it does show your RTX 3060, move on to the next checks.
2. Confirm you're using the correct Jupyter kernel
VS Code often auto-selects a default Python kernel, which might not be the virtual environment where you installed your GPU-enabled TensorFlow.
- Look at the top-right corner of your Notebook—click the kernel name (e.g., "Python 3.8") and switch to the environment you set up for deep learning. Restart the kernel and re-run your training code.
3. Check for CPU-forcing code in your training script
Double-check your training code for any lines that explicitly force TensorFlow to use the CPU. For example:
# This would force all operations to run on CPU with tf.device('/CPU:0'): # Training loop here
Since your code works on Colab, it's unlikely, but it's worth ruling out if someone modified the code for local testing and forgot to revert it.
4. Test a simple GPU-bound task
To isolate if the GPU itself is functional with TensorFlow in VS Code, run this quick matrix multiplication:
import tensorflow as tf # Force the operation to run on GPU with tf.device('/GPU:0'): large_matrix_1 = tf.random.normal((10000, 10000)) large_matrix_2 = tf.random.normal((10000, 10000)) result = tf.matmul(large_matrix_1, large_matrix_2) print("Matrix multiplication completed. Shape:", result.shape)
While this runs, check nvidia-smi—if your GPU usage spikes, your GPU is working with TensorFlow, and the issue might be specific to how your training script is executing in VS Code (e.g., accidental early termination, missing data paths, or output truncation making it seem like training finished).
Final Notes
Most likely, the problem stems from TensorFlow not accessing your GPU or using the wrong kernel—these are the most common pitfalls when moving from Colab to a local server setup. Once you confirm TensorFlow can see your GPU, the training should behave similarly to how it does in Colab.
内容的提问来源于stack exchange,提问作者Minsoo Jeong

