TensorFlow中欧氏距离计算未占用GPU的原因及优化方案
Why Your TensorFlow Euclidean Distance Code Is Slow & GPU Idle (And How to Fix It)
First, let’s break down exactly why your current code is dragging its feet and leaving your GPU unused:
Key Issues With Your Original Code
- Repeated graph construction + tiny
sess.run()calls: You’re creating new TensorFlow placeholders and operations inside nested loops. Everysess.run()call adds overhead for graph setup, CPU-GPU data transfer, and running a single tiny calculation. GPUs excel at batch parallelism—this approach gives them nothing substantial to process, so they sit idle while the CPU wastes time on loop logistics. - NumPy array bottleneck: Initializing a NumPy array
Dand filling it one value at a time is a CPU-bound task that disconnects your GPU computations from the final result, adding unnecessary slowdowns. - No batch processing: Your code calculates distances one pair at a time, completely ignoring TensorFlow’s core strength of parallelizing computations across multiple inputs simultaneously.
Optimized Solution
The fix focuses on leveraging TensorFlow’s vectorized batch processing, minimizing graph overhead, and keeping computations on the GPU as much as possible. Here’s how to rewrite it:
For TensorFlow 1.x (Session-Based)
import tensorflow as tf import numpy as np def compute_distance_matrix(ir_set): # Convert your list of arrays into a single TensorFlow tensor # Shape becomes: (m, k+1, *dim_image) where m=len(ir_set), k=qtd_images-1 ir_tensor = tf.convert_to_tensor(ir_set, dtype=tf.float32) # Split into k comparison images and the target xr_j (last image in each column) k_images = ir_tensor[:, :-1, :] # Shape: (m, k, *dim_image) xr_j = ir_tensor[:, -1:, :] # Shape: (m, 1, *dim_image) (broadcastable to k_images shape) # Compute Euclidean distance in bulk using broadcasting squared_diff = tf.square(k_images - xr_j) # Sum over all image dimensions (adjust axes if your images have more/less dimensions) sum_squared = tf.reduce_sum(squared_diff, axis=[2, 3]) distances = tf.sqrt(sum_squared) # Transpose to get the desired (k, m) shape D = tf.transpose(distances) # Run the computation once (no loops!) with tf.Session() as sess: D_np = sess.run(D) print(f'\nAll distances computed. Matrix D shape: {D_np.shape}') return D_np
For TensorFlow 2.x (Eager Execution)
TF2.x simplifies this even more with eager execution enabled by default—no sessions required:
import tensorflow as tf import numpy as np def compute_distance_matrix(ir_set): ir_tensor = tf.convert_to_tensor(ir_set, dtype=tf.float32) # Shape: (m, k+1, *dim_image) k_images = ir_tensor[:, :-1, :] xr_j = ir_tensor[:, -1:, :] # Handle any number of image dimensions dynamically image_dims = tuple(range(2, ir_tensor.ndim)) squared_diff = tf.square(k_images - xr_j) sum_squared = tf.reduce_sum(squared_diff, axis=image_dims) distances = tf.sqrt(sum_squared) # Transpose and convert to NumPy array directly D = tf.transpose(distances).numpy() print(f'\nAll distances computed. Matrix D shape: {D.shape}') return D
What This Fix Does
- Single graph construction: We build the computation graph once, not thousands of times in loops—eliminating massive setup overhead.
- Full GPU utilization: Using TensorFlow’s broadcasting, we compute all (m*k) distances in parallel. This keeps your GPU’s cores busy with meaningful, parallelizable work.
- Minimal data transfer: Data moves from CPU to GPU once, not once per distance calculation—cutting down on expensive cross-device IO.
- No slow NumPy loops: We directly convert the final GPU-computed tensor to a NumPy array, avoiding CPU-side element-wise assignments that kill performance.
Quick Extra Tips
- Ensure all input data uses
float32to match TensorFlow’s GPU-friendly default dtype (avoids on-the-fly type conversions). - If
ir_setis already a NumPy array (not a list), pass it directly totf.convert_to_tensorto skip list-to-array conversion overhead.
内容的提问来源于stack exchange,提问作者Gabriel Machado
相关产品推荐
相关产品推荐

