如何基于TensorFlow 2.0利用多GPU运行图像向量化任务?
Hi there! Let's break down how to get your 3 NVIDIA GPUs working efficiently for your image vectorization task. I'll address both of your questions directly.
Context Recap
You're running on Ubuntu with Python 3.7 and TensorFlow 2.0 (using v1-compatible sessions), processing images to generate vector embeddings. Your single-GPU setup takes 29 seconds for 100 images, but attempts with MirroredStrategy didn't leverage all GPUs, and manual GPU assignment caused overhead without speed gains.
1) Can we fix tf.distribute.MirroredStrategy() to use all 3 GPUs?
The issue with your initial MirroredStrategy attempt is that you loaded your frozen graph and created the session outside the strategy scope. TensorFlow can't distribute operations that were already defined on the default device (GPU 0) to other GPUs retroactively.
Here's the corrected approach to make MirroredStrategy work:
- Load your frozen graph inside the strategy scope so TensorFlow replicates the graph across all GPUs.
- Use batch processing (critical for distributed efficiency—single-image inference has too much overhead to see gains) and let the strategy distribute batches across GPUs.
import tensorflow as tf import os from PIL import Image import numpy as np # Initialize MirroredStrategy with all 3 GPUs mirrored_strategy = tf.distribute.MirroredStrategy(devices=["/gpu:0", "/gpu:1", "/gpu:2"]) with mirrored_strategy.scope(): # Load frozen graph INSIDE the strategy scope to replicate across GPUs def load_graph(frozen_graph_filename): with tf.io.gfile.GFile(frozen_graph_filename, "rb") as f: graph_def = tf.compat.v1.GraphDef() graph_def.ParseFromString(f.read()) with tf.Graph().as_default() as graph: tf.import_graph_def(graph_def, name="") return graph GRAPH = load_graph(os.path.join(settings.IMAGENET_PATH['PATH'], 'classify_image_graph_def.pb')) INPUT_TENSOR = GRAPH.get_tensor_by_name('DecodeJpeg:0') # Update to your input tensor name POOL_TENSOR = GRAPH.get_tensor_by_name('pool_3:0') # Update to your output tensor name # Configure session with GPU memory settings config = tf.compat.v1.ConfigProto() config.gpu_options.per_process_gpu_memory_fraction = 0.9 config.gpu_options.allow_growth = True SESSION = tf.compat.v1.Session(graph=GRAPH, config=config) # Wrap inference in a tf.function for distributed execution @tf.function def run_inference(image_data_batch): return SESSION.run(POOL_TENSOR, {INPUT_TENSOR: image_data_batch}) # Preprocess a single image def preprocess_image(image_path): with Image.open(image_path) as f: return f.convert('RGB') # Process images in batches (adjust batch size based on your GPU memory) batch_size = 6 # Example: 2 images per GPU (3 GPUs × 2 = 6) total_images = len(image_list) for batch_start in range(0, total_images, batch_size): batch_end = min(batch_start + batch_size, total_images) current_batch = image_list[batch_start:batch_end] # Preprocess all images in the batch batch_data = [preprocess_image(img) for img in current_batch] # Distribute batch across GPUs and run inference distributed_results = mirrored_strategy.run(run_inference, args=(batch_data,)) # Combine results from all GPUs and convert to numpy feature_sets = tf.concat(distributed_results.values, axis=0).numpy() # Save each vector for idx, feature_set in enumerate(feature_sets): feature_vector = np.squeeze(feature_set) out_filename = os.path.basename(current_batch[idx]) + ".vc" out_path = os.path.join(settings.VECTORS_DIR_PATH['PATH'], out_filename) np.savetxt(out_path, feature_vector, delimiter=',')
Key notes:
- Loading the graph inside the strategy scope ensures TensorFlow copies operations to all GPUs.
- Batch processing minimizes the overhead of distributing work across GPUs—you'll see meaningful speed gains here compared to single-image inference.
2) If MirroredStrategy isn't viable, how to use all GPUs asynchronously?
A reliable alternative is to use multi-processing, where each GPU gets its own dedicated process handling a subset of images. This avoids GPU switching overhead and lets each GPU run at full capacity independently.
Here's a working implementation:
import tensorflow as tf import os from PIL import Image import numpy as np import multiprocessing as mp def process_image_subset(gpu_id, image_subset): # Bind this process to the specified GPU os.environ["CUDA_VISIBLE_DEVICES"] = str(gpu_id) # Load graph and create session for this process (each process has its own session) def load_graph(frozen_graph_filename): with tf.io.gfile.GFile(frozen_graph_filename, "rb") as f: graph_def = tf.compat.v1.GraphDef() graph_def.ParseFromString(f.read()) with tf.Graph().as_default() as graph: tf.import_graph_def(graph_def, name="") return graph GRAPH = load_graph(os.path.join(settings.IMAGENET_PATH['PATH'], 'classify_image_graph_def.pb')) INPUT_TENSOR = GRAPH.get_tensor_by_name('DecodeJpeg:0') POOL_TENSOR = GRAPH.get_tensor_by_name('pool_3:0') config = tf.compat.v1.ConfigProto() config.gpu_options.per_process_gpu_memory_fraction = 0.9 config.gpu_options.allow_growth = True sess = tf.compat.v1.Session(graph=GRAPH, config=config) # Process each image in the subset for image in image_subset: with Image.open(image) as f: image_data = f.convert('RGB') feature_set = sess.run(POOL_TENSOR, {INPUT_TENSOR: image_data}) feature_vector = np.squeeze(feature_set) out_filename = os.path.basename(image) + ".vc" out_path = os.path.join(settings.VECTORS_DIR_PATH['PATH'], out_filename) np.savetxt(out_path, feature_vector, delimiter=',') sess.close() if __name__ == "__main__": num_gpus = 3 # Split image list into equal subsets for each GPU image_subsets = [image_list[i::num_gpus] for i in range(num_gpus)] # Start a process for each GPU processes = [] for gpu_id in range(num_gpus): proc = mp.Process(target=process_image_subset, args=(gpu_id, image_subsets[gpu_id])) processes.append(proc) proc.start() # Wait for all processes to finish for proc in processes: proc.join()
Key notes:
- Each process loads its own graph and session, so there's no interference between GPUs.
- Splitting the image list evenly ensures each GPU gets roughly the same amount of work.
- Wrap the main logic in
if __name__ == "__main__":to avoid issues with multiprocessing in Python.
内容的提问来源于stack exchange,提问作者Dmitriy Kisil

