You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于TensorFlow 2.0利用多GPU运行图像向量化任务?

Multi-GPU Image Vectorization in TensorFlow 2.0

Hi there! Let's break down how to get your 3 NVIDIA GPUs working efficiently for your image vectorization task. I'll address both of your questions directly.

Context Recap

You're running on Ubuntu with Python 3.7 and TensorFlow 2.0 (using v1-compatible sessions), processing images to generate vector embeddings. Your single-GPU setup takes 29 seconds for 100 images, but attempts with MirroredStrategy didn't leverage all GPUs, and manual GPU assignment caused overhead without speed gains.


1) Can we fix tf.distribute.MirroredStrategy() to use all 3 GPUs?

The issue with your initial MirroredStrategy attempt is that you loaded your frozen graph and created the session outside the strategy scope. TensorFlow can't distribute operations that were already defined on the default device (GPU 0) to other GPUs retroactively.

Here's the corrected approach to make MirroredStrategy work:

  • Load your frozen graph inside the strategy scope so TensorFlow replicates the graph across all GPUs.
  • Use batch processing (critical for distributed efficiency—single-image inference has too much overhead to see gains) and let the strategy distribute batches across GPUs.
import tensorflow as tf
import os
from PIL import Image
import numpy as np

# Initialize MirroredStrategy with all 3 GPUs
mirrored_strategy = tf.distribute.MirroredStrategy(devices=["/gpu:0", "/gpu:1", "/gpu:2"])

with mirrored_strategy.scope():
    # Load frozen graph INSIDE the strategy scope to replicate across GPUs
    def load_graph(frozen_graph_filename):
        with tf.io.gfile.GFile(frozen_graph_filename, "rb") as f:
            graph_def = tf.compat.v1.GraphDef()
            graph_def.ParseFromString(f.read())
        with tf.Graph().as_default() as graph:
            tf.import_graph_def(graph_def, name="")
        return graph

    GRAPH = load_graph(os.path.join(settings.IMAGENET_PATH['PATH'], 'classify_image_graph_def.pb'))
    INPUT_TENSOR = GRAPH.get_tensor_by_name('DecodeJpeg:0')  # Update to your input tensor name
    POOL_TENSOR = GRAPH.get_tensor_by_name('pool_3:0')  # Update to your output tensor name

    # Configure session with GPU memory settings
    config = tf.compat.v1.ConfigProto()
    config.gpu_options.per_process_gpu_memory_fraction = 0.9
    config.gpu_options.allow_growth = True
    SESSION = tf.compat.v1.Session(graph=GRAPH, config=config)

# Wrap inference in a tf.function for distributed execution
@tf.function
def run_inference(image_data_batch):
    return SESSION.run(POOL_TENSOR, {INPUT_TENSOR: image_data_batch})

# Preprocess a single image
def preprocess_image(image_path):
    with Image.open(image_path) as f:
        return f.convert('RGB')

# Process images in batches (adjust batch size based on your GPU memory)
batch_size = 6  # Example: 2 images per GPU (3 GPUs × 2 = 6)
total_images = len(image_list)

for batch_start in range(0, total_images, batch_size):
    batch_end = min(batch_start + batch_size, total_images)
    current_batch = image_list[batch_start:batch_end]
    
    # Preprocess all images in the batch
    batch_data = [preprocess_image(img) for img in current_batch]
    
    # Distribute batch across GPUs and run inference
    distributed_results = mirrored_strategy.run(run_inference, args=(batch_data,))
    
    # Combine results from all GPUs and convert to numpy
    feature_sets = tf.concat(distributed_results.values, axis=0).numpy()
    
    # Save each vector
    for idx, feature_set in enumerate(feature_sets):
        feature_vector = np.squeeze(feature_set)
        out_filename = os.path.basename(current_batch[idx]) + ".vc"
        out_path = os.path.join(settings.VECTORS_DIR_PATH['PATH'], out_filename)
        np.savetxt(out_path, feature_vector, delimiter=',')

Key notes:

  • Loading the graph inside the strategy scope ensures TensorFlow copies operations to all GPUs.
  • Batch processing minimizes the overhead of distributing work across GPUs—you'll see meaningful speed gains here compared to single-image inference.

2) If MirroredStrategy isn't viable, how to use all GPUs asynchronously?

A reliable alternative is to use multi-processing, where each GPU gets its own dedicated process handling a subset of images. This avoids GPU switching overhead and lets each GPU run at full capacity independently.

Here's a working implementation:

import tensorflow as tf
import os
from PIL import Image
import numpy as np
import multiprocessing as mp

def process_image_subset(gpu_id, image_subset):
    # Bind this process to the specified GPU
    os.environ["CUDA_VISIBLE_DEVICES"] = str(gpu_id)
    
    # Load graph and create session for this process (each process has its own session)
    def load_graph(frozen_graph_filename):
        with tf.io.gfile.GFile(frozen_graph_filename, "rb") as f:
            graph_def = tf.compat.v1.GraphDef()
            graph_def.ParseFromString(f.read())
        with tf.Graph().as_default() as graph:
            tf.import_graph_def(graph_def, name="")
        return graph

    GRAPH = load_graph(os.path.join(settings.IMAGENET_PATH['PATH'], 'classify_image_graph_def.pb'))
    INPUT_TENSOR = GRAPH.get_tensor_by_name('DecodeJpeg:0')
    POOL_TENSOR = GRAPH.get_tensor_by_name('pool_3:0')
    
    config = tf.compat.v1.ConfigProto()
    config.gpu_options.per_process_gpu_memory_fraction = 0.9
    config.gpu_options.allow_growth = True
    sess = tf.compat.v1.Session(graph=GRAPH, config=config)
    
    # Process each image in the subset
    for image in image_subset:
        with Image.open(image) as f:
            image_data = f.convert('RGB')
        feature_set = sess.run(POOL_TENSOR, {INPUT_TENSOR: image_data})
        feature_vector = np.squeeze(feature_set)
        
        out_filename = os.path.basename(image) + ".vc"
        out_path = os.path.join(settings.VECTORS_DIR_PATH['PATH'], out_filename)
        np.savetxt(out_path, feature_vector, delimiter=',')
    
    sess.close()

if __name__ == "__main__":
    num_gpus = 3
    # Split image list into equal subsets for each GPU
    image_subsets = [image_list[i::num_gpus] for i in range(num_gpus)]
    
    # Start a process for each GPU
    processes = []
    for gpu_id in range(num_gpus):
        proc = mp.Process(target=process_image_subset, args=(gpu_id, image_subsets[gpu_id]))
        processes.append(proc)
        proc.start()
    
    # Wait for all processes to finish
    for proc in processes:
        proc.join()

Key notes:

  • Each process loads its own graph and session, so there's no interference between GPUs.
  • Splitting the image list evenly ensures each GPU gets roughly the same amount of work.
  • Wrap the main logic in if __name__ == "__main__": to avoid issues with multiprocessing in Python.

内容的提问来源于stack exchange,提问作者Dmitriy Kisil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:02:00