在OSX系统下利用AMD Radeon eGPU(Metal)加速Python图像/掩码转numpy数组
Great question! Let's break this down and figure out how to optimize your image preprocessing step, even if we can't directly offload every part to your Vega 56 eGPU.
First, let's clarify why your current code isn't using the eGPU:
- The
image.load_img(from Keras/PIL) andimage.img_to_arrayoperations are CPU-bound — these libraries don't have support for Metal/GPU acceleration, so they'll always run on your MacBook's CPU. - Your serial loop processing images one-by-one is also inefficient, as it doesn't leverage parallelism either on CPU or GPU.
That said, there are ways to speed up this process, and even offload parts of it to your eGPU indirectly. Here are the best approaches:
1. Use TensorFlow's tf.data.Dataset with Metal Acceleration (Recommended)
TensorFlow natively supports Metal on Mac, so using tf.data to handle your preprocessing will let you parallelize operations and offload image resizing/normalization to your eGPU. Plus, it eliminates the need to convert everything to a numpy array upfront, reducing data copy overhead between CPU and GPU.
Here's how to refactor your code:
import tensorflow as tf from pathlib import Path def preprocess_image(file_path, label): # Read image file (still CPU-bound, but parallelized) img = tf.io.read_file(file_path) # Decode JPEG (can be accelerated by Metal) img = tf.image.decode_jpeg(img, channels=channels) # Resize image (Metal-accelerated) img = tf.image.resize(img, [img_height, img_width]) # Normalize and cast to float (GPU-accelerated) img = tf.cast(img, tf.float32) / 255.0 return img, label # Prepare file paths and labels from your DataFrame file_paths = [str(Path(train_path) / row["ImageId"]) for _, row in train_df.iterrows()] labels = train_df["Category"].tolist() # Build parallelized dataset train_dataset = tf.data.Dataset.from_tensor_slices((file_paths, labels)) # Map preprocessing with automatic parallelism train_dataset = train_dataset.map(preprocess_image, num_parallel_calls=tf.data.AUTOTUNE) # Batch for training (adjust batch size to fit your eGPU memory) train_dataset = train_dataset.batch(32)
You can then feed this dataset directly to your Keras model's fit() method — no need for numpy arrays. Most of the heavy lifting (resizing, normalization) will run on your eGPU, and the file reading is parallelized across CPU cores to keep up.
2. Parallelize CPU Processing with Multiprocessing
If you prefer sticking with PIL/Keras image utilities, you can speed up the serial loop using Python's multiprocessing to leverage your MacBook's CPU cores. This won't use the eGPU, but it will drastically reduce preprocessing time so your GPU doesn't sit idle waiting for data.
Example code:
from concurrent.futures import ProcessPoolExecutor from keras.preprocessing import image from pathlib import Path def process_image(row): path = Path(train_path) / row["ImageId"] img = image.load_img(path, target_size=(img_height, img_width, channels)) img = image.img_to_array(img) return img / 255 # Use all available CPU cores to process images in parallel with ProcessPoolExecutor() as executor: train_image = list(executor.map(process_image, [row for _, row in train_df.iterrows()])) train_img = np.array(train_image)
3. Offload Post-Loading Operations to PlaidML GPU
While you can't load images directly to the eGPU, you can convert the numpy arrays to PlaidML tensors immediately after loading, so operations like normalization run on the GPU instead of the CPU. This also avoids copying data back and forth between CPU and GPU during training.
Example code:
import plaidml.keras plaidml.keras.install_backend() from keras.preprocessing import image from pathlib import Path def process_image_plaid(row): path = Path(train_path) / row["ImageId"] img = image.load_img(path, target_size=(img_height, img_width, channels)) img_array = image.img_to_array(img) # Convert to PlaidML tensor (stored in eGPU memory) plaid_tensor = plaidml.tensor.as_tensor(img_array) # Normalize directly on the GPU plaid_tensor = plaid_tensor / 255 return plaid_tensor # Process images and stack into a batch tensor train_tensors = [process_image_plaid(row) for _, row in train_df.iterrows()] train_plaid_batch = plaidml.tensor.stack(train_tensors)
You can then use this PlaidML batch tensor directly in your model training.
Key Takeaway
You can't accelerate the actual image.load_img or img_to_array steps with your eGPU (since those libraries don't support Metal), but you can:
- Use TensorFlow's
tf.datato offload most preprocessing work to the eGPU - Parallelize CPU processing to eliminate bottlenecks
- Use PlaidML tensors to handle post-loading operations on the eGPU and reduce data transfer overhead
内容的提问来源于stack exchange,提问作者CCC

