预训练Keras模型跨设备部署咨询:CPU存变量与多GPU拆分
Hey there, let's break down your two questions with practical, actionable steps—dealing with large pre-trained models that don't fit into a single GPU is super common, so I’ve got some hands-on tips for you.
1. Storing all tf.Variables on CPU while running computations on GPU (for pre-trained Keras models)
The key here is to initialize the model's variables on the CPU first, then run the forward/backward passes on the GPU. TensorFlow handles cross-device data transfer automatically, though you should be aware of potential overhead from moving data between CPU and GPU.
Here's how to do it:
- Wrap the model loading code in a
tf.device("/CPU:0")context. This ensures all trainable and non-trainable variables (like pre-trained weights) are stored on the CPU. - When running inference or training, move your input data to the GPU and execute the model within a
tf.device("/GPU:0")context. The computations will run on the GPU, pulling variables from the CPU as needed.
Example code:
import tensorflow as tf from tensorflow.keras.applications import ResNet50 # Load the pre-trained model entirely on CPU to store all variables there with tf.device("/CPU:0"): model = ResNet50(weights="imagenet", include_top=True) # Prepare input data (you can generate or load real data here) input_data = tf.random.uniform((1, 224, 224, 3)) # Run computations on GPU with tf.device("/GPU:0"): # Move input to GPU first (optional, but explicit is better) gpu_input = input_data.gpu() predictions = model(gpu_input)
Note: This approach will add some latency since variables are fetched from CPU to GPU for each computation. It’s a trade-off between fitting the model into memory and speed.
2. Splitting pre-trained model layers across GPUs on different machines
This is called model parallelism across multi-worker devices, and it requires setting up a TensorFlow cluster to coordinate between machines. You’ll need to manually split the model into segments and assign each segment to a GPU on a different worker machine.
Step 1: Set up the cluster
First, define your cluster configuration (replace the IPs and ports with your actual machine details). Each machine will run a server instance to join the cluster.
On Machine 1 (IP: 192.168.0.1, task index 0):
import tensorflow as tf # Define the cluster spec cluster_spec = tf.train.ClusterSpec({ "worker": ["192.168.0.1:2222", "192.168.0.2:2222"] }) # Start the server server = tf.distribute.Server(cluster_spec, job_name="worker", task_index=0) server.join() # Keep the server running
On Machine 2 (IP: 192.168.0.2, task index 1):
import tensorflow as tf cluster_spec = tf.train.ClusterSpec({ "worker": ["192.168.0.1:2222", "192.168.0.2:2222"] }) server = tf.distribute.Server(cluster_spec, job_name="worker", task_index=1) server.join()
Step 2: Split and assign model layers
On a client machine (or one of the worker machines), load the pre-trained model and split its layers across the cluster’s GPUs. Use tf.device() with the full device path (e.g., /job:worker/task:0/GPU:0) to assign each segment.
Example code:
import tensorflow as tf from tensorflow.keras.applications import ResNet50 # Connect to the existing cluster cluster_resolver = tf.distribute.cluster_resolver.SimpleClusterResolver( cluster_spec=tf.train.ClusterSpec({ "worker": ["192.168.0.1:2222", "192.168.0.2:2222"] }) ) # Load the base model first (we'll split it later) with tf.device("/CPU:0"): base_model = ResNet50(weights="imagenet", include_top=False, input_shape=(224, 224, 3)) # Assign lower layers to Machine 1's GPU with tf.device("/job:worker/task:0/GPU:0"): # Extract layers up to a specific point (use model.summary() to find layer names) lower_model = tf.keras.Model( inputs=base_model.input, outputs=base_model.get_layer("conv4_block6_out").output ) # Assign upper layers to Machine 2's GPU with tf.device("/job:worker/task:1/GPU:0"): # Create input matching the output of the lower model upper_input = tf.keras.Input(shape=lower_model.output_shape[1:]) x = upper_input # Add layers starting from the next layer after our split point start_layer_idx = base_model.layers.index(base_model.get_layer("conv4_block6_out")) + 1 for layer in base_model.layers[start_layer_idx:]: x = layer(x) upper_model = tf.keras.Model(inputs=upper_input, outputs=x) # Combine into a single model full_model = tf.keras.Model(inputs=lower_model.input, outputs=upper_model(lower_model.output)) # Test with sample input input_data = tf.random.uniform((1, 224, 224, 3)) predictions = full_model(input_data)
Key Notes:
- Use
model.summary()to get the exact names and output shapes of layers, so you can split the model cleanly. - Model parallelism across machines has significant network overhead, so it’s best suited for models where the split point minimizes data transfer between layers.
- Ensure all machines have the same TensorFlow version and network access (firewalls open for the cluster port).
内容的提问来源于stack exchange,提问作者Mohamed Amine Ouali

