如何用训练好的TensorFlow-Keras模型识别实拍手势图片?
Hey there! Awesome work getting your model trained and saved. To predict on your thumbs_up.jpg image, you absolutely need to use the handrecognition_model.h5 file you created. Let's walk through each step with code that matches your original model setup:
1. Import Required Libraries & Load the Saved Model
First, we'll load the model you saved. Make sure you have the same Keras/TensorFlow versions installed as when you trained the model to avoid compatibility issues.
# Import necessary libraries from keras.models import load_model from PIL import Image import numpy as np # Load your trained model model = load_model('handrecognition_model.h5')
2. Preprocess Your Input Image
Your model was trained on images with the shape (120, 320, 1) (height, width, grayscale channel). We need to convert your thumbs_up.jpg to match this exact format:
# Load the image using PIL img = Image.open('thumbs_up.jpg') # Convert to grayscale (matches your model's single-channel input) img_gray = img.convert('L') # Resize to match the model's input dimensions (width=320, height=120) # Note: PIL uses (width, height) order for resize, which aligns with our (120,320) input shape img_resized = img_gray.resize((320, 120)) # Convert image to numpy array and normalize pixel values to 0-1 (standard for image models) img_array = np.array(img_resized) / 255.0 # Add a batch dimension (model expects input in shape (batch_size, height, width, channels)) img_input = np.expand_dims(img_array, axis=0) # Add the single channel dimension to match the model's input shape img_input = np.expand_dims(img_input, axis=-1)
Quick note: If you normalized your training data differently (e.g., didn't divide by 255), adjust this step to match exactly what you did during training. Consistency here is critical for accurate predictions!
3. Run Prediction & Interpret the Result
Now we can feed the preprocessed image into the model and get the prediction:
# Get prediction probabilities for each of the 10 classes predictions = model.predict(img_input, verbose=0) # Find the class index with the highest probability predicted_class = np.argmax(predictions[0]) # Map the class index to your actual gesture label (fill this in based on your training data!) # Example mapping—adjust to match your dataset's labels: gesture_labels = { 0: "Thumbs Up", 1: "Peace Sign", 2: "Fist", 3: "Open Palm", # Add all 10 of your gesture labels here } # Print the final result print(f"Predicted Gesture: {gesture_labels[predicted_class]}") print(f"Confidence Score: {round(predictions[0][predicted_class] * 100, 2)}%") print(f"All Class Probabilities: {predictions[0]}")
Quick Troubleshooting Tips:
- Label Mapping: Double-check that your
gesture_labelsdictionary matches how you labeled your training data (which index corresponds to which gesture). - Image Orientation: If predictions are off, confirm the resized image isn't rotated/flipped compared to your training dataset. Use
img.rotate()to adjust orientation if needed. - Normalization: If you skipped normalization during training, remove the
/ 255.0line in the preprocessing step.
内容的提问来源于stack exchange,提问作者Maria Rona

