如何基于已训练完成的TensorFlow模型对给定的未见过的字符串数据预测分类类别
Hey there! Nice work getting your model trained successfully. To use it for classifying unseen string data, you need to make sure your new data matches the exact format the model was trained on—here's a step-by-step guide to do that:
1. Save/Load Your Trained Model
First, if you haven't already, save your trained model so you can reuse it later (otherwise, the model is lost once your session ends):
# Save the model right after training finishes model.save('text_classification_model.h5')
When you're ready to make predictions later, load the saved model back:
from tensorflow.keras.models import load_model model = load_model('text_classification_model.h5')
2. Preprocess Unseen Strings (Critical Step!)
Your model can't understand raw text—it expects numerical sequences just like your training data. You must use the exact same preprocessing logic and tokenizer that you used during training.
Example Preprocessing (Matching Your Training Workflow)
Suppose during training, you used this code to convert text to model-ready sequences:
from tensorflow.keras.preprocessing.text import Tokenizer from tensorflow.keras.preprocessing.sequence import pad_sequences # These values must match what you defined for training! max_features = 10000 # Replace with your actual max_features maxlen = 200 # Replace with your actual maxlen # Fit tokenizer on your training text (you did this during training) tokenizer = Tokenizer(num_words=max_features) tokenizer.fit_on_texts(your_training_texts)
For new unseen strings, never re-fit the tokenizer—use the saved one to convert and pad your data:
# Save the tokenizer for later use (do this after training) import pickle with open('text_tokenizer.pickle', 'wb') as handle: pickle.dump(tokenizer, handle, protocol=pickle.HIGHEST_PROTOCOL) # Load the saved tokenizer when predicting with open('text_tokenizer.pickle', 'rb') as handle: tokenizer = pickle.load(handle) # Your new unseen strings to classify new_unseen_strings = [ "This is a brand new text we need to categorize.", "Another example string the model hasn't seen before." ] # Convert raw text to numerical sequences new_sequences = tokenizer.texts_to_sequences(new_unseen_strings) # Pad sequences to the fixed length used in training new_padded_sequences = pad_sequences(new_sequences, maxlen=maxlen)
3. Run Predictions & Interpret Results
Now feed the preprocessed data into your model to get predictions, then map the output to readable class labels:
# Get predicted probabilities for each class predictions = model.predict(new_padded_sequences) # Each prediction is an array of 3 probabilities (since your model outputs 3 classes) # Grab the index of the highest probability to get the predicted class import numpy as np predicted_class_indices = np.argmax(predictions, axis=1) # Map indices to your actual class names (replace with your own labels) class_labels = ["Class A", "Class B", "Class C"] predicted_classes = [class_labels[idx] for idx in predicted_class_indices] # Print out the results for text, cls, probs in zip(new_unseen_strings, predicted_classes, predictions): print(f"Text: {text}\nPredicted Class: {cls}\nClass Probabilities: {probs.round(4)}\n")
Key Reminders
- Always use the same tokenizer and preprocessing parameters (
max_features,maxlen) as training—using new values will lead to incorrect predictions. - If you forgot to save your tokenizer, you'll need to re-run your training preprocessing code (without fitting on new data) to retrieve the correct tokenizer.
内容的提问来源于stack exchange,提问作者Abdullah Zia

