如何保存TfidfVectorizer并复用?以及如何保存含该组件的文本分类器?
Got it, let's walk through this step by step. First, your current setup for TfidfVectorizer looks good—you're using a custom tokenizer, filtering English stop words, and limiting features to 1100, which is a solid starting point for text classification.
Now, when it comes to saving the vectorizer so you can reuse it later (for example, predicting on new text without re-fitting the vectorizer from scratch), we have two reliable options using Python's built-in or scikit-learn-friendly libraries.
Saving the TfidfVectorizer
Option 1: Use joblib (Recommended for Scikit-Learn Objects)
joblib is optimized for objects that store large numerical arrays (like the vocabulary and IDF weights in TfidfVectorizer), so it's the go-to choice here.
First, import the library:
import joblib
After you've fitted your vectorizer (you already did this with fit_transform), save it to a file:
# Save the trained vectorizer to a file joblib.dump(vectorizer, 'tfidf_vectorizer.joblib')
Option 2: Use pickle
If you prefer using Python's standard pickle module, that works too:
import pickle # Save the vectorizer via pickle with open('tfidf_vectorizer.pkl', 'wb') as file: pickle.dump(vectorizer, file)
Loading and Reusing the Vectorizer
When you need to apply the vectorizer to new text data (like for predictions), load it first—and make sure to use transform() instead of fit_transform() (you don't want to retrain the vectorizer on new data, which would break consistency with your original model):
# Load with joblib loaded_vectorizer = joblib.load('tfidf_vectorizer.joblib') # OR load with pickle with open('tfidf_vectorizer.pkl', 'rb') as file: loaded_vectorizer = pickle.load(file) # Transform new text to match the original feature space new_text_samples = ["This is a new text sample to classify", "Another piece of unlabeled text"] new_features = loaded_vectorizer.transform(new_text_samples) # Now pass these features to your trained classifier (calibrated_svc) predictions = calibrated_svc.predict(new_features)
A Quick Best Practice Note for Your Existing Code
I noticed you're fitting the vectorizer on both training and test data combined:
corpus_data_features = vectorizer.fit_transform(train_data_df.Text.tolist() + test_data_df.Text.tolist())
This is a common pitfall called data leakage—you're letting the model see test data during the preprocessing step, which can lead to overoptimistic performance metrics. Instead, you should only fit the vectorizer on training data, then transform both train and test sets:
# Fit ONLY on training data vectorizer.fit(train_data_df.Text.tolist()) # Transform both training and test data using the fitted vectorizer train_features = vectorizer.transform(train_data_df.Text.tolist()) test_features = vectorizer.transform(test_data_df.Text.tolist()) # Convert to ndarray if needed train_features_nd = train_features.toarray() test_features_nd = test_features.toarray() # Train your model on the training features calibrated_svc.fit(X=train_features_nd, y=train_data_df.YOUR_LABEL_COLUMN.tolist())
This keeps your pipeline consistent with real-world usage, where you won't have access to test data when deploying the model.
内容的提问来源于stack exchange,提问作者mrmrn

