You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何保存TfidfVectorizer并复用?以及如何保存含该组件的文本分类器?

Got it, let's walk through this step by step. First, your current setup for TfidfVectorizer looks good—you're using a custom tokenizer, filtering English stop words, and limiting features to 1100, which is a solid starting point for text classification.

Now, when it comes to saving the vectorizer so you can reuse it later (for example, predicting on new text without re-fitting the vectorizer from scratch), we have two reliable options using Python's built-in or scikit-learn-friendly libraries.

Saving the TfidfVectorizer

joblib is optimized for objects that store large numerical arrays (like the vocabulary and IDF weights in TfidfVectorizer), so it's the go-to choice here.

First, import the library:

import joblib

After you've fitted your vectorizer (you already did this with fit_transform), save it to a file:

# Save the trained vectorizer to a file
joblib.dump(vectorizer, 'tfidf_vectorizer.joblib')

Option 2: Use pickle

If you prefer using Python's standard pickle module, that works too:

import pickle

# Save the vectorizer via pickle
with open('tfidf_vectorizer.pkl', 'wb') as file:
    pickle.dump(vectorizer, file)

Loading and Reusing the Vectorizer

When you need to apply the vectorizer to new text data (like for predictions), load it first—and make sure to use transform() instead of fit_transform() (you don't want to retrain the vectorizer on new data, which would break consistency with your original model):

# Load with joblib
loaded_vectorizer = joblib.load('tfidf_vectorizer.joblib')

# OR load with pickle
with open('tfidf_vectorizer.pkl', 'rb') as file:
    loaded_vectorizer = pickle.load(file)

# Transform new text to match the original feature space
new_text_samples = ["This is a new text sample to classify", "Another piece of unlabeled text"]
new_features = loaded_vectorizer.transform(new_text_samples)

# Now pass these features to your trained classifier (calibrated_svc)
predictions = calibrated_svc.predict(new_features)

A Quick Best Practice Note for Your Existing Code

I noticed you're fitting the vectorizer on both training and test data combined:

corpus_data_features = vectorizer.fit_transform(train_data_df.Text.tolist() + test_data_df.Text.tolist())

This is a common pitfall called data leakage—you're letting the model see test data during the preprocessing step, which can lead to overoptimistic performance metrics. Instead, you should only fit the vectorizer on training data, then transform both train and test sets:

# Fit ONLY on training data
vectorizer.fit(train_data_df.Text.tolist())

# Transform both training and test data using the fitted vectorizer
train_features = vectorizer.transform(train_data_df.Text.tolist())
test_features = vectorizer.transform(test_data_df.Text.tolist())

# Convert to ndarray if needed
train_features_nd = train_features.toarray()
test_features_nd = test_features.toarray()

# Train your model on the training features
calibrated_svc.fit(X=train_features_nd, y=train_data_df.YOUR_LABEL_COLUMN.tolist())

This keeps your pipeline consistent with real-world usage, where you won't have access to test data when deploying the model.

内容的提问来源于stack exchange,提问作者mrmrn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:31:13