如何将未见过的测试数据传入交叉验证函数并预测标签?
Got it, let's walk through how to adapt your code to predict labels on that unseen test data. First, a quick clarification: your existing cross_val() function is for evaluating model performance on splits of your training data, not generating predictions on external test sets. We’ll keep that for validation, but add a separate workflow to train a final model and make predictions.
Step 1: Keep Cross-Validation for Performance Checks
Your cross-validation code is solid for assessing how well your model generalizes. We’ll leave that as-is to confirm performance before moving to predictions.
Step 2: Add a Training & Prediction Workflow
Here’s the updated code that combines your existing setup with the logic to predict on X2:
import pandas as pd from sklearn.model_selection import KFold, cross_val_score from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.pipeline import make_pipeline from sklearn import preprocessing from sklearn.svm import SVC # Load your datasets df = pd.read_csv('./output/csv_sanitized_16_.csv', dtype=str) X = df['description_plus'] y = df['category_id'] df_2 = pd.read_csv('./output/csv_sanitized_2.csv', dtype=str) X2 = df_2['description_plus'] # Your original cross-validation function (unchanged) def cross_val(): cv = KFold(n_splits=20) vectorizer = TfidfVectorizer(sublinear_tf=True, max_df=0.5, stop_words='english') X_train = vectorizer.fit_transform(X) clf = make_pipeline(preprocessing.StandardScaler(with_mean=False), SVC(C=1)) scores = cross_val_score(clf, X_train, y, cv=cv) print(scores) print("Accuracy: %0.2f (+/- %0.2f)" % (scores.mean(), scores.std() * 2)) # Run cross-validation to validate model performance cross_val() # Function to train on full training data and predict on unseen test data def train_and_predict_unseen(): # Fit the TF-IDF vectorizer ONLY on training data (critical!) vectorizer = TfidfVectorizer(sublinear_tf=True, max_df=0.5, stop_words='english') X_train_transformed = vectorizer.fit_transform(X) # Build and train the full pipeline on all training data clf = make_pipeline(preprocessing.StandardScaler(with_mean=False), SVC(C=1)) clf.fit(X_train_transformed, y) # Transform test data using the pre-fitted vectorizer (don't re-fit!) X2_transformed = vectorizer.transform(X2) # Generate predictions for the unseen data predicted_labels = clf.predict(X2_transformed) # Optional: Attach predictions back to your test dataframe and save df_2['predicted_category_id'] = predicted_labels df_2.to_csv('./output/test_with_predictions.csv', index=False) return predicted_labels # Run the prediction function test_predictions = train_and_predict_unseen() print("Generated predictions for unseen test data:", test_predictions)
Key Things to Note
- Don’t fit on test data: Always use
vectorizer.transform(X2)for test data, notfit_transform. This ensures we use the same vocabulary and TF-IDF weights learned from the training set—usingfit_transformon test data would leak information and invalidate your predictions. - Cross-validation ≠ prediction: Cross-validation is for estimating performance. To get real predictions on new data, you need to train a model on the entire training dataset (since cross-val uses splits of training data for testing, not external data).
- Pipeline consistency: Using the same pipeline (vectorizer → scaler → SVC) for both cross-validation and final training ensures your preprocessing steps are applied the same way everywhere.
内容的提问来源于stack exchange,提问作者neelmeg
相关产品推荐
相关产品推荐

