You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将未见过的测试数据传入交叉验证函数并预测标签?

Got it, let's walk through how to adapt your code to predict labels on that unseen test data. First, a quick clarification: your existing cross_val() function is for evaluating model performance on splits of your training data, not generating predictions on external test sets. We’ll keep that for validation, but add a separate workflow to train a final model and make predictions.

Step 1: Keep Cross-Validation for Performance Checks

Your cross-validation code is solid for assessing how well your model generalizes. We’ll leave that as-is to confirm performance before moving to predictions.

Step 2: Add a Training & Prediction Workflow

Here’s the updated code that combines your existing setup with the logic to predict on X2:

import pandas as pd
from sklearn.model_selection import KFold, cross_val_score
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.pipeline import make_pipeline
from sklearn import preprocessing
from sklearn.svm import SVC

# Load your datasets
df = pd.read_csv('./output/csv_sanitized_16_.csv', dtype=str)
X = df['description_plus']
y = df['category_id']

df_2 = pd.read_csv('./output/csv_sanitized_2.csv', dtype=str)
X2 = df_2['description_plus']

# Your original cross-validation function (unchanged)
def cross_val():
    cv = KFold(n_splits=20)
    vectorizer = TfidfVectorizer(sublinear_tf=True, max_df=0.5, stop_words='english')
    X_train = vectorizer.fit_transform(X)
    clf = make_pipeline(preprocessing.StandardScaler(with_mean=False), SVC(C=1))
    scores = cross_val_score(clf, X_train, y, cv=cv)
    print(scores)
    print("Accuracy: %0.2f (+/- %0.2f)" % (scores.mean(), scores.std() * 2))

# Run cross-validation to validate model performance
cross_val()

# Function to train on full training data and predict on unseen test data
def train_and_predict_unseen():
    # Fit the TF-IDF vectorizer ONLY on training data (critical!)
    vectorizer = TfidfVectorizer(sublinear_tf=True, max_df=0.5, stop_words='english')
    X_train_transformed = vectorizer.fit_transform(X)
    
    # Build and train the full pipeline on all training data
    clf = make_pipeline(preprocessing.StandardScaler(with_mean=False), SVC(C=1))
    clf.fit(X_train_transformed, y)
    
    # Transform test data using the pre-fitted vectorizer (don't re-fit!)
    X2_transformed = vectorizer.transform(X2)
    
    # Generate predictions for the unseen data
    predicted_labels = clf.predict(X2_transformed)
    
    # Optional: Attach predictions back to your test dataframe and save
    df_2['predicted_category_id'] = predicted_labels
    df_2.to_csv('./output/test_with_predictions.csv', index=False)
    
    return predicted_labels

# Run the prediction function
test_predictions = train_and_predict_unseen()
print("Generated predictions for unseen test data:", test_predictions)

Key Things to Note

  • Don’t fit on test data: Always use vectorizer.transform(X2) for test data, not fit_transform. This ensures we use the same vocabulary and TF-IDF weights learned from the training set—using fit_transform on test data would leak information and invalidate your predictions.
  • Cross-validation ≠ prediction: Cross-validation is for estimating performance. To get real predictions on new data, you need to train a model on the entire training dataset (since cross-val uses splits of training data for testing, not external data).
  • Pipeline consistency: Using the same pipeline (vectorizer → scaler → SVC) for both cross-validation and final training ensures your preprocessing steps are applied the same way everywhere.

内容的提问来源于stack exchange,提问作者neelmeg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 20:59:09