如何使用支持向量机(SVM)优化二分类数据集的预测效果?
Hey there! Let's figure out how to get better performance from your SVM on that 1000x1000 binary classification dataset—58% accuracy definitely has plenty of room to grow. I'll walk you through key steps to optimize your model, from data prep to choosing the right SVM settings.
Right now, you're using model.score(x,y) to check accuracy, which tests the model on the same data it was trained on. This doesn't tell you how well it will perform on unseen test data, and in your case, even the training score is low—so we need to address both underfitting potential and proper evaluation.
Instead, use cross-validation to get a reliable estimate of your model's true performance. For example, cross_val_score from sklearn will split your training data into folds, train on some, test on others, and give you an average score.
SVMs (especially linear ones) are extremely sensitive to feature scales, and high-dimensional data (1000 features!) can introduce noise and slow down training. Here's what to do:
- Standardize features: Use
StandardScalerto scale each feature to have mean 0 and variance 1. This ensures no single feature dominates the SVM's decision function. - Remove redundant features: Drop features with 0 variance (they don't add any information) or use feature selection methods like
SelectKBestto keep only the most predictive features. You could also try PCA to reduce dimensionality while retaining most of the data's variance.
You're using svm.SVC(kernel='linear')—but for high-dimensional data, LinearSVC is a better choice because it's optimized for linear kernels and runs faster. If linear isn't working, try non-linear kernels:
- RBF Kernel: The default for
SVC, great for capturing non-linear relationships in data. This is often the first non-linear kernel to try. - Polynomial Kernel: Useful if you suspect polynomial relationships, but it can be slower and more prone to overfitting than RBF.
The default parameters (C=1, linear kernel) might not be optimal for your data. Use grid or random search to find the best settings:
- For linear models (LinearSVC/SVC linear): Tune
C(regularization strength—smaller values mean more regularization, larger values mean the model will try to fit the training data closely). - For RBF kernel: Tune both
C(regularization) andgamma(controls how far the influence of a single training example reaches—small gamma = wider influence, large gamma = narrower influence).
Here's a quick example using GridSearchCV to tune an RBF SVM:
from sklearn.model_selection import GridSearchCV from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline # Create a pipeline that scales data first, then trains SVM pipe = Pipeline([ ('scaler', StandardScaler()), ('svm', svm.SVC()) ]) # Define parameter grid to search param_grid = { 'svm__kernel': ['linear', 'rbf'], 'svm__C': [0.1, 1, 10, 100], 'svm__gamma': ['scale', 'auto', 0.001, 0.01] } # Set up grid search with cross-validation grid_search = GridSearchCV(pipe, param_grid, cv=5, scoring='accuracy') grid_search.fit(x, y) # Print best parameters and score print("Best parameters:", grid_search.best_params_) print("Best cross-validation accuracy:", grid_search.best_score_)
If your dataset has way more of one class (e.g., 90% class 0 and 10% class 1), 58% accuracy might just be better than random guesses but not actually good. Check the class distribution with y.value_counts(). If it's imbalanced:
- Use
class_weight='balanced'in your SVM (this adjusts the model to give more weight to the minority class). - Try resampling techniques (oversample the minority class or undersample the majority class) using libraries like imbalanced-learn.
Here's how you could revise your original code with these fixes:
import numpy as np import pandas as pd from sklearn import svm from sklearn.preprocessing import StandardScaler from sklearn.model_selection import cross_val_score from sklearn.pipeline import Pipeline # Read data data = pd.read_csv('train.csv') x = data.loc[:, 'D_0':'D_1023'] y = data['Class'] test = pd.read_csv('test.csv') # Create pipeline with scaling and SVM pipe = Pipeline([ ('scaler', StandardScaler()), ('svm', svm.LinearSVC(class_weight='balanced', C=10)) ]) # Evaluate with cross-validation cv_scores = cross_val_score(pipe, x, y, cv=5) print("Cross-validation accuracy scores:", cv_scores) print("Mean cross-validation accuracy:", np.mean(cv_scores)) # Train on full data and predict pipe.fit(x, y) predictions = pipe.predict(test) pd.DataFrame(predictions, columns=['PredictedScore']).to_csv('prediction.csv', index=False)
内容的提问来源于stack exchange,提问作者user3019158

