You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用支持向量机(SVM)优化二分类数据集的预测效果?

Hey there! Let's figure out how to get better performance from your SVM on that 1000x1000 binary classification dataset—58% accuracy definitely has plenty of room to grow. I'll walk you through key steps to optimize your model, from data prep to choosing the right SVM settings.

1. Fix Your Evaluation First (You're Testing on Training Data!)

Right now, you're using model.score(x,y) to check accuracy, which tests the model on the same data it was trained on. This doesn't tell you how well it will perform on unseen test data, and in your case, even the training score is low—so we need to address both underfitting potential and proper evaluation.

Instead, use cross-validation to get a reliable estimate of your model's true performance. For example, cross_val_score from sklearn will split your training data into folds, train on some, test on others, and give you an average score.

2. Preprocess Your High-Dimensional Data

SVMs (especially linear ones) are extremely sensitive to feature scales, and high-dimensional data (1000 features!) can introduce noise and slow down training. Here's what to do:

  • Standardize features: Use StandardScaler to scale each feature to have mean 0 and variance 1. This ensures no single feature dominates the SVM's decision function.
  • Remove redundant features: Drop features with 0 variance (they don't add any information) or use feature selection methods like SelectKBest to keep only the most predictive features. You could also try PCA to reduce dimensionality while retaining most of the data's variance.
3. Pick the Right SVM Class & Kernel

You're using svm.SVC(kernel='linear')—but for high-dimensional data, LinearSVC is a better choice because it's optimized for linear kernels and runs faster. If linear isn't working, try non-linear kernels:

  • RBF Kernel: The default for SVC, great for capturing non-linear relationships in data. This is often the first non-linear kernel to try.
  • Polynomial Kernel: Useful if you suspect polynomial relationships, but it can be slower and more prone to overfitting than RBF.
4. Tune Hyperparameters (Don't Guess!)

The default parameters (C=1, linear kernel) might not be optimal for your data. Use grid or random search to find the best settings:

  • For linear models (LinearSVC/SVC linear): Tune C (regularization strength—smaller values mean more regularization, larger values mean the model will try to fit the training data closely).
  • For RBF kernel: Tune both C (regularization) and gamma (controls how far the influence of a single training example reaches—small gamma = wider influence, large gamma = narrower influence).

Here's a quick example using GridSearchCV to tune an RBF SVM:

from sklearn.model_selection import GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline

# Create a pipeline that scales data first, then trains SVM
pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('svm', svm.SVC())
])

# Define parameter grid to search
param_grid = {
    'svm__kernel': ['linear', 'rbf'],
    'svm__C': [0.1, 1, 10, 100],
    'svm__gamma': ['scale', 'auto', 0.001, 0.01]
}

# Set up grid search with cross-validation
grid_search = GridSearchCV(pipe, param_grid, cv=5, scoring='accuracy')
grid_search.fit(x, y)

# Print best parameters and score
print("Best parameters:", grid_search.best_params_)
print("Best cross-validation accuracy:", grid_search.best_score_)
5. Check for Class Imbalance

If your dataset has way more of one class (e.g., 90% class 0 and 10% class 1), 58% accuracy might just be better than random guesses but not actually good. Check the class distribution with y.value_counts(). If it's imbalanced:

  • Use class_weight='balanced' in your SVM (this adjusts the model to give more weight to the minority class).
  • Try resampling techniques (oversample the minority class or undersample the majority class) using libraries like imbalanced-learn.
Example of Improved Code

Here's how you could revise your original code with these fixes:

import numpy as np
import pandas as pd
from sklearn import svm
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import Pipeline

# Read data
data = pd.read_csv('train.csv')
x = data.loc[:, 'D_0':'D_1023']
y = data['Class']
test = pd.read_csv('test.csv')

# Create pipeline with scaling and SVM
pipe = Pipeline([
    ('scaler', StandardScaler()),
    ('svm', svm.LinearSVC(class_weight='balanced', C=10))
])

# Evaluate with cross-validation
cv_scores = cross_val_score(pipe, x, y, cv=5)
print("Cross-validation accuracy scores:", cv_scores)
print("Mean cross-validation accuracy:", np.mean(cv_scores))

# Train on full data and predict
pipe.fit(x, y)
predictions = pipe.predict(test)
pd.DataFrame(predictions, columns=['PredictedScore']).to_csv('prediction.csv', index=False)

内容的提问来源于stack exchange,提问作者user3019158

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:51:59