SVM测试集评分问题:使用scikit-learn时遭遇ValueError
Hey there, I spotted the issue in your code that's triggering that ValueError—let's get it sorted out!
The Root Cause
You made a small but critical mistake when defining your test set labels:
y_test_1 = dataset[:,15:16] # Wrong! You're using the training dataset here
Instead of pulling labels from your test dataset (test_dataset_1), you're accidentally using the training dataset (dataset). This means your test features (X_test_1) and test labels (y_test_1) almost certainly have different numbers of samples, which scikit-learn can't handle when calculating the model score.
Corrected Code
Here's the fixed version of your full code:
from sklearn.svm import SVC import numpy dataset = numpy.loadtxt("training.txt", delimiter="\t") X = dataset[:,0:15] y = dataset[:,15:16] y = y.ravel() test_dataset_1 = numpy.loadtxt("test_14-15.txt", delimiter="\t") X_test_1 = test_dataset_1[:,0:15] y_test_1 = test_dataset_1[:,15:16] # Fixed: use test_dataset_1 here y_test_1 = y_test_1.ravel() model = SVC(kernel='linear', C=75) model.fit(X, y) score_1 = model.score(X_test_1, y_test_1)
Quick Debug Tip for Future
To avoid this kind of mismatch issue down the line, add quick shape checks after loading your data:
print(f"Test features shape: {X_test_1.shape}") print(f"Test labels shape: {y_test_1.shape}")
If the first number (sample count) doesn't match between the two, you'll immediately know there's a problem with how you're loading your labels.
内容的提问来源于stack exchange,提问作者AnDeep

