基于SVM的情感检测模型训练准确率100%问题排查求助
Hey there, let's dive straight into fixing this—since you're up against a deadline, I’ll focus on the critical issues first and give you actionable fixes. The root cause here is a mix of incorrect label handling and severe overfitting, but we can resolve this step by step.
1. Core Problem: You’re Treating a Multi-Label Task as Single-Label
Looking at your label examples:
neutral,positive,neutral,unclear Tällä hetkellä Liptonin vihreä sitrushedelmätee . ...
Your labels are comma-separated multi-label combinations (each sample can have multiple sentiment tags), but you’re using LinearSVC as a single-class classifier (it assumes each sample belongs to exactly one category).
What’s happening:
- During training, the model just memorizes every unique label string from the training set (like "neutral,positive,neutral") because these strings are almost never repeated. That’s why your train accuracy hits 100%—it’s just matching exact label strings it’s seen before.
- On the dev set, the label strings are almost entirely new, so the model has no way to generalize, leading to that abysmal ~8.7% accuracy.
2. Step 1: Fix Your Label Format for Multi-Label Classification
First, we need to convert your comma-separated labels into a binary matrix that multi-label models can understand. Use MultiLabelBinarizer from scikit-learn:
from sklearn.preprocessing import MultiLabelBinarizer # Split each label string into a list of individual sentiment tags labels_list = [label.split(',') for label in labels] # Convert to binary multi-label matrix (e.g., "neutral,positive" becomes [1,1,0,0] for 4 categories) mlb = MultiLabelBinarizer() binary_labels = mlb.fit_transform(labels_list) # Now split your data with the corrected labels train_texts, dev_texts, train_labels, dev_labels = train_test_split(context, binary_labels, test_size=0.2)
3. Step 2: Use a Multi-Label Compatible Model
LinearSVC doesn’t support multi-label classification out of the box—wrap it with OneVsRestClassifier to handle multiple labels:
from sklearn.multiclass import OneVsRestClassifier from sklearn.svm import LinearSVC from sklearn.metrics import accuracy_score, f1_score # Initialize vectorizer (we'll tweak this next to reduce overfitting) vectorizer = CountVectorizer( max_features=5000, # Reduce feature count to fight overfitting binary=True, ngram_range=(1,2), lowercase=True, # Normalize text to lowercase stop_words=None # Add Finnish stopwords here if you can (see note below) ) # Fit on training texts, transform both train and dev feature_matrix_train = vectorizer.fit_transform(train_texts) feature_matrix_dev = vectorizer.transform(dev_texts) # Wrap LinearSVC for multi-label classification classifier = OneVsRestClassifier(LinearSVC(C=0.1, verbose=1)) classifier.fit(feature_matrix_train, train_labels) # Evaluate with appropriate metrics (accuracy works, but F1 is better for multi-label) train_preds = classifier.predict(feature_matrix_train) dev_preds = classifier.predict(feature_matrix_dev) print("TRAIN Accuracy:", accuracy_score(train_labels, train_preds)) print("DEV Accuracy:", accuracy_score(dev_labels, dev_preds)) print("DEV Macro F1 Score:", f1_score(dev_labels, dev_preds, average='macro'))
4. Step 3: Reduce Overfitting (Your Train Accuracy Was 100%!)
Your model was severely overfitting the training data. Here are quick fixes to mitigate this:
- Lower
max_features: I reduced it from 100000 to 5000—fewer features mean less noise for the model to memorize. - Adjust
Cparameter:C=0.009was too small (weak regularization). Try0.1or1.0to balance fitting and generalization. - Add stopwords: Your text is in Finnish—use Finnish stopwords to remove low-value words like "ja", "on", "että". You can get these via
nltk.corpus.stopwords(download withnltk.download('stopwords')and usestop_words=stopwords.words('finnish')). - Switch to TF-IDF: Replace
CountVectorizerwithTfidfVectorizer—it weights words by their importance, reducing the impact of frequent, uninformative words.
5. Bonus: Check Your Data Size
You only have 980 total samples, which is small for a text classification task. If you can, add more training data, but for a university assignment, fixing the multi-label handling and reducing overfitting should get you a reasonable dev accuracy.
内容的提问来源于stack exchange,提问作者somerandomdude

