You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于SVM的情感检测模型训练准确率100%问题排查求助

Fixing Your Sentiment Detector: Train Accuracy 100%, Dev Accuracy ~8.7%

Hey there, let's dive straight into fixing this—since you're up against a deadline, I’ll focus on the critical issues first and give you actionable fixes. The root cause here is a mix of incorrect label handling and severe overfitting, but we can resolve this step by step.

1. Core Problem: You’re Treating a Multi-Label Task as Single-Label

Looking at your label examples:

neutral,positive,neutral,unclear Tällä hetkellä Liptonin vihreä sitrushedelmätee . ...

Your labels are comma-separated multi-label combinations (each sample can have multiple sentiment tags), but you’re using LinearSVC as a single-class classifier (it assumes each sample belongs to exactly one category).

What’s happening:

  • During training, the model just memorizes every unique label string from the training set (like "neutral,positive,neutral") because these strings are almost never repeated. That’s why your train accuracy hits 100%—it’s just matching exact label strings it’s seen before.
  • On the dev set, the label strings are almost entirely new, so the model has no way to generalize, leading to that abysmal ~8.7% accuracy.

2. Step 1: Fix Your Label Format for Multi-Label Classification

First, we need to convert your comma-separated labels into a binary matrix that multi-label models can understand. Use MultiLabelBinarizer from scikit-learn:

from sklearn.preprocessing import MultiLabelBinarizer

# Split each label string into a list of individual sentiment tags
labels_list = [label.split(',') for label in labels]

# Convert to binary multi-label matrix (e.g., "neutral,positive" becomes [1,1,0,0] for 4 categories)
mlb = MultiLabelBinarizer()
binary_labels = mlb.fit_transform(labels_list)

# Now split your data with the corrected labels
train_texts, dev_texts, train_labels, dev_labels = train_test_split(context, binary_labels, test_size=0.2)

3. Step 2: Use a Multi-Label Compatible Model

LinearSVC doesn’t support multi-label classification out of the box—wrap it with OneVsRestClassifier to handle multiple labels:

from sklearn.multiclass import OneVsRestClassifier
from sklearn.svm import LinearSVC
from sklearn.metrics import accuracy_score, f1_score

# Initialize vectorizer (we'll tweak this next to reduce overfitting)
vectorizer = CountVectorizer(
    max_features=5000,  # Reduce feature count to fight overfitting
    binary=True,
    ngram_range=(1,2),
    lowercase=True,  # Normalize text to lowercase
    stop_words=None  # Add Finnish stopwords here if you can (see note below)
)

# Fit on training texts, transform both train and dev
feature_matrix_train = vectorizer.fit_transform(train_texts)
feature_matrix_dev = vectorizer.transform(dev_texts)

# Wrap LinearSVC for multi-label classification
classifier = OneVsRestClassifier(LinearSVC(C=0.1, verbose=1))
classifier.fit(feature_matrix_train, train_labels)

# Evaluate with appropriate metrics (accuracy works, but F1 is better for multi-label)
train_preds = classifier.predict(feature_matrix_train)
dev_preds = classifier.predict(feature_matrix_dev)

print("TRAIN Accuracy:", accuracy_score(train_labels, train_preds))
print("DEV Accuracy:", accuracy_score(dev_labels, dev_preds))
print("DEV Macro F1 Score:", f1_score(dev_labels, dev_preds, average='macro'))

4. Step 3: Reduce Overfitting (Your Train Accuracy Was 100%!)

Your model was severely overfitting the training data. Here are quick fixes to mitigate this:

  • Lower max_features: I reduced it from 100000 to 5000—fewer features mean less noise for the model to memorize.
  • Adjust C parameter: C=0.009 was too small (weak regularization). Try 0.1 or 1.0 to balance fitting and generalization.
  • Add stopwords: Your text is in Finnish—use Finnish stopwords to remove low-value words like "ja", "on", "että". You can get these via nltk.corpus.stopwords (download with nltk.download('stopwords') and use stop_words=stopwords.words('finnish')).
  • Switch to TF-IDF: Replace CountVectorizer with TfidfVectorizer—it weights words by their importance, reducing the impact of frequent, uninformative words.

5. Bonus: Check Your Data Size

You only have 980 total samples, which is small for a text classification task. If you can, add more training data, but for a university assignment, fixing the multi-label handling and reducing overfitting should get you a reasonable dev accuracy.


内容的提问来源于stack exchange,提问作者somerandomdude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 10:47:58