基于字幕的电影评级分类准确率不佳,寻求Python优化方案
Hey there! Let’s break down your movie rating classification problem step by step—first identifying potential missteps in your current process, then diving into actionable fixes to lift those accuracy numbers.
First, Let’s Flag Clear Issues in Your Current Approach
Incorrect Clustering Evaluation
You used accuracy to assess KMeans performance, which is a mistake. Clustering algorithms assign arbitrary cluster labels that don’t map directly to your rating categories (e.g., KMeans might group R and PG-13 films together, but you can’t just force those clusters to match your true labels). For clustering, use metrics like Adjusted Rand Index (ARI) or Normalized Mutual Information (NMI) instead—these measure how well clusters align with true labels without assuming direct one-to-one mapping.
Overly High-Dimensional Features for a Small Dataset
Your dataset only has 130 films, but you’re using 5000 features (1000 per category from TF-IDF). This is a classic case of the curse of dimensionality: with more features than samples, models struggle to learn meaningful patterns, leading to either overfitting (like your Logistic Regression) or poor generalization (like Naive Bayes).TF-IDF Implementation Misstep
Generating TF-IDF features per category and merging them is not standard practice. You should fit a single TF-IDF vectorizer on your entire dataset, then select top N features (e.g., 500-1000 total) usingmax_features. Per-category TF-IDF creates redundant, noisy features that don’t help models learn cross-category patterns.Severe Overfitting in Logistic Regression
A training accuracy of 0.96 vs. test accuracy of 0.53 means your model has memorized training data instead of learning generalizable patterns. This is almost certainly due to high feature dimensionality and lack of regularization.
Actionable Fixes to Boost Accuracy
1. Fix Dataset & Class Imbalance Issues
- Expand Your Dataset: 130 films split across 5 categories averages just 26 samples per class—way too small for reliable text classification. Look for public datasets with movie subtitles + MPAA ratings (e.g., scrape IMDb for ratings paired with open-source subtitle data).
- Address Class Imbalance: If some ratings (like G or NR) have far fewer samples, use:
class_weight='balanced'parameter in models like Logistic Regression or SVC to give more weight to underrepresented classes.- SMOTE (carefully—text data is sparse, so use text-specific oversampling tools) or undersampling of overrepresented classes.
2. Optimize Feature Engineering
- Simplify TF-IDF:
- Fit one TF-IDF vectorizer on all subtitles, set
ngram_range=(1,2)to capture meaningful phrases (e.g., "graphic violence" or "sexual content" are far more predictive than single words like "violence"). - Reduce
max_featuresto 500-800 total features—this cuts noise and reduces overfitting.
- Fit one TF-IDF vectorizer on all subtitles, set
- Add Targeted Features:
- Count occurrences of rating-relevant terms (e.g., swear words, violence-related vocabulary) as separate features.
- Include metadata like movie runtime (longer films might have more content that pushes them to higher ratings).
- Reduce Dimensionality:
- Use
TruncatedSVD(better for sparse text data than PCA) to compress TF-IDF features into 100-200 dimensions—this helps models like KNN perform better by mitigating the curse of dimensionality.
- Use
3. Tune Models & Fix Overfitting
- Logistic Regression:
- Add regularization with
penalty='l2'(default) orpenalty='l1', and tune theCparameter (smaller values = stronger regularization) usingGridSearchCV. Example:from sklearn.model_selection import GridSearchCV from sklearn.linear_model import LogisticRegression param_grid = {'C': [0.01, 0.1, 1, 10], 'penalty': ['l1', 'l2']} grid = GridSearchCV(LogisticRegression(max_iter=1000), param_grid, cv=5) grid.fit(X_train, y_train)
- Add regularization with
- KNN:
- Reduce feature dimensions first (via TruncatedSVD), then tune
n_neighbors(try 3,5,7) andweights(e.g.,'distance'instead of'uniform') to improve performance.
- Reduce feature dimensions first (via TruncatedSVD), then tune
- Naive Bayes:
- Use
MultinomialNBinstead of GaussianNB—it’s designed for sparse text data like TF-IDF. Tune thealphasmoothing parameter to reduce overfitting.
- Use
- Try Ensemble Models:
- Models like Random Forest, XGBoost, or LightGBM often perform better on small datasets because they use built-in regularization and ensemble multiple weak learners. Enable
class_weight='balanced'here too.
- Models like Random Forest, XGBoost, or LightGBM often perform better on small datasets because they use built-in regularization and ensemble multiple weak learners. Enable
4. Refine Preprocessing
- Validate Text Cleaning:
- Double-check your stopword list—make sure you’re not removing rating-relevant terms (e.g., "violence" shouldn’t be a stopword).
- Strip out non-subtitle noise like timestamps (
[00:15:22]), character names, or scene directions from your text.
- Improve Lemmatization:
- Ensure your POS tagging is accurate (spaCy’s lemmatizer is often more reliable than NLTK’s for this task). Correct POS tags lead to better lemmatization, which helps group similar terms (e.g., "violent" → "violence").
Final Check: Evaluate Properly
- For classification, don’t rely solely on accuracy—look at confusion matrices to see which categories are being misclassified (e.g., PG-13 vs. R is a common mix-up). This can guide you to add more targeted features for those categories.
- For clustering, switch to ARI or NMI to get a true measure of how well your clusters align with ratings.
内容的提问来源于stack exchange,提问作者Harshita Vemula

