You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字幕的电影评级分类准确率不佳,寻求Python优化方案

How to Boost Movie Rating Classification Accuracy & Fix Workflow Gaps

Hey there! Let’s break down your movie rating classification problem step by step—first identifying potential missteps in your current process, then diving into actionable fixes to lift those accuracy numbers.

First, Let’s Flag Clear Issues in Your Current Approach

  1. Incorrect Clustering Evaluation

    You used accuracy to assess KMeans performance, which is a mistake. Clustering algorithms assign arbitrary cluster labels that don’t map directly to your rating categories (e.g., KMeans might group R and PG-13 films together, but you can’t just force those clusters to match your true labels). For clustering, use metrics like Adjusted Rand Index (ARI) or Normalized Mutual Information (NMI) instead—these measure how well clusters align with true labels without assuming direct one-to-one mapping.

  2. Overly High-Dimensional Features for a Small Dataset
    Your dataset only has 130 films, but you’re using 5000 features (1000 per category from TF-IDF). This is a classic case of the curse of dimensionality: with more features than samples, models struggle to learn meaningful patterns, leading to either overfitting (like your Logistic Regression) or poor generalization (like Naive Bayes).

  3. TF-IDF Implementation Misstep
    Generating TF-IDF features per category and merging them is not standard practice. You should fit a single TF-IDF vectorizer on your entire dataset, then select top N features (e.g., 500-1000 total) using max_features. Per-category TF-IDF creates redundant, noisy features that don’t help models learn cross-category patterns.

  4. Severe Overfitting in Logistic Regression
    A training accuracy of 0.96 vs. test accuracy of 0.53 means your model has memorized training data instead of learning generalizable patterns. This is almost certainly due to high feature dimensionality and lack of regularization.


Actionable Fixes to Boost Accuracy

1. Fix Dataset & Class Imbalance Issues

  • Expand Your Dataset: 130 films split across 5 categories averages just 26 samples per class—way too small for reliable text classification. Look for public datasets with movie subtitles + MPAA ratings (e.g., scrape IMDb for ratings paired with open-source subtitle data).
  • Address Class Imbalance: If some ratings (like G or NR) have far fewer samples, use:
    • class_weight='balanced' parameter in models like Logistic Regression or SVC to give more weight to underrepresented classes.
    • SMOTE (carefully—text data is sparse, so use text-specific oversampling tools) or undersampling of overrepresented classes.

2. Optimize Feature Engineering

  • Simplify TF-IDF:
    • Fit one TF-IDF vectorizer on all subtitles, set ngram_range=(1,2) to capture meaningful phrases (e.g., "graphic violence" or "sexual content" are far more predictive than single words like "violence").
    • Reduce max_features to 500-800 total features—this cuts noise and reduces overfitting.
  • Add Targeted Features:
    • Count occurrences of rating-relevant terms (e.g., swear words, violence-related vocabulary) as separate features.
    • Include metadata like movie runtime (longer films might have more content that pushes them to higher ratings).
  • Reduce Dimensionality:
    • Use TruncatedSVD (better for sparse text data than PCA) to compress TF-IDF features into 100-200 dimensions—this helps models like KNN perform better by mitigating the curse of dimensionality.

3. Tune Models & Fix Overfitting

  • Logistic Regression:
    • Add regularization with penalty='l2' (default) or penalty='l1', and tune the C parameter (smaller values = stronger regularization) using GridSearchCV. Example:
      from sklearn.model_selection import GridSearchCV
      from sklearn.linear_model import LogisticRegression
      
      param_grid = {'C': [0.01, 0.1, 1, 10], 'penalty': ['l1', 'l2']}
      grid = GridSearchCV(LogisticRegression(max_iter=1000), param_grid, cv=5)
      grid.fit(X_train, y_train)
      
  • KNN:
    • Reduce feature dimensions first (via TruncatedSVD), then tune n_neighbors (try 3,5,7) and weights (e.g., 'distance' instead of 'uniform') to improve performance.
  • Naive Bayes:
    • Use MultinomialNB instead of GaussianNB—it’s designed for sparse text data like TF-IDF. Tune the alpha smoothing parameter to reduce overfitting.
  • Try Ensemble Models:
    • Models like Random Forest, XGBoost, or LightGBM often perform better on small datasets because they use built-in regularization and ensemble multiple weak learners. Enable class_weight='balanced' here too.

4. Refine Preprocessing

  • Validate Text Cleaning:
    • Double-check your stopword list—make sure you’re not removing rating-relevant terms (e.g., "violence" shouldn’t be a stopword).
    • Strip out non-subtitle noise like timestamps ([00:15:22]), character names, or scene directions from your text.
  • Improve Lemmatization:
    • Ensure your POS tagging is accurate (spaCy’s lemmatizer is often more reliable than NLTK’s for this task). Correct POS tags lead to better lemmatization, which helps group similar terms (e.g., "violent" → "violence").

Final Check: Evaluate Properly

  • For classification, don’t rely solely on accuracy—look at confusion matrices to see which categories are being misclassified (e.g., PG-13 vs. R is a common mix-up). This can guide you to add more targeted features for those categories.
  • For clustering, switch to ARI or NMI to get a true measure of how well your clusters align with ratings.

内容的提问来源于stack exchange,提问作者Harshita Vemula

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:07:36