解决文本分类样本偏差:Python实现ADASYN及零售评论分类问询
Hey there! Dealing with class imbalance in a Random Forest classification task is super common when working with real-world review datasets like yours—90% of reviews leaning toward 'Delivery' definitely skews the model’s learning. Let’s break down practical, Python-implementable fixes you can apply to correct this bias and build a more reliable classifier:
1. Resample Your Dataset
Resampling adjusts the class distribution to give your model equal exposure to both classes. You have two main options here:
Undersample the Majority Class ('Delivery')
Randomly remove samples from the overrepresented class to match the size of the minority class. Just be careful not to discard critical data—use stratified sampling to preserve the class’s underlying distribution as much as possible.
from imblearn.under_sampling import RandomUnderSampler from sklearn.model_selection import train_test_split # Split your data first (always split before resampling to avoid data leakage!) X_train, X_test, y_train, y_test = train_test_split(features, labels, test_size=0.2, stratify=labels) # Initialize undersampler rus = RandomUnderSampler(random_state=42) X_train_resampled, y_train_resampled = rus.fit_resample(X_train, y_train) # Now train your Random Forest on the resampled data from sklearn.ensemble import RandomForestClassifier rf = RandomForestClassifier(random_state=42) rf.fit(X_train_resampled, y_train_resampled)
Oversample the Minority Class ('Customer Service')
Instead of removing data, generate synthetic samples for the underrepresented class (using methods like SMOTE) to match the majority class size. This avoids losing information from the majority class.
from imblearn.over_sampling import SMOTE # Again, split first! X_train, X_test, y_train, y_test = train_test_split(features, labels, test_size=0.2, stratify=labels) # Initialize SMOTE smote = SMOTE(random_state=42) X_train_resampled, y_train_resampled = smote.fit_resample(X_train, y_train) # Train Random Forest on resampled data rf = RandomForestClassifier(random_state=42) rf.fit(X_train_resampled, y_train_resampled)
2. Adjust Class Weights Directly in Random Forest
Random Forest has a built-in class_weight parameter that assigns higher weights to minority classes, so the model penalizes misclassifying them more heavily. This is a quick fix without modifying your dataset.
from sklearn.ensemble import RandomForestClassifier # Use 'balanced' to let the model calculate weights based on class frequencies rf_weighted = RandomForestClassifier(class_weight='balanced', random_state=42) rf_weighted.fit(X_train, y_train) # Or define custom weights (e.g., if you want to prioritize minority class even more) custom_weights = {'Delivery': 1, 'Customer Service': 9} # Since 90% are Delivery, weight minority 9x rf_custom = RandomForestClassifier(class_weight=custom_weights, random_state=42) rf_custom.fit(X_train, y_train)
3. Stop Using Accuracy—Use Imbalance-Aware Metrics
Accuracy is misleading for imbalanced data (your model could just guess 'Delivery' 90% of the time and get high accuracy). Instead, use metrics that focus on the minority class:
from sklearn.metrics import f1_score, precision_recall_curve, auc, confusion_matrix # Get predictions y_pred = rf_weighted.predict(X_test) # Calculate key metrics print(f"F1-Score (Customer Service): {f1_score(y_test, y_pred, pos_label='Customer Service'):.2f}") print(f"Confusion Matrix:\n{confusion_matrix(y_test, y_pred)}") # Precision-Recall AUC (better than ROC AUC for imbalanced data) precision, recall, _ = precision_recall_curve(y_test, rf_weighted.predict_proba(X_test)[:,1], pos_label='Customer Service') pr_auc = auc(recall, precision) print(f"Precision-Recall AUC: {pr_auc:.2f}")
4. Try Imbalance-Focused Ensemble Models
Libraries like imblearn have specialized ensemble models built for imbalanced data, which combine resampling with ensemble learning:
from imblearn.ensemble import BalancedRandomForestClassifier # This model automatically undersamples majority class in each tree's bootstrap sample brf = BalancedRandomForestClassifier(random_state=42) brf.fit(X_train, y_train) # Evaluate with the same metrics as above y_pred_brf = brf.predict(X_test) print(f"Balanced RF F1-Score: {f1_score(y_test, y_pred_brf, pos_label='Customer Service'):.2f}")
Pro Tip
Combine methods for better results—for example, use SMOTE to oversample the minority class, then train a weighted Random Forest, and validate with Precision-Recall AUC. Always test different approaches to see what works best for your specific review data!
内容的提问来源于stack exchange,提问作者Christopher Loynes

