You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决文本分类样本偏差:Python实现ADASYN及零售评论分类问询

Hey there! Dealing with class imbalance in a Random Forest classification task is super common when working with real-world review datasets like yours—90% of reviews leaning toward 'Delivery' definitely skews the model’s learning. Let’s break down practical, Python-implementable fixes you can apply to correct this bias and build a more reliable classifier:

Key Fixes for Class Imbalance in Your Random Forest Model

1. Resample Your Dataset

Resampling adjusts the class distribution to give your model equal exposure to both classes. You have two main options here:

Undersample the Majority Class ('Delivery')

Randomly remove samples from the overrepresented class to match the size of the minority class. Just be careful not to discard critical data—use stratified sampling to preserve the class’s underlying distribution as much as possible.

from imblearn.under_sampling import RandomUnderSampler
from sklearn.model_selection import train_test_split

# Split your data first (always split before resampling to avoid data leakage!)
X_train, X_test, y_train, y_test = train_test_split(features, labels, test_size=0.2, stratify=labels)

# Initialize undersampler
rus = RandomUnderSampler(random_state=42)
X_train_resampled, y_train_resampled = rus.fit_resample(X_train, y_train)

# Now train your Random Forest on the resampled data
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(random_state=42)
rf.fit(X_train_resampled, y_train_resampled)

Oversample the Minority Class ('Customer Service')

Instead of removing data, generate synthetic samples for the underrepresented class (using methods like SMOTE) to match the majority class size. This avoids losing information from the majority class.

from imblearn.over_sampling import SMOTE

# Again, split first!
X_train, X_test, y_train, y_test = train_test_split(features, labels, test_size=0.2, stratify=labels)

# Initialize SMOTE
smote = SMOTE(random_state=42)
X_train_resampled, y_train_resampled = smote.fit_resample(X_train, y_train)

# Train Random Forest on resampled data
rf = RandomForestClassifier(random_state=42)
rf.fit(X_train_resampled, y_train_resampled)

2. Adjust Class Weights Directly in Random Forest

Random Forest has a built-in class_weight parameter that assigns higher weights to minority classes, so the model penalizes misclassifying them more heavily. This is a quick fix without modifying your dataset.

from sklearn.ensemble import RandomForestClassifier

# Use 'balanced' to let the model calculate weights based on class frequencies
rf_weighted = RandomForestClassifier(class_weight='balanced', random_state=42)
rf_weighted.fit(X_train, y_train)

# Or define custom weights (e.g., if you want to prioritize minority class even more)
custom_weights = {'Delivery': 1, 'Customer Service': 9}  # Since 90% are Delivery, weight minority 9x
rf_custom = RandomForestClassifier(class_weight=custom_weights, random_state=42)
rf_custom.fit(X_train, y_train)

3. Stop Using Accuracy—Use Imbalance-Aware Metrics

Accuracy is misleading for imbalanced data (your model could just guess 'Delivery' 90% of the time and get high accuracy). Instead, use metrics that focus on the minority class:

from sklearn.metrics import f1_score, precision_recall_curve, auc, confusion_matrix

# Get predictions
y_pred = rf_weighted.predict(X_test)

# Calculate key metrics
print(f"F1-Score (Customer Service): {f1_score(y_test, y_pred, pos_label='Customer Service'):.2f}")
print(f"Confusion Matrix:\n{confusion_matrix(y_test, y_pred)}")

# Precision-Recall AUC (better than ROC AUC for imbalanced data)
precision, recall, _ = precision_recall_curve(y_test, rf_weighted.predict_proba(X_test)[:,1], pos_label='Customer Service')
pr_auc = auc(recall, precision)
print(f"Precision-Recall AUC: {pr_auc:.2f}")

4. Try Imbalance-Focused Ensemble Models

Libraries like imblearn have specialized ensemble models built for imbalanced data, which combine resampling with ensemble learning:

from imblearn.ensemble import BalancedRandomForestClassifier

# This model automatically undersamples majority class in each tree's bootstrap sample
brf = BalancedRandomForestClassifier(random_state=42)
brf.fit(X_train, y_train)

# Evaluate with the same metrics as above
y_pred_brf = brf.predict(X_test)
print(f"Balanced RF F1-Score: {f1_score(y_test, y_pred_brf, pos_label='Customer Service'):.2f}")

Pro Tip

Combine methods for better results—for example, use SMOTE to oversample the minority class, then train a weighted Random Forest, and validate with Precision-Recall AUC. Always test different approaches to see what works best for your specific review data!

内容的提问来源于stack exchange,提问作者Christopher Loynes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:59:20