You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Random Forest准确率为0.0问题求助(Decision Tree准确率99.99)

Hey there, let’s figure out why your Random Forest is hitting 0% accuracy while your Decision Tree works perfectly. Looking at your output details, there are several key issues we can tackle step by step:

1. You’re not using the optimal parameters from GridSearchCV

Right now, after running GridSearchCV, you’re manually initializing a new RandomForestClassifier with hardcoded parameters (n_estimators=50, max_depth=5), but you’re ignoring the actual best parameters found by the grid search. This is a critical mistake—your manual parameters might be way off from what the grid search determined was optimal, leading to a poorly performing model.

Fix it like this:
Instead of manually creating rfc1, use the pre-trained best estimator directly from GridSearchCV:

# After running CV_rfc.fit(X_train, y_train)
print("Best parameters found:", CV_rfc.best_params_)
rfc1 = CV_rfc.best_estimator_  # Use the already trained optimal model
# No need to call rfc1.fit(X_train, y_train) again—GridSearchCV already did this!

2. Extreme class imbalance is breaking your model

Looking at your classification report, class 1 dominates the test set (1.7M samples) while most other classes have only a handful of samples (some even have just 1-2). Your decision tree might have stumbled on a single feature that perfectly separates classes, but Random Forest’s default settings can struggle with extreme imbalance—especially if it’s not configured to prioritize minority classes.

Fix it:
Add class weighting to your Random Forest to force it to pay attention to smaller classes:

# Initialize the base model with class weighting
rfc = RandomForestClassifier(random_state=42, class_weight='balanced_subsample')
# Then run GridSearchCV as before

balanced_subsample adjusts weights based on the class distribution in each bootstrap sample, which works better for imbalanced multi-class problems than the default uniform weighting.

3. Check for training/test set label mismatches

Your confusion matrix shows all predictions are failing, and the classification report lists class 0 with 0 support (meaning no samples in the test set belong to class 0). If your model is predicting class 0 (or any class not present in the test set), it will naturally have 0% accuracy.

Verify this:
Check if your training set contains classes that aren’t in the test set (or vice versa), and confirm label encoding is consistent between splits:

import pandas as pd

print("Training set class distribution:")
print(pd.Series(y_train).value_counts())

print("\nTest set class distribution:")
print(pd.Series(y_test).value_counts())

# Also check what your model is actually predicting
preds = rfc1.predict(X_test)
print("\nFirst 10 predictions:", preds[:10])
print("First 10 true labels:", y_test[:10])

If the model is predicting classes that don’t exist in the test set, you’ll need to align your training and test label sets (e.g., merge rare classes or collect more samples for underrepresented classes).

4. Ensure consistent data preprocessing

If you applied any preprocessing (like scaling, normalization, or feature selection) to X_train, you must apply the exact same transformation to X_test—using the fit from the training set, not re-fitting on the test set. For example:

from sklearn.preprocessing import StandardScaler

# Fit scaler ONLY on training data
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
# Transform test data using the already fitted scaler
X_test_scaled = scaler.transform(X_test)

# Now use X_train_scaled and X_test_scaled for model training/testing

Forgetting this step can completely break model performance, as the test data will be in a different scale than what the model learned on.

5. Simplify the model for debugging

If the above steps don’t fix the issue, temporarily simplify your Random Forest to rule out parameter-related problems:

# Use a basic, un-tuned Random Forest
simple_rfc = RandomForestClassifier(random_state=42, n_estimators=10, class_weight='balanced')
simple_rfc.fit(X_train, y_train)
print("Simple RF test accuracy:", accuracy_score(y_test, simple_rfc.predict(X_test)))

If this simple model performs better, your grid search parameter range might be too restrictive (e.g., max_depth=5 is too shallow to capture patterns in your data).


内容的提问来源于stack exchange,提问作者Maryem Samet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:53:06