You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何确定Python中随机森林模型的最优分类阈值

Got it, let's translate that R logic into Python step by step! You're looking to find the optimal cutoff for your Random Forest model to tweak predictions (moving away from the default 0.5) and boost performance—here's exactly how to do it, mirroring your R workflow:

Python Implementation for Optimal Random Forest Cutoff

Step 1: Define the Metric Calculation Function

First, let's create a function that replicates your R perform_fn_rf logic. This will take a cutoff value, predicted probabilities, true labels, and return key metrics (sensitivity, specificity, accuracy).

from sklearn.metrics import confusion_matrix
import numpy as np

def perform_fn_rf(cutoff, y_proba, y_true, positive_class="YES"):
    # Convert predicted probabilities to class labels using the cutoff
    predicted_response = np.where(y_proba[:, 1] >= cutoff, "YES", "NO")
    # Generate confusion matrix (ensure labels match your class order)
    cm = confusion_matrix(y_true, predicted_response, labels=["NO", "YES"])
    # Extract TN, FP, FN, TP from the confusion matrix
    tn, fp, fn, tp = cm.ravel()
    
    # Calculate metrics (add checks to avoid division by zero)
    sensitivity = tp / (tp + fn) if (tp + fn) != 0 else 0.0
    specificity = tn / (tn + fp) if (tn + fp) != 0 else 0.0
    accuracy = (tp + tn) / (tp + tn + fp + fn) if (tp + tn + fp + fn) != 0 else 0.0
    
    # Return metrics in an easy-to-handle dictionary
    return {
        "cutoff": round(cutoff, 2),
        "sensitivity": round(sensitivity, 4),
        "specificity": round(specificity, 4),
        "accuracy": round(accuracy, 4)
    }

Step 2: Test Multiple Cutoffs & Find the Optimal One

Next, we'll loop through a range of cutoff values (e.g., 0.1 to 0.9 in 0.01 increments), calculate metrics for each, then pick the cutoff that meets your goal—whether that's maximizing the number of "YES" predictions, accuracy, or a balance of sensitivity/specificity.

import pandas as pd

# Assume you have your trained RandomForest model and validation data ready
# First, get predicted probabilities for the positive class ("YES")
# (Note: predict_proba returns [prob_NO, prob_YES], so we take index 1)
y_proba = rf_model.predict_proba(train_validation.drop("Outcome.Status", axis=1))

# Define the range of cutoffs to test
cutoffs = np.arange(0.1, 0.9, 0.01)

# Collect metrics for each cutoff
results = []
for cutoff in cutoffs:
    metrics = perform_fn_rf(cutoff, y_proba, train_validation["Outcome.Status"])
    results.append(metrics)

# Convert results to a DataFrame for easy analysis
results_df = pd.DataFrame(results)

### Option 1: Find cutoff for maximum predicted "YES"
# Calculate how many "YES" predictions each cutoff produces
results_df["predicted_yes_count"] = [np.sum(y_proba[:,1] >= c) for c in results_df["cutoff"]]
optimal_cutoff_max_yes = results_df.loc[results_df["predicted_yes_count"].idxmax(), "cutoff"]

### Option 2: Find cutoff for maximum accuracy
optimal_cutoff_max_acc = results_df.loc[results_df["accuracy"].idxmax(), "cutoff"]

### Option 3: Find cutoff that balances sensitivity and specificity
# For example, where sensitivity and specificity are closest to each other
results_df["sens_spec_diff"] = abs(results_df["sensitivity"] - results_df["specificity"])
optimal_cutoff_balanced = results_df.loc[results_df["sens_spec_diff"].idxmin(), "cutoff"]

# Print out your results
print(f"Cutoff for maximum predicted YES: {optimal_cutoff_max_yes}")
print(f"Cutoff for maximum accuracy: {optimal_cutoff_max_acc}")
print(f"Cutoff for balanced sensitivity/specificity: {optimal_cutoff_balanced}")

# View the top 5 cutoffs by accuracy
print("\nTop 5 Cutoffs by Accuracy:")
print(results_df.sort_values(by="accuracy", ascending=False).head())

Key Notes

  • Double-check that your true labels (train_validation["Outcome.Status"]) match the class labels used in the function ("NO" and "YES"). If you're using numeric labels (0/1), adjust the np.where condition and confusion matrix labels accordingly.
  • The predict_proba output order depends on your model's classes_ attribute—verify with rf_model.classes_ that y_proba[:,1] corresponds to the "YES" class.
  • "Maximizing YES predictions" will typically lead to a lower cutoff, but this might increase false positives. Always balance this against your model's performance metrics for your specific use case.

内容的提问来源于stack exchange,提问作者Sagar Jaiswal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:44:45