如何确定Python中随机森林模型的最优分类阈值
Got it, let's translate that R logic into Python step by step! You're looking to find the optimal cutoff for your Random Forest model to tweak predictions (moving away from the default 0.5) and boost performance—here's exactly how to do it, mirroring your R workflow:
Step 1: Define the Metric Calculation Function
First, let's create a function that replicates your R perform_fn_rf logic. This will take a cutoff value, predicted probabilities, true labels, and return key metrics (sensitivity, specificity, accuracy).
from sklearn.metrics import confusion_matrix import numpy as np def perform_fn_rf(cutoff, y_proba, y_true, positive_class="YES"): # Convert predicted probabilities to class labels using the cutoff predicted_response = np.where(y_proba[:, 1] >= cutoff, "YES", "NO") # Generate confusion matrix (ensure labels match your class order) cm = confusion_matrix(y_true, predicted_response, labels=["NO", "YES"]) # Extract TN, FP, FN, TP from the confusion matrix tn, fp, fn, tp = cm.ravel() # Calculate metrics (add checks to avoid division by zero) sensitivity = tp / (tp + fn) if (tp + fn) != 0 else 0.0 specificity = tn / (tn + fp) if (tn + fp) != 0 else 0.0 accuracy = (tp + tn) / (tp + tn + fp + fn) if (tp + tn + fp + fn) != 0 else 0.0 # Return metrics in an easy-to-handle dictionary return { "cutoff": round(cutoff, 2), "sensitivity": round(sensitivity, 4), "specificity": round(specificity, 4), "accuracy": round(accuracy, 4) }
Step 2: Test Multiple Cutoffs & Find the Optimal One
Next, we'll loop through a range of cutoff values (e.g., 0.1 to 0.9 in 0.01 increments), calculate metrics for each, then pick the cutoff that meets your goal—whether that's maximizing the number of "YES" predictions, accuracy, or a balance of sensitivity/specificity.
import pandas as pd # Assume you have your trained RandomForest model and validation data ready # First, get predicted probabilities for the positive class ("YES") # (Note: predict_proba returns [prob_NO, prob_YES], so we take index 1) y_proba = rf_model.predict_proba(train_validation.drop("Outcome.Status", axis=1)) # Define the range of cutoffs to test cutoffs = np.arange(0.1, 0.9, 0.01) # Collect metrics for each cutoff results = [] for cutoff in cutoffs: metrics = perform_fn_rf(cutoff, y_proba, train_validation["Outcome.Status"]) results.append(metrics) # Convert results to a DataFrame for easy analysis results_df = pd.DataFrame(results) ### Option 1: Find cutoff for maximum predicted "YES" # Calculate how many "YES" predictions each cutoff produces results_df["predicted_yes_count"] = [np.sum(y_proba[:,1] >= c) for c in results_df["cutoff"]] optimal_cutoff_max_yes = results_df.loc[results_df["predicted_yes_count"].idxmax(), "cutoff"] ### Option 2: Find cutoff for maximum accuracy optimal_cutoff_max_acc = results_df.loc[results_df["accuracy"].idxmax(), "cutoff"] ### Option 3: Find cutoff that balances sensitivity and specificity # For example, where sensitivity and specificity are closest to each other results_df["sens_spec_diff"] = abs(results_df["sensitivity"] - results_df["specificity"]) optimal_cutoff_balanced = results_df.loc[results_df["sens_spec_diff"].idxmin(), "cutoff"] # Print out your results print(f"Cutoff for maximum predicted YES: {optimal_cutoff_max_yes}") print(f"Cutoff for maximum accuracy: {optimal_cutoff_max_acc}") print(f"Cutoff for balanced sensitivity/specificity: {optimal_cutoff_balanced}") # View the top 5 cutoffs by accuracy print("\nTop 5 Cutoffs by Accuracy:") print(results_df.sort_values(by="accuracy", ascending=False).head())
Key Notes
- Double-check that your true labels (
train_validation["Outcome.Status"]) match the class labels used in the function ("NO" and "YES"). If you're using numeric labels (0/1), adjust thenp.wherecondition and confusion matrix labels accordingly. - The
predict_probaoutput order depends on your model'sclasses_attribute—verify withrf_model.classes_thaty_proba[:,1]corresponds to the "YES" class. - "Maximizing YES predictions" will typically lead to a lower cutoff, but this might increase false positives. Always balance this against your model's performance metrics for your specific use case.
内容的提问来源于stack exchange,提问作者Sagar Jaiswal

