sklearn回归器的impurity指什么?与criterion参数是否一致?
Great question—this is a common point of confusion since the docs don’t spell it out explicitly, but let’s break this down clearly:
1. Impurity is directly tied to the criterion parameter
Yes, the "impurity" referenced in min_impurity_decrease is exactly the metric you set via the criterion argument. There’s no separate, hidden impurity definition here—they’re one and the same. This terminology is a holdover from classification trees (where impurity refers to metrics like Gini or entropy), but for regression trees, it’s repurposed to mean the error metric you’re optimizing for.
2. When criterion='mae', MAE is used for impurity calculations
If you set criterion='mae', the impurity of a node will be calculated as the mean absolute error of the target values in that node. The min_impurity_decrease threshold uses this MAE-based impurity to decide whether splitting the node is worthwhile:
- First, calculate the MAE of the parent node
- Calculate the weighted average MAE of the two child nodes (weighted by the number of samples in each child)
- If the parent’s MAE minus this weighted child MAE is greater than or equal to
min_impurity_decrease, the split is allowed
Quick recap for other criteria
To cover all bases, here’s how impurity maps to other common criterion choices:
criterion='mse'(default): Impurity = Mean Squared Error of the node’s target valuescriterion='friedman_mse': Impurity = Friedman’s modified MSE (adjusted to account for potential splits)
内容的提问来源于stack exchange,提问作者karu

