LightGBM及二阶近似提升树处理L1损失的技术问询
Understanding LightGBM's Handling of L1 Losses (MAE, MAPE, Huber)
Let’s unpack your questions one by one—this is a really nuanced area where the usual second-order boosting rules get bent a bit, so it’s totally reasonable to have these doubts.
1. Does LightGBM degrade to traditional gradient boosting for MAE?
Short answer: Not exactly, but it aligns far closer to traditional GBM behavior than it does for L2 loss.
Here’s the breakdown:
- For L2 loss, LightGBM leans heavily on both first (gradient) and second (Hessian) derivatives to calculate split gains, using the second-order Taylor approximation of the loss function to drive efficient splits.
- For MAE, the gradient is just
sign(y - y_pred)(constant magnitude ±1), and the Hessian is mathematically 0 (since the absolute value function isn’t differentiable at 0, and its derivative is constant elsewhere). As you noted, LightGBM works around this by forcing the Hessian to 1 for all samples. - With Hessian=1, the split gain formula simplifies from the standard second-order version:
to a first-order-like calculation:gain = (left_grad_sum² / left_hess_sum) + (right_grad_sum² / right_hess_sum) - (total_grad_sum² / total_hess_sum)
Since gradients are ±1, this gain measures how well a split separates positive and negative residuals (i.e., samples where the model underpredicts vs overpredicts). This is similar to traditional GBM, which splits based on minimizing residual error for L1 loss.gain = left_grad_sum² + right_grad_sum² - total_grad_sum² - The key difference is that LightGBM still retains its optimized histogram-based splitting and leaf-wise growth (instead of level-wise), so it’s faster than traditional GBM even for MAE. The leaf node output calculation (using the median/quantile of residuals) is also directly borrowed from traditional GBM’s logic for L1 loss.
2. How does GOSS work with constant-gradient losses like MAE?
Great catch—GOSS (Gradient-based One-Side Sampling) was built to prioritize samples with large gradients, since those drive more model improvement for L2 loss. But for MAE, every sample has a gradient of ±1 (same absolute magnitude), so GOSS’s core logic breaks down:
- There’s no way to distinguish "large gradient" vs "small gradient" samples—all gradients are identical in size.
- In this scenario, GOSS effectively reduces to random sampling of the "small gradient" group (which is all samples), since there’s no gradient magnitude to filter on. This means GOSS provides no unique benefit over random sampling for MAE; it might even hurt performance by introducing unnecessary variance.
- LightGBM’s default behavior accounts for this: when using L1-based losses, it automatically disables GOSS (or falls back to a neutral sampling strategy if needed). If you explicitly set
boosting_type='goss'with MAE, you’ll likely see no improvement (or worse results) compared to the standardgbdtmode.
Quick Recap of LightGBM’s L1 Loss Pipeline
To tie it all together:
- Split phase: Uses gradient (
sign(y - y_pred)) and fixed Hessian=1 to calculate split gains, focusing on separating underpredicted/overpredicted samples. - Leaf output phase: Instead of the standard
sum(grad) / sum(hess)formula (which would be useless for MAE), it computes the quantile of residuals (50% for MAE, the median) that minimizes the L1 loss. - Optimizations: Histogram splitting and leaf-wise growth still apply, but GOSS is effectively irrelevant here.
内容的提问来源于stack exchange,提问作者Francisco Rodriguez
相关产品推荐
相关产品推荐

