使用拉普拉斯平滑时是否可能得到负信息增益?
Short answer: No, information gain (IG) will never be negative when using Laplace smoothing. Here's a breakdown of why this is the case:
1. The Core Property of Information Gain
First, recall the definition of information gain:
$$IG = H(Y) - H(Y|X)$$
where $H(Y)$ is the marginal entropy of target variable $Y$, and $H(Y|X)$ is the conditional entropy of $Y$ given feature $X$.
A foundational rule of entropy is that conditional entropy can never exceed marginal entropy:
$$H(Y|X) \leq H(Y)$$
This holds for any valid probability distribution (where probabilities are non-negative and sum to 1), regardless of how you estimate those probabilities. Equality only occurs when $X$ and $Y$ are completely independent—meaning knowing $X$ gives no useful information about $Y$.
2. Laplace Smoothing Produces Valid Probabilities
Laplace smoothing fixes the problem of zero probabilities (e.g., when a feature value $x$ has never been observed, $n_x=0$). The formula you provided:
$$P(Y=y|X=x) = \frac{n_{y|x} + 1}{n_x + |X|}$$
(Note: Standard Laplace smoothing for conditional probabilities uses $|Y|$ (number of class labels) in the denominator to ensure the sum over all $y$ equals 1, but even with your specified denominator, the key validity check still applies.)
This smoothed probability is a valid distribution:
- Every term is positive (numerator and denominator are positive integers)
- For any fixed $x$, the sum of $P(Y=y|X=x)$ over all $y$ equals 1 (assuming the denominator is chosen to maintain this property, like $n_x + |Y|$ in standard smoothing)
3. Why Negative IG Is Impossible
Since Laplace smoothing generates valid probabilities, the core entropy rule $H(Y|X) \leq H(Y)$ remains intact. Subtracting a value that’s less than or equal to $H(Y)$ from $H(Y)$ will always result in a non-negative number.
Even in edge cases (like an unseen feature value $x$), Laplace smoothing assigns uniform-like probabilities to $Y$ given $x$. This might reduce the information gain towards 0, but it will never make $H(Y|X)$ larger than $H(Y)$—so IG can never dip below 0.
内容的提问来源于stack exchange,提问作者Ankit Shubham

