关于二进制对称信道(BSC)互信息的技术疑问
Great question—this is a common point of confusion when first diving into information theory, so let's unpack it clearly.
First, let's clarify why $H(Y) = \log_2 2 = 1$ in that BSC derivation: this only holds when the input distribution $X$ is uniform (i.e., $P(X=0) = P(X=1) = 0.5$). When calculating the channel capacity (the maximum possible mutual information $I(X,Y)$ for a given channel), we optimize the input distribution to maximize $I(X,Y)$. For a BSC, the uniform input distribution achieves this maximum.
In this scenario, let's verify the output probabilities:
- $P(Y=0) = P(X=0)(1-f) + P(X=1)f = 0.5(1-f) + 0.5f = 0.5$
- $P(Y=1) = P(X=1)(1-f) + P(X=0)f = 0.5(1-f) + 0.5f = 0.5$
Since $Y$ is also uniformly distributed over the 2-symbol alphabet, its entropy $H(Y)$ is indeed $\log_2 2 = 1$.
输出熵并不总是等于字母表大小的对数
This equality only holds if the output distribution is uniform over its alphabet. If either the input distribution is non-uniform, or the channel's structure leads to non-uniform output probabilities, $H(Y)$ will be less than $\log_2 |\mathcal{Y}|$ (where $|\mathcal{Y}|$ is the size of the output alphabet). Here are two concrete examples:
例1:非均匀输入的BSC
Take the same BSC with flip probability $f$, but let the input $X$ have a skewed distribution: $P(X=0) = 0.9$, $P(X=1) = 0.1$. Now calculate the output probabilities:
- $P(Y=0) = 0.9(1-f) + 0.1f = 0.9 - 0.8f$
- $P(Y=1) = 0.1(1-f) + 0.9f = 0.1 + 0.8f$
Unless $f=0.5$ (a fully noisy channel where output is independent of input), these probabilities won't be 0.5. For example, if $f=0.1$:
- $P(Y=0) = 0.90.9 + 0.10.1 = 0.82$
- $P(Y=1) = 0.10.9 + 0.90.1 = 0.18$
The entropy $H(Y) = -0.82\log_2 0.82 - 0.18\log_2 0.18 \approx 0.76$, which is much less than 1.
例2:二进制删除信道(BEC)
Consider a BEC with output alphabet ${0,1,e}$ (where $e$ represents a "deleted" symbol) and deletion probability $d=0.2$. Let the input $X$ be uniform ($P(X=0)=P(X=1)=0.5$). The output probabilities are:
- $P(Y=0) = 0.5*(1-0.2) = 0.4$
- $P(Y=1) = 0.5*(1-0.2) = 0.4$
- $P(Y=e) = 0.2$
The output entropy here is:
$H(Y) = -0.4\log_2 0.4 -0.4\log_2 0.4 -0.2\log_2 0.2 \approx 1.52$, which is less than $\log_2 3 \approx 1.585$ (the entropy of a uniform 3-symbol distribution).
总结
The equality $H(Y) = \log_2 |\mathcal{Y}|$ is a special case, not a rule. It only occurs when the output distribution is perfectly uniform, which typically happens when the input is uniform and the channel preserves that uniformity (like a BSC with uniform input). In most other cases, the output entropy will be lower due to non-uniform probabilities.
内容的提问来源于stack exchange,提问作者tomwesolowski

