R中向量标准化后均值为何是接近0的极小值而非0?
Great question—this is such a common gotcha when working with floating-point numbers, and it’s not a flaw in R’s algorithms! Let’s break down what’s going on:
First, let’s recap your steps to spot where the tiny error creeps in:
- You defined
t <- c(1,2,3,7)and calculated its mean as 3.25. That’s a clean value—13 divided by 4, which can be represented perfectly in binary floating-point (since 4 is a power of 2, no rounding needed here). - Subtracting the mean gives
[-2.25, -1.25, -0.25, 3.75]—all these values are also perfect binary floats, no issues at this stage. - The problem hits when calculating the standard deviation:
- The sum of squared deviations is
(1-3.25)² + (2-3.25)² + (3-3.25)² + (7-3.25)² = 5.0625 + 1.5625 + 0.0625 + 14.0625 = 20.75 - Variance is that sum divided by
n-1(3), so20.75/3 ≈ 6.916666666666667 - The standard deviation is the square root of that variance, which is an irrational number—it can’t be represented exactly as a finite binary floating-point value. R has to store an approximation of this number instead.
- The sum of squared deviations is
When you divide each element of your centered vector by this approximate standard deviation, each division introduces a tiny rounding error. By the time you sum all those standardized values and divide by the length, those tiny errors add up to a mean that’s extremely close to 0 (like -1.734723e-17), but not exactly 0.
This isn’t unique to R—every programming language that uses binary floating-point arithmetic (Python, C++, Java, you name it) will have this same behavior. It’s just how computers handle decimal numbers behind the scenes!
If you want to confirm that the mean is effectively 0 for practical purposes, try running all.equal(mean(standardized_t), 0) in R. It’ll return TRUE because R uses a small tolerance to account for these tiny floating-point discrepancies.
内容的提问来源于stack exchange,提问作者Lu Zhang

