为何需要浮点相对精度?MATLAB迭代收敛检查中sqrt(eps)作用解析
First off, let's ground this in how floating-point numbers work. Unlike integers, real numbers can't all be represented exactly in a finite amount of memory—double-precision floats (the default in MATLAB) use 52 bits for the mantissa, which means there's a hard limit to how precise we can get.
The floating-point relative precision (eps for double-precision, equal to 2^-52 ≈ 2.2e-16) is the smallest positive number such that 1 + eps can be distinguished from 1 in floating-point arithmetic. Put more generally, for any non-zero real number x, the closest representable float to x will have a relative error no larger than eps/2.
We need this value for a few key reasons:
- Error estimation: It tells us the inherent precision limit of our calculations—any result can't be trusted beyond this relative error margin.
- Convergence checks: As you saw in the MATLAB code, it helps set a meaningful threshold for when an iterative process has "converged"—we don't want to keep iterating just because of floating-point noise.
- Numerical stability: It guides us in avoiding operations that would amplify floating-point errors beyond acceptable levels.
sqrt(eps) Let's look at the code again:
dw = max(max(abs(W-Wnew) / (sqrt(eps)+max(max(abs(Wnew))))));
This calculates the maximum relative change between the old parameter matrix W and the updated Wnew. The core idea is to measure how much W has changed relative to the magnitude of Wnew—but there's a critical problem if Wnew is very close to zero: the denominator max(max(abs(Wnew))) would approach zero, leading to a huge (and meaningless) value for dw. That's where sqrt(eps) comes in.
Here's why it's included:
- Prevent division by (near) zero: When
Wnewis tiny (e.g., all elements are near machine epsilon), addingsqrt(eps)gives the denominator a hard lower bound. This stops the division from blowing up to infinity, which would incorrectly signal that the iteration hasn't converged. - Balance between stability and accuracy: Why
sqrt(eps)instead of justeps? If we usedepsdirectly, whenWnewhas a magnitude aroundsqrt(eps), theepsterm would dominate the denominator, skewing the relative change calculation.sqrt(eps)(≈1.4e-8 for double-precision) is small enough that it doesn't affect the relative error calculation whenWnewis of normal magnitude, but large enough to prevent numerical instability whenWnewis near zero. - Robust convergence checking: This trick ensures the convergence check behaves consistently across all possible magnitudes of
Wnew, so we don't get false positives (thinking we haven't converged when we have) or false negatives (stopping too early) due to floating-point quirks.
内容的提问来源于stack exchange,提问作者user3086871

