数据点回归直线的理论确定方法及给定数据的回归直线求解
Got it, let's tackle this regression line problem from both theory and practice—super straightforward once you break it down!
At its core, a regression line is all about minimizing the sum of squared errors (SSE)—this is the "least squares" method, the gold standard for linear regression. Here's the step-by-step theory:
First, let's define our terms clearly:
- We have
ndata points:{(x₁,y₁), (x₂,y₂), ..., (xₙ,yₙ)} - Our predicted line (hypothesis function) is
h(x) = w₀ + w₁x, wherew₀is the y-intercept andw₁is the slope (I think you might have had a typo withw + hx—this is the standard notation used in most regression contexts).
The squared error loss function we want to minimize is:
L(w₀, w₁) = Σᵢ₌₁ⁿ (yᵢ - h(xᵢ))² = Σᵢ₌₁ⁿ (yᵢ - w₀ - w₁xᵢ)²
To find the w₀ and w₁ that make this loss as small as possible, we take the partial derivatives of L with respect to each coefficient, set them equal to zero, and solve the resulting system of equations.
Step 1: Solve for
w₀
Take the partial derivative ofLwith respect tow₀, set it to 0:∂L/∂w₀ = -2Σᵢ₌₁ⁿ (yᵢ - w₀ - w₁xᵢ) = 0Simplify this, and you'll find:
w₀ = ȳ - w₁x̄Where
x̄is the mean of allxvalues, andȳis the mean of allyvalues. A handy sanity check: this tells us the regression line always passes through the point(x̄, ȳ).Step 2: Solve for
w₁
Now take the partial derivative ofLwith respect tow₁, set it to 0, and substitute thew₀we just found:∂L/∂w₁ = -2Σᵢ₌₁ⁿ xᵢ(yᵢ - w₀ - w₁xᵢ) = 0After simplifying, the slope
w₁comes out to:w₁ = [Σᵢ₌₁ⁿ (xᵢ - x̄)(yᵢ - ȳ)] / [Σᵢ₌₁ⁿ (xᵢ - x̄)²]You might also recognize this as the ratio of the covariance between
xandyto the variance ofx(Cov(x,y)/Var(x)).
The key idea here: this line is the one where the total squared vertical distance from each data point to the line is smaller than any other possible line. If we assume our errors are normally distributed (a common regression assumption), this least squares solution is also the maximum likelihood estimate.
Let's walk through a concrete example to make this tangible—if your dataset is different, just swap in your numbers!
Suppose we have these data points: (1,2), (2,3), (3,5), (4,6), (5,8)
Step 1: Compute basic summary statistics
- Number of points:
n = 5 - Mean of x values:
x̄ = (1+2+3+4+5)/5 = 3 - Mean of y values:
ȳ = (2+3+5+6+8)/5 = 4.8 - Compute the numerator for
w₁:Σ(xᵢ - x̄)(yᵢ - ȳ)(1-3)(2-4.8) + (2-3)(3-4.8) + (3-3)(5-4.8) + (4-3)(6-4.8) + (5-3)(8-4.8) = (-2)(-2.8) + (-1)(-1.8) + 0*0.2 + 1*1.2 + 2*3.2 = 5.6 + 1.8 + 0 + 1.2 + 6.4 = 15 - Compute the denominator for
w₁:Σ(xᵢ - x̄)²(1-3)² + (2-3)² + (3-3)² + (4-3)² + (5-3)² = 4 + 1 + 0 + 1 + 4 = 10
Step 2: Calculate coefficients
- Slope:
w₁ = 15/10 = 1.5 - Intercept:
w₀ = 4.8 - (1.5 * 3) = 4.8 - 4.5 = 0.3
Step 3: Final regression function
Our line that minimizes squared error is:
h(x) = 0.3 + 1.5x
A quick edge case to note: if all your x values are identical (variance of x is 0), the regression line is just h(x) = ȳ—a horizontal line, since there's no relationship between x and y to model.
内容的提问来源于stack exchange,提问作者Deepak Garg

