如何手动计算逻辑回归系数?二元自变量模型系数及数据集X1/X2计算问询
Great question—manual calculation of logistic regression coefficients is a fantastic way to build intuition for how the model works, since unlike linear regression, there’s no neat closed-form solution. Let’s break this down step by step.
1. Core Idea: Maximum Likelihood Estimation (MLE)
Logistic regression uses MLE to find coefficients that maximize the probability of observing your dataset. Here’s the high-level workflow:
Step 1: Define the Likelihood & Log-Likelihood
For each sample (i) with outcome (y_i) (0 or 1), the predicted probability of (y_i=1) is:p_i = 1 / (1 + e^-(B0 + B1*X1i + ... + Bk*Xki))The likelihood function (L) multiplies the probabilities of all observed outcomes:
L = product(p_i^y_i * (1-p_i)^(1-y_i))Taking the natural log turns products into sums (easier to work with), giving the log-likelihood (LL):
LL = sum(y_i * ln(p_i) + (1-y_i) * ln(1-p_i))Step 2: Partial Derivatives & Normal Equations
To maximize LL, we take partial derivatives with respect to each coefficient (B_j) and set them to 0. This gives a system of equations:- For the intercept (B0):
sum(y_i - p_i) = 0 - For each predictor (Bj) (j ≥1):
sum((y_i - p_i)*Xji) = 0
These equations can’t be solved directly (thanks to the nonlinear logistic function), so we use iterative methods.
- For the intercept (B0):
Step 3: Newton-Raphson Iteration (Most Common Method)
This is the standard approach to converge on the optimal coefficients:- Start with initial coefficient values (usually all 0s).
- Calculate (p_i) for every sample using the current coefficients.
- Build a diagonal weight matrix (W) where (Wii = p_i*(1-p_i)) (this accounts for the variance of each predicted probability).
- Create the design matrix (X): first column is all 1s (for (B0)), followed by columns for each predictor.
- Update coefficients using this formula:
B_new = B_old + (X^T * W * X)^(-1) * X^T * (y - p) - Repeat steps 2-5 until coefficients change by less than a small threshold (e.g., 1e-6) or the log-likelihood stops improving.
2. Two-Predictor Logistic Regression: General Calculation
For your model logit(π) = B0 + B1*X1 + B2*X2 (where (π = P(Y=1))), the process is identical to the above—only the design matrix (X) has 3 columns (1, X1, X2). Here’s a condensed step-by-step:
- Gather all samples with ((X1_i, X2_i, y_i)) values.
- Initialize (B0=0, B1=0, B2=0).
- Iterate:
- Compute (p_i = 1/(1 + e^-(B0 + B1X1_i + B2X2_i))) for each sample.
- Calculate residuals (r_i = y_i - p_i).
- Build the weight matrix (W) (diagonal entries (p_i*(1-p_i))).
- Compute the 3x3 matrix (X^T * W * X) using sums of (Wii), (WiiX1_i), (WiiX2_i), (WiiX1_i²), (WiiX1_iX2_i), (WiiX2_i²).
- Compute the 3x1 vector (X^T * r) using sums of (r_i), (r_iX1_i), (r_iX2_i).
- Invert (X^T * W * X) (make sure it’s invertible—no perfect multicollinearity!).
- Update coefficients with the formula from Step 3.5.
- Stop when coefficients stabilize.
3. Example with a Sample Dataset
You mentioned a "specified dataset below" but it wasn’t included, so let’s use a small example to walk through the process. This will help you adapt it to your actual data.
Sample Dataset
| y (Outcome) | X1 | X2 |
|---|---|---|
| 1 | 2 | 3 |
| 0 | 1 | 5 |
| 1 | 3 | 2 |
| 0 | 0 | 4 |
| 1 | 4 | 1 |
Step-by-Step Calculation
- Initialization: (B0=0, B1=0, B2=0)
- First Iteration:
- All (p_i = 1/(1+e^0) = 0.5) (since the linear combination is 0).
- Residuals: (r_i = y_i - 0.5) → [0.5, -0.5, 0.5, -0.5, 0.5]
- Weight matrix entries: (Wii = 0.5*0.5 = 0.25) for all samples.
- Compute (X^T * W * X):
[1.25 2.5 3.75] [2.5 7.5 5.25] [3.75 5.25 13.75] - Compute (X^T * r): [0.5, 4, -1.5]
- Invert the 3x3 matrix (using standard matrix inversion methods):
[ 2.7692 -0.6923 -0.4615] [-0.6923 0.2308 0.0769] [-0.4615 0.0769 0.2308] - Update coefficients:
- (B0_{new} = 0 + (2.76920.5) + (-0.69234) + (-0.4615*(-1.5)) = -0.6923)
- (B1_{new} = 0 + (-0.69230.5) + (0.23084) + (0.0769*(-1.5)) = 0.4617)
- (B2_{new} = 0 + (-0.46150.5) + (0.07694) + (0.2308*(-1.5)) = -0.2694)
- Subsequent Iterations: Repeat the process with the new coefficients. For example, next time you’ll calculate (p_i) using (B0=-0.6923, B1=0.4617, B2=-0.2694), and keep going until coefficients change by less than your threshold (e.g., 0.0001).
Note on X1 & X2 Values
If you’re asking how to get X1/X2 values from your dataset—they’re just the observed predictor values for each sample! If you need to preprocess them (like standardizing or centering), that’s a separate step. For example, standardizing X1 would look like:
X1_std = (X1 - mean(X1)) / std(X1)
But in coefficient calculation, you use the X values as they are (or preprocessed, depending on your modeling choice).
内容的提问来源于stack exchange,提问作者Divyanshu

