识别并‘捕获’多元回归中使用的观测值
Hey there! This is such a common scenario when comparing regression results—good call wanting to keep the exact same sample for both models to make your comparison valid. Let me break down how to do this in the most popular stats tools, since you didn’t specify which one you’re using:
R
When you run a regression with lm(), the resulting model object stores all the info you need about the used observations. Here are two straightforward ways:
Use the model's built-in data subset: The
$modelcomponent of the regression object contains only the rows that were actually used (no missing values for any of the variables in the multiple regression). You can feed this directly into your simple regression:# First run your multiple regression mult_reg <- lm(y ~ x1 + x2 + x3, data = my_dataset) # Run simple regression on the exact same sample simple_reg <- lm(y ~ x1, data = mult_reg$model)Extract row indices to reuse: If you want to keep working with your original dataset, you can get the indices of the used rows either from the model's row names or by excluding the rows that were dropped due to missing values:
# Get row indices of observations used in the multiple regression used_rows <- rownames(mult_reg$model) # Alternatively, use the na.action attribute to find excluded rows used_rows <- setdiff(1:nrow(my_dataset), mult_reg$na.action) # Run simple regression on the subset simple_reg <- lm(y ~ x1, data = my_dataset[used_rows, ])
Stata
Stata makes this super easy with its built-in sample tracking. After running a regression, the e(sample) function flags which observations were used. You can turn this into a permanent indicator variable to reuse the sample:
// Run your multiple regression first regress y x1 x2 x3 // Create a variable marking observations used in the regression gen used_in_mult = e(sample) // Now run the simple regression only on those observations regress y x1 if used_in_mult == 1
Python (with statsmodels)
Statsmodels stores the used sample in the model object's data components. You can extract a mask or the filtered dataset directly:
import statsmodels.api as sm import pandas as pd # Assume your data is in a pandas DataFrame called my_data # First run the multiple regression X_mult = sm.add_constant(my_data[['x1', 'x2', 'x3']]) mult_model = sm.OLS(my_data['y'], X_mult).fit() # Get the mask of used observations (excludes rows with missing values) used_mask = mult_model.model.data.row_labels # Filter your original data to only used rows used_data = my_data.loc[used_mask] # Run the simple regression on the filtered data X_simple = sm.add_constant(used_data['x1']) simple_model = sm.OLS(used_data['y'], X_simple).fit()
A Quick Note
The key advantage of using the model's built-in tracking (instead of manually subsetting with complete.cases() on your multiple regression variables) is that it accounts for any other exclusions the software might have applied (like if you used weights, clustering, or other filters in your original regression). This ensures your sample is exactly identical across both models.
备注:内容来源于stack exchange,提问作者AKJ

