You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

识别并‘捕获’多元回归中使用的观测值

识别并‘捕获’多元回归中使用的观测值

Hey there! This is such a common scenario when comparing regression results—good call wanting to keep the exact same sample for both models to make your comparison valid. Let me break down how to do this in the most popular stats tools, since you didn’t specify which one you’re using:

R

When you run a regression with lm(), the resulting model object stores all the info you need about the used observations. Here are two straightforward ways:

  • Use the model's built-in data subset: The $model component of the regression object contains only the rows that were actually used (no missing values for any of the variables in the multiple regression). You can feed this directly into your simple regression:

    # First run your multiple regression
    mult_reg <- lm(y ~ x1 + x2 + x3, data = my_dataset)
    
    # Run simple regression on the exact same sample
    simple_reg <- lm(y ~ x1, data = mult_reg$model)
    
  • Extract row indices to reuse: If you want to keep working with your original dataset, you can get the indices of the used rows either from the model's row names or by excluding the rows that were dropped due to missing values:

    # Get row indices of observations used in the multiple regression
    used_rows <- rownames(mult_reg$model)
    # Alternatively, use the na.action attribute to find excluded rows
    used_rows <- setdiff(1:nrow(my_dataset), mult_reg$na.action)
    
    # Run simple regression on the subset
    simple_reg <- lm(y ~ x1, data = my_dataset[used_rows, ])
    

Stata

Stata makes this super easy with its built-in sample tracking. After running a regression, the e(sample) function flags which observations were used. You can turn this into a permanent indicator variable to reuse the sample:

// Run your multiple regression first
regress y x1 x2 x3

// Create a variable marking observations used in the regression
gen used_in_mult = e(sample)

// Now run the simple regression only on those observations
regress y x1 if used_in_mult == 1

Python (with statsmodels)

Statsmodels stores the used sample in the model object's data components. You can extract a mask or the filtered dataset directly:

import statsmodels.api as sm
import pandas as pd

# Assume your data is in a pandas DataFrame called my_data
# First run the multiple regression
X_mult = sm.add_constant(my_data[['x1', 'x2', 'x3']])
mult_model = sm.OLS(my_data['y'], X_mult).fit()

# Get the mask of used observations (excludes rows with missing values)
used_mask = mult_model.model.data.row_labels

# Filter your original data to only used rows
used_data = my_data.loc[used_mask]

# Run the simple regression on the filtered data
X_simple = sm.add_constant(used_data['x1'])
simple_model = sm.OLS(used_data['y'], X_simple).fit()

A Quick Note

The key advantage of using the model's built-in tracking (instead of manually subsetting with complete.cases() on your multiple regression variables) is that it accounts for any other exclusions the software might have applied (like if you used weights, clustering, or other filters in your original regression). This ensures your sample is exactly identical across both models.

备注:内容来源于stack exchange,提问作者AKJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.23 07:09:52