You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言OLS回归中排除数据集中的98、99特定值?

Fixing OLS Regression by Excluding Invalid e5 Values in R

Hey there! I get that dealing with wonky response values can feel frustrating when you're new to R. Let's walk through a few straightforward solutions to exclude those 98 ("don't know") and 99 ("refuse to answer") entries from your regression without messing up your workflow.

Solution 1: Filter directly in the lm() function (no dataset changes)

The quickest fix is to use the subset argument in lm() to only include rows where e5 falls in the valid 1-5 range. This way you don't have to touch your original dataset at all:

# Run regression with only valid e5 responses
PolReg <- lm(e5 ~ age + gender + f1 + turnout05, 
             data = Data, 
             subset = e5 %in% 1:5)

The %in% 1:5 check ensures e5 is exactly 1, 2, 3, 4, or 5—automatically cutting out 98 and 99. Perfect for quick, one-off analyses.

Solution 2: Create a cleaned dataset for repeated use

If you plan to run multiple analyses with the filtered data, it's smarter to make a separate cleaned version first:

# Keep only rows where e5 is valid (don't forget the comma at the end!)
clean_Data <- Data[Data$e5 %in% 1:5, ]

# Run regression on the cleaned dataset
PolReg <- lm(e5 ~ age + gender + f1 + turnout05, data = clean_Data)

Pro tip: That trailing comma is key—it tells R to keep all columns, just filter the rows. If your earlier row deletion attempt failed, this might have been the missing piece!

Solution 3: Convert invalid values to NA and use R's missing value handling

Another standard R approach is to turn 98 and 99 into NA (missing values), then tell lm() to exclude those rows:

# Replace invalid e5 values with NA
Data$e5[Data$e5 %in% c(98, 99)] <- NA

# Run regression, excluding rows with NA in e5
PolReg <- lm(e5 ~ age + gender + f1 + turnout05, 
             data = Data, 
             na.action = na.exclude)

The na.action = na.exclude ensures rows with missing e5 values are dropped, and it also tracks these excluded rows for things like residual calculations later on.

All three methods will give you the same reliable regression results—pick the one that fits how you work best! If you struggled with row deletion before, double-check that you were specifying the row condition correctly (the %in% operator is way more reliable than stacking multiple != checks here).

内容的提问来源于stack exchange,提问作者Soppers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:42:15