You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中使用线性回归填充数据框var1的缺失值(NA)

Fill Missing Values in var1 Using Linear Regression in R

Got it, let's work through this problem step by step to fill those missing var1 values with a linear regression model. Looking at your dataset, var2 is completely intact, and there’s a clear linear relationship between var1 and var2 in the non-missing rows—perfect for regression-based imputation.

Step 1: Explore the Variable Relationship

First, let’s confirm the relationship between var1 and var2 (since var2 has no NAs, it’s our best predictor here):

# Load your dataset
df <- structure(list(time = c(15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26), var1 = c(20.4, 31.5, NA, 53.7, 64.8, NA, NA, NA, NA, 120.3, NA, 142.5), var2 = c(30.6, 47.25, 63.9, 80.55, 97.2, 113.85, 130.5, 147.15, 163.8, 180.45, 197.1, 213.75)), .Names = c("time", "var1", "var2"), row.names = c(NA, -12L), class = c("tbl_df", "tbl", "data.frame"))

# Check correlation between var1 and var2 (only using complete observations)
cor(df$var1, df$var2, use = "complete.obs")

# Plot the relationship to visualize
plot(df$var2, df$var1, na.rm = TRUE, xlab = "var2", ylab = "var1", main = "var1 vs var2")
abline(lm(var1 ~ var2, data = df, na.action = na.exclude), col = "red")

You’ll see the correlation is nearly 1, which means the linear relationship is extremely strong—ideal for imputation.

Step 2: Build the Linear Regression Model

Next, we’ll fit a linear regression model using the complete pairs of var1 and var2:

# Fit the model (exclude NA values during fitting)
lm_model <- lm(var1 ~ var2, data = df, na.action = na.exclude)

# Check model summary to confirm fit
summary(lm_model)

From your sample data, the model will simplify to var1 = 0.6 * var2—all non-missing var1 values are exactly 60% of var2, so the model will be a perfect fit here.

Step 3: Predict and Fill Missing Values

Now use the model to predict the missing var1 values and update your dataset:

# Predict missing values and fill them in
df$var1[is.na(df$var1)] <- predict(lm_model, newdata = df[is.na(df$var1), ])

# View the updated dataset
df

This will replace all NAs in var1 with values predicted from the linear relationship with var2.

Alternative: Using Time as a Predictor

If you wanted to use time instead (or alongside var2), you could fit a model like lm(var1 ~ time, data = df, na.action = na.exclude) and predict similarly. But since var2 is directly tied to time and has a stronger direct relationship with var1, using var2 is more accurate here.

内容的提问来源于stack exchange,提问作者Kathiravan Meeran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:21:00