在R中使用线性回归填充数据框var1的缺失值(NA)
var1 Using Linear Regression in R Got it, let's work through this problem step by step to fill those missing var1 values with a linear regression model. Looking at your dataset, var2 is completely intact, and there’s a clear linear relationship between var1 and var2 in the non-missing rows—perfect for regression-based imputation.
Step 1: Explore the Variable Relationship
First, let’s confirm the relationship between var1 and var2 (since var2 has no NAs, it’s our best predictor here):
# Load your dataset df <- structure(list(time = c(15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26), var1 = c(20.4, 31.5, NA, 53.7, 64.8, NA, NA, NA, NA, 120.3, NA, 142.5), var2 = c(30.6, 47.25, 63.9, 80.55, 97.2, 113.85, 130.5, 147.15, 163.8, 180.45, 197.1, 213.75)), .Names = c("time", "var1", "var2"), row.names = c(NA, -12L), class = c("tbl_df", "tbl", "data.frame")) # Check correlation between var1 and var2 (only using complete observations) cor(df$var1, df$var2, use = "complete.obs") # Plot the relationship to visualize plot(df$var2, df$var1, na.rm = TRUE, xlab = "var2", ylab = "var1", main = "var1 vs var2") abline(lm(var1 ~ var2, data = df, na.action = na.exclude), col = "red")
You’ll see the correlation is nearly 1, which means the linear relationship is extremely strong—ideal for imputation.
Step 2: Build the Linear Regression Model
Next, we’ll fit a linear regression model using the complete pairs of var1 and var2:
# Fit the model (exclude NA values during fitting) lm_model <- lm(var1 ~ var2, data = df, na.action = na.exclude) # Check model summary to confirm fit summary(lm_model)
From your sample data, the model will simplify to var1 = 0.6 * var2—all non-missing var1 values are exactly 60% of var2, so the model will be a perfect fit here.
Step 3: Predict and Fill Missing Values
Now use the model to predict the missing var1 values and update your dataset:
# Predict missing values and fill them in df$var1[is.na(df$var1)] <- predict(lm_model, newdata = df[is.na(df$var1), ]) # View the updated dataset df
This will replace all NAs in var1 with values predicted from the linear relationship with var2.
Alternative: Using Time as a Predictor
If you wanted to use time instead (or alongside var2), you could fit a model like lm(var1 ~ time, data = df, na.action = na.exclude) and predict similarly. But since var2 is directly tied to time and has a stronger direct relationship with var1, using var2 is more accurate here.
内容的提问来源于stack exchange,提问作者Kathiravan Meeran

