predict.lm返回结果长度与测试集不符的原因解析
问题描述
通过选取数据集80%行索引拆分训练集与测试集,训练线性模型后用测试集预测时,predict.lm返回的预测结果长度与训练集一致(17290),而非测试集的4324,导致计算误差时出现长度不匹配警告。
相关代码及信息如下:
数据集拆分代码
# 移除无用变量 df <- df[,!(colnames(df) %in% c("sqm_lot", "sqm_lot15"))] # 拆分训练集与测试集 df_train <- df[1:as.integer(nrow(df)*0.8),] df_test <- df[as.integer(nrow(df)*0.8):nrow(df),] nrow(df_test) ## [1] 4324 nrow(df_train) ## [1] 17290
数据集基本信息
## 'data.frame': 21613 obs. of 16 variables: ## $ price_eur : num 206367 500340 167400 561720 474300 ... ## $ bedrooms : int 3 3 2 4 3 4 3 3 3 3 ... ## $ bathrooms : int 1 2 1 3 2 4 2 1 1 2 ... ## $ sqm_living : num 109.6 238.8 71.5 182.1 156.1 ... ## $ sqm_lot : num 525 673 929 464 751 ... ## $ floors : int 1 2 1 1 1 1 2 1 1 2 ... ## $ waterfront : int 0 0 0 0 0 0 0 0 0 0 ... ## $ view : int 0 0 0 0 0 0 0 0 0 0 ... ## $ condition : int 3 3 3 5 3 3 3 3 3 3 ... ## $ sqm_basement: num 0 37.2 0 84.5 0 ... ## $ yr_built : int 1955 1951 1933 1965 1987 2001 1995 1963 1960 2003 ... ## $ yr_renovated: int 0 1991 0 0 0 0 0 0 0 0 ... ## $ sqm_living15: num 124 157 253 126 167 ... ## $ sqm_lot15 : num 525 710 749 464 697 ...
模型训练与预测代码
x_train <- df_train[,!(colnames(df_train) %in% "price_eur")] y_train <- df_train$price_eur # 训练线性模型 modelo.mlineal <- lm(formula = price_eur ~ ., data = df_train) modelo.mlineal ## ## Call: ## lm(formula = price_eur ~ ., data = df_train) ## ## Coefficients: ## (Intercept) bedrooms bathrooms sqm_living sqm_lot ## 5.696e+06 -4.366e+04 5.886e+04 2.371e+03 3.063e-01 ## floors waterfront view condition sqm_basement ## 2.524e+04 5.327e+05 4.727e+04 2.125e+04 -3.477e+02 ## yr_built yr_renovated sqm_living15 sqm_lot15 ## -3.000e+03 1.571e+01 1.039e+03 -8.048e+00
# 提取测试集标签与特征 y_test <- df_test$price_eur x_test <- df_test[,!(colnames(df_test) %in% "price_eur")] # 尝试用测试集预测 prediccion_prueba <- modelo.mlineal %>% predict.lm(data = df_test ) print(length(prediccion_prueba)) ## [1] 17290 print(length(y_test)) ## [1] 4324 errores <- prediccion_prueba - y_test ## Warning in prediccion_prueba - y_test: length of larger object is not a multiple of the length of a smaller one
问题原因
- 参数名错误:
predict.lm的参数中没有data,指定预测数据集的正确参数名是newdata。由于参数名错误,df_test未被识别为预测用的新数据,predict.lm默认使用训练时的df_train生成预测结果,导致长度与训练集一致。 - (潜在问题):模型输出中出现了已删除的
sqm_lot和sqm_lot15变量,说明数据集处理与模型训练的代码执行可能存在顺序问题,需确保训练集和测试集中已移除目标变量。
解决办法
修正预测代码
将参数名改为newdata,两种可行写法:
写法一:保留管道符
prediccion_prueba <- modelo.mlineal %>% predict.lm(newdata = df_test)
写法二:常规调用(更直观)
prediccion_prueba <- predict(modelo.mlineal, newdata = df_test)
执行后,length(prediccion_prueba)会返回4324,与y_test长度匹配,计算误差时不再出现警告。
确认数据集一致性
重新运行数据集处理与模型训练的代码,确保df_train和df_test中已移除sqm_lot和sqm_lot15变量,避免模型训练与预测时的变量不匹配问题。
内容的提问来源于stack exchange,提问作者ffriast
相关产品推荐
相关产品推荐

