R中多元多项式回归预测异常:结果维度与测试集不符
这问题我太熟了!核心原因就是你调用predict()时的参数写法完全错了——你直接把一堆poly()的结果加起来传进去,这不是在给模型指定测试数据集,而是在做矩阵/向量加法,最终得到的是和训练集长度一致的特征集合,所以predict()默认用训练集的数据计算,结果自然和训练集维度一样。
反观你用简单线性回归时的写法predict(model_lm, test)是正确的:你把整个测试数据集传进去,predict()会自动根据模型的公式去提取对应变量、生成需要的特征。
为什么多项式回归这么写不行?
当你在lm()里写poly(training$x1, degree=2, raw=TRUE)时,模型确实会生成对应的二次项特征,但predict.lm()的newdata参数需要的是包含原始变量的数据框,而不是你手动计算好的特征矩阵。你直接传poly(test$x1,...)+...,函数根本识别不出这是测试数据,反而会用训练集的特征来计算预测值。
两种解决方案
方案一:用公式+数据框的标准写法(推荐)
建模时不要直接用training$x1这种形式,而是通过data参数指定数据集,让模型自动关联变量。预测时直接传测试数据集即可,predict()会自动帮你生成对应的多项式特征:
# 构建多项式回归模型,用data参数指定训练集 model_poly = lm(y ~ poly(x1, degree=2, raw=TRUE) + poly(x2, degree=2, raw=TRUE) + poly(x3, degree=2, raw=TRUE) + poly(x4, degree=2, raw=TRUE) + poly(x5, degree=2, raw=TRUE) + poly(x6, degree=2, raw=TRUE) + poly(x7, degree=2, raw=TRUE) + poly(x8, degree=2, raw=TRUE) + poly(x9, degree=2, raw=TRUE) + poly(x10, degree=2, raw=TRUE), data = training) # 预测时直接传入测试数据集 poly_predictions = predict(model_poly, newdata = test)
这种写法不仅简洁,还能避免手动生成特征时可能出现的列名不匹配、维度不一致等问题。
方案二:手动生成多项式特征数据框
如果你一定要手动处理特征,可以先把训练集和测试集的多项式特征整理成结构一致的数据框,再建模和预测:
# 生成训练集特征数据框,包含响应变量y train_poly_features = data.frame( y = training$y, poly(x1, degree=2, raw=TRUE), poly(x2, degree=2, raw=TRUE), poly(x3, degree=2, raw=TRUE), poly(x4, degree=2, raw=TRUE), poly(x5, degree=2, raw=TRUE), poly(x6, degree=2, raw=TRUE), poly(x7, degree=2, raw=TRUE), poly(x8, degree=2, raw=TRUE), poly(x9, degree=2, raw=TRUE), poly(x10, degree=2, raw=TRUE) ) # 基于特征数据框建模 model_poly = lm(y ~ ., data = train_poly_features) # 生成测试集特征数据框(注意不要包含y,列名要和训练集完全一致) test_poly_features = data.frame( poly(test$x1, degree=2, raw=TRUE), poly(test$x2, degree=2, raw=TRUE), poly(test$x3, degree=2, raw=TRUE), poly(test$x4, degree=2, raw=TRUE), poly(test$x5, degree=2, raw=TRUE), poly(test$x6, degree=2, raw=TRUE), poly(test$x7, degree=2, raw=TRUE), poly(test$x8, degree=2, raw=TRUE), poly(test$x9, degree=2, raw=TRUE), poly(test$x10, degree=2, raw=TRUE) ) # 预测 poly_predictions = predict(model_poly, newdata = test_poly_features)
这种方法要特别注意训练集和测试集的特征列名必须完全相同,否则predict()会报错找不到变量。
关键知识点回顾
predict.lm()的newdata参数要求传入一个包含模型公式中所有原始变量的数据框,函数会自动根据公式生成所需的衍生特征(比如多项式、交互项)。如果你直接传计算好的特征矩阵,函数会默认使用训练集的model.frame来计算,这就是你得到训练集维度预测结果的根本原因。
内容的提问来源于stack exchange,提问作者drakedog

