在R中拟合ANCOVA时,lm()与aov()为何输出结果不同?
ANCOVA在R中lm()与aov()输出差异的原因
问题描述
我了解到ANCOVA本质上就是多元线性回归的另一种表述,二者在数学上完全等价。但在R中用lm()和aov()拟合同一个ANCOVA模型时,输出结果看起来不一样,这是为什么?
附上我的代码及输出:
library(dplyr) library(stats) # 将mtcars中的cyl转为因子 mtcars <- mtcars %>% tibble() %>% mutate(cyl = factor(cyl)) # 用lm()拟合ANCOVA lm_cars <- lm(mpg ~ wt + cyl, mtcars) summary(lm_cars) # lm()的输出结果: Call: lm(formula = mpg ~ wt + cyl, data = mtcars) Residuals: Min 1Q Median 3Q Max -4.5890 -1.2357 -0.5159 1.3845 5.7915 Coefficients: Estimate Std. Error t value Pr(>|t|) (Intercept) 33.9908 1.8878 18.006 < 2e-16 *** wt -3.2056 0.7539 -4.252 0.000213 *** cyl6 -4.2556 1.3861 -3.070 0.004718 ** cyl8 -6.0709 1.6523 -3.674 0.000999 *** --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 Residual standard error: 2.557 on 28 degrees of freedom Multiple R-squared: 0.8374, Adjusted R-squared: 0.82 F-statistic: 48.08 on 3 and 28 DF, p-value: 3.594e-11 # 用aov()拟合ANCOVA aov_cars <- aov(mpg ~ wt + cyl, mtcars) summary(aov_cars) TukeyHSD(aov_cars, "cyl") # aov()的输出结果: > summary(aov_cars) Df Sum Sq Mean Sq F value Pr(>F) wt 1 847.7 847.7 129.665 5.08e-12 *** cyl 2 95.3 47.6 7.286 0.00284 ** Residuals 28 183.1 6.5 --- Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1 > TukeyHSD(aov_cars, "cyl") Tukey multiple comparisons of means 95% family-wise confidence level Fit: aov(formula = mpg ~ wt + cyl, data = mtcars) $cyl diff lwr upr p adj 6-4 -2.47730226 -5.536223 0.5816181 0.1298897 8-4 -2.40595373 -4.955054 0.1431466 0.0671811 8-6 0.07134853 -2.857345 3.0000418 0.9979988 Warning message: In replications(paste("~", xx), data = mf) : non-factors ignored: wt
核心结论:两个函数拟合的是完全相同的模型
lm()和aov()在底层都是用普通最小二乘法(OLS)拟合线性模型,数学上没有任何差异,差异仅在于输出结果的展示形式,以及附加功能的不同。
1. 输出差异的具体原因
summary(lm())的定位:聚焦于模型的系数解读,展示每个自变量(包括因子变量生成的哑变量)的回归系数、标准误、t检验结果,同时给出模型整体的拟合优度(R²)和整体F检验结果。
比如你看到的cyl6和cyl8系数,是分别与基准水平cyl4的差值,t检验是判断单个水平与基准水平的差异是否显著。summary(aov())的定位:聚焦于方差分解,把因子变量作为一个整体,计算其对因变量的解释方差占比,用F检验判断这个因子整体是否对因变量有显著影响。
比如cyl的DF=2(因为3个水平,自由度为水平数-1),对应的F检验是判断“所有cyl水平对mpg的影响是否整体显著”,而不是单个水平的差异。
2. 验证模型一致性的方法
你可以通过以下代码验证两个模型完全一致:
# 查看aov模型的系数,和lm的系数完全相同 coef(aov_cars) # 对lm模型做方差分析,得到和aov summary完全一致的结果 anova(lm_cars)
3. TukeyHSD的差异说明
aov()可以直接配合TukeyHSD()做事后多重比较,而lm()需要借助emmeans等包实现,但这只是工具功能的差异,模型本身的参数估计完全相同。
另外你看到的警告“non-factors ignored: wt”,是因为TukeyHSD()只针对因子变量做多重比较,会自动忽略连续变量wt,属于正常提示。
内容的提问来源于stack exchange,提问作者dalexco
相关产品推荐
相关产品推荐

