You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中拟合ANCOVA时,lm()与aov()为何输出结果不同?

ANCOVA在R中lm()与aov()输出差异的原因

问题描述

我了解到ANCOVA本质上就是多元线性回归的另一种表述,二者在数学上完全等价。但在R中用lm()和aov()拟合同一个ANCOVA模型时,输出结果看起来不一样,这是为什么?

附上我的代码及输出:

library(dplyr)
library(stats)

# 将mtcars中的cyl转为因子
mtcars <-
  mtcars %>%
  tibble() %>%
  mutate(cyl = factor(cyl))

# 用lm()拟合ANCOVA
lm_cars <-
  lm(mpg ~ wt + cyl, mtcars)

summary(lm_cars)

# lm()的输出结果:
Call:
lm(formula = mpg ~ wt + cyl, data = mtcars)

Residuals:
    Min      1Q  Median      3Q     Max 
-4.5890 -1.2357 -0.5159  1.3845  5.7915 

Coefficients:
            Estimate Std. Error t value Pr(>|t|)    
(Intercept)  33.9908     1.8878  18.006  < 2e-16 ***
wt           -3.2056     0.7539  -4.252 0.000213 ***
cyl6         -4.2556     1.3861  -3.070 0.004718 ** 
cyl8         -6.0709     1.6523  -3.674 0.000999 ***
---
Signif. codes:  0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Residual standard error: 2.557 on 28 degrees of freedom
Multiple R-squared:  0.8374,    Adjusted R-squared:   0.82 
F-statistic: 48.08 on 3 and 28 DF,  p-value: 3.594e-11

# 用aov()拟合ANCOVA
aov_cars <-
  aov(mpg ~ wt + cyl, mtcars)

summary(aov_cars)

TukeyHSD(aov_cars, "cyl")

# aov()的输出结果:
> summary(aov_cars)
            Df Sum Sq Mean Sq F value   Pr(>F)    
wt           1  847.7   847.7 129.665 5.08e-12 ***
cyl          2   95.3    47.6   7.286  0.00284 ** 
Residuals   28  183.1     6.5                     
---
Signif. codes:  0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
> TukeyHSD(aov_cars, "cyl")
  Tukey multiple comparisons of means
    95% family-wise confidence level

Fit: aov(formula = mpg ~ wt + cyl, data = mtcars)

$cyl
           diff       lwr       upr     p adj
6-4 -2.47730226 -5.536223 0.5816181 0.1298897
8-4 -2.40595373 -4.955054 0.1431466 0.0671811
8-6  0.07134853 -2.857345 3.0000418 0.9979988

Warning message:
In replications(paste("~", xx), data = mf) : non-factors ignored: wt

核心结论:两个函数拟合的是完全相同的模型

lm()和aov()在底层都是用普通最小二乘法(OLS)拟合线性模型,数学上没有任何差异,差异仅在于输出结果的展示形式,以及附加功能的不同。

1. 输出差异的具体原因

  • summary(lm())的定位:聚焦于模型的系数解读,展示每个自变量(包括因子变量生成的哑变量)的回归系数、标准误、t检验结果,同时给出模型整体的拟合优度(R²)和整体F检验结果。
    比如你看到的cyl6和cyl8系数,是分别与基准水平cyl4的差值,t检验是判断单个水平与基准水平的差异是否显著。

  • summary(aov())的定位:聚焦于方差分解,把因子变量作为一个整体,计算其对因变量的解释方差占比,用F检验判断这个因子整体是否对因变量有显著影响。
    比如cyl的DF=2(因为3个水平,自由度为水平数-1),对应的F检验是判断“所有cyl水平对mpg的影响是否整体显著”,而不是单个水平的差异。

2. 验证模型一致性的方法

你可以通过以下代码验证两个模型完全一致:

# 查看aov模型的系数,和lm的系数完全相同
coef(aov_cars)

# 对lm模型做方差分析,得到和aov summary完全一致的结果
anova(lm_cars)

3. TukeyHSD的差异说明

aov()可以直接配合TukeyHSD()做事后多重比较,而lm()需要借助emmeans等包实现,但这只是工具功能的差异,模型本身的参数估计完全相同。
另外你看到的警告“non-factors ignored: wt”,是因为TukeyHSD()只针对因子变量做多重比较,会自动忽略连续变量wt,属于正常提示。

内容的提问来源于stack exchange,提问作者dalexco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 23:45:27