You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中对iris数据集按Species分类执行线性回归并获取统计量

问题:针对iris数据集按物种分组执行线性回归并提取关键统计量

假设我正在R中使用iris数据集:

data(iris)

summary(iris)
# 输出结果:
#   Sepal.Length    Sepal.Width     Petal.Length    Petal.Width   
#  Min.   :4.300   Min.   :2.000   Min.   :1.000   Min.   :0.100  
#  1st Qu.:5.100   1st Qu.:2.800   1st Qu.:1.600   1st Qu.:0.300  
#  Median :5.800   Median :3.000   Median :4.350   Median :1.300  
#  Mean   :5.843   Mean   :3.057   Mean   :3.758   Mean   :1.199  
#  3rd Qu.:6.400   3rd Qu.:3.300   3rd Qu.:5.100   3rd Qu.:1.800  
#  Max.   :7.900   Max.   :4.400   Max.   :6.900   Max.   :2.500  
#        Species  
#  setosa    :50  
#  versicolor:50  
#  virginica :50

我希望执行以Petal.Length为因变量、Sepal.Length为自变量的线性回归,如何在R中一次性针对每个Species类别执行该回归,并获取每个检验的P值、R²值和F值?


方法一:使用基础R的by()函数

用by()函数按Species分组处理数据,自定义函数提取所需统计量:

# 定义提取统计量的函数
extract_stats <- function(data) {
  # 拟合线性回归模型
  model <- lm(Petal.Length ~ Sepal.Length, data = data)
  mod_sum <- summary(model)
  
  # 提取各统计量
  r_squared <- mod_sum$r.squared
  f_stat <- mod_sum$fstatistic[1]
  f_pval <- pf(f_stat, mod_sum$fstatistic[2], mod_sum$fstatistic[3], lower.tail = FALSE)
  
  # 返回结构化结果
  data.frame(
    Species = unique(data$Species),
    R_squared = round(r_squared, 4),
    F_statistic = round(f_stat, 4),
    F_p_value = round(f_pval, 6)
  )
}

# 按物种分组执行并合并结果
result <- do.call(rbind, by(iris, iris$Species, extract_stats))
print(result)

运行后会得到如下结构化结果:

Species R_squared F_statistic F_p_value
setosa       setosa    0.0171       0.853    0.3584
versicolor versicolor    0.2792      18.910    0.0001
virginica   virginica    0.2038      12.560    0.0008

方法二:使用tidyverse工具链(dplyr + broom)

如果习惯tidyverse风格,用dplyr分组配合broom包提取模型结果更简洁:
首先安装并加载依赖包:

install.packages(c("dplyr", "broom"))
library(dplyr)
library(broom)

然后执行分组回归并提取统计量:

result <- iris %>%
  group_by(Species) %>%
  do(model = lm(Petal.Length ~ Sepal.Length, data = .)) %>%
  mutate(
    r_squared = glance(model)$r.squared,
    f_statistic = glance(model)$statistic,
    f_p_value = glance(model)$p.value
  ) %>%
  select(Species, r_squared, f_statistic, f_p_value) %>%
  ungroup() %>%
  mutate(across(c(r_squared, f_statistic), round, 4),
         across(f_p_value, round, 6))

print(result)

该方法会输出与基础R方法一致的结构化结果,更便于后续数据处理。


内容的提问来源于stack exchange,提问作者user18443305

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 01:05:24