You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr与kruskal.test分析iris数据集时遇错误的技术咨询

解决iris数据集Kruskal-Wallis检验的报错问题

错误原因

你代码里的group_by(Species)是问题核心:分组后每个子数据集仅包含单一Species类别的观测,而Kruskal-Wallis检验要求至少两个不同的组才能计算统计量,因此触发all observations are in the same group错误。

修正后的代码

方式一:循环+dplyr

直接在整个数据集上执行检验,无需分组:

library(dplyr)

for (col in colnames(iris)[1:4]) {
  iris %>%
    summarise(
      zone = col,
      pvalue = kruskal.test(.data[[col]] ~ Species)$p.value
    ) %>%
    print()
}
  • 用.data[[col]]替代get(col),更符合tidyverse规范,避免环境变量冲突;
  • 直接指定colnames(iris)[1:4]筛选数值列,比1:ncol(iris)-1更直观。

方式二:purrr批量处理(更简洁)

用purrr遍历所有数值列,一次性输出所有检验结果:

library(dplyr)
library(purrr)

iris %>%
  select(where(is.numeric)) %>%
  map_dfr(function(col) {
    tibble(
      zone = cur_column(),
      pvalue = kruskal.test(col ~ iris$Species)$p.value
    )
  })
  • select(where(is.numeric))动态筛选数值列,适配任意包含数值变量的数据集;
  • map_dfr自动将结果合并为一个数据框,无需循环打印。

内容的提问来源于stack exchange,提问作者david

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 13:50:14