You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按plot分组计算dataframe中各物种列的均值

嗨,我来帮你搞定这个按plot分组计算物种均值的问题!

首先先确认下你的原始数据:

df=data.frame(plot=c(1000, 1000, 1000, 1005, 1005, 1005, 1009, 1009, 1009), 
              speciesA=c(5, 0.5, 10, 7, 8, 45, 0.2, 3, 17), 
              speciesB = c(1, 11, 46, 98, 0.2, 14, 40, 37, 22), 
              speciesC = c(0.7, 72, 17, 0, 14, 8, 0, 9, 0.9))

你想要的是按plot分组后,计算每个物种列的均值,不过你给出的目标df2里有两个小笔误:比如1009地块的speciesA均值应该是(0.2+3+17)/3≈6.4(不是20.2),1000地块的speciesC均值是(0.7+72+17)/3=29.9(不是89.7),咱们按正确的计算结果来实现~

方法一:用dplyr(更直观易读)

你之前用的summarise_each是dplyr旧版本的函数,现在已经被废弃了,这大概率是程序没响应的原因。现在dplyr推荐用summarise()结合across()来批量处理多列:

library(dplyr)

# 分组计算均值,选中所有以species开头的列,同时处理可能的缺失值
df2 <- df %>%
  group_by(plot) %>%
  summarise(across(starts_with("species"), mean, na.rm = TRUE)) %>%
  as.data.frame()  # 可选:把tibble转回普通data.frame格式

# 查看结果
print(df2)

运行后会得到:

plot speciesA speciesB speciesC
1  1000  5.166667 19.33333 29.90000
2  1005 20.000000 37.40000  7.33333
3  1009  6.400000 33.00000  3.30000

方法二:用data.table(处理大数据更快)

你之前只会用data.table拼接字符串,其实只要把拼接的函数换成mean就行,非常简单:

library(data.table)

# 先把普通data.frame转换成data.table格式
setDT(df)

# 按plot分组,对每个分组的子数据集(.SD)的所有列计算均值
df2 <- df[, lapply(.SD, mean, na.rm = TRUE), by = plot]

# 查看结果
print(df2)

这个方法得到的结果和dplyr完全一致,而且如果你的数据量很大,data.table的速度会更有优势。

这样两种方法都能完美实现你想要的分组均值计算啦!

内容的提问来源于stack exchange,提问作者Kactus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:23:39