You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按采样日期计算化合物占比并对3棵树数据取平均值?

数据与需求说明

原始数据格式

DateTreeCompoundCompound_mg%mg
13.1aC215
13.1bC214x
13.1cC219
20.2aC216
20.2bC215x
20.2cC2110
13.1aC236
13.1bC236x
13.1cC2310
20.2aC235
20.2bC234x
20.2cC239

仅最后一列%mg为待计算列。对3棵树(N=3)采样,在不同日期提取多种化合物,需计算:

  • 每个时间点下,C21占所有化合物的%mg比例(按3棵树取平均值)
  • C23的计算逻辑与C21完全一致

示例数据集(R代码)

Df <- data.frame(
  rating = 1:12,
  Date = c('13.1','13.1','13.1','20.2','20.2','20.2','13.1','13.1','13.1','20.2','20.2','20.2'),
  Tree = c('a', 'b', 'c', 'a', 'b', 'c', 'a', 'b', 'c', 'a', 'b', 'c'),
  Compound = c("C21","C21","C21","C21","C21","C21","C23","C23","C23","C23","C23","C23"),
  Compound_mg = c(5, 4, 9, 6, 5, 10, 6, 6, 10, 5, 4, 9)
)

尝试的代码(存在问题)

Df2 <- Df %>%
group_by(Compound) %>%
mutate(perc = Compound(mg) / sum(Compound(mg))

解决方案

核心逻辑是先计算单棵树在单个日期下的化合物占比,再按日期+化合物分组取3棵树的平均占比:

library(dplyr)

# 1. 计算单棵树在单个日期下的化合物占比
Df_with_perc <- Df %>%
  group_by(Date, Tree) %>%
  mutate(perc = Compound_mg / sum(Compound_mg)) %>%  # 单树单日的化合物占比
  ungroup()

# 2. 按日期+化合物分组,取3棵树的平均占比(转成百分比格式)
avg_perc <- Df_with_perc %>%
  group_by(Date, Compound) %>%
  summarise(avg_%mg = mean(perc) * 100, .groups = "drop")

# 可选:将平均占比合并回原始数据,方便对应查看
Df_final <- Df_with_perc %>%
  left_join(avg_perc, by = c("Date", "Compound"))

结果说明

以13.1日期的C21为例:

  • 树a的C21占比:5/(5+6)≈45.45%
  • 树b的C21占比:4/(4+6)=40%
  • 树c的C21占比:9/(9+10)≈47.37%
  • 平均占比:(45.45+40+47.37)/3≈44.27%,与代码计算结果一致

内容的提问来源于stack exchange,提问作者Sofi A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 06:37:06