You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何基于多列体长频次数据创建长度分箱表

R实现物种体长观测数据的分箱频次统计

原始数据结构

示例输入数据如下:

df <- data.frame(
  species = c("A","A","A","B","B","B"),
  station = c(1:3,1:3),
  CLS1 = 11:16,
  Freq1 = c(1,1,1,1,1,2),
  CLS2 = c(0, 14, 0, 16,0,17),
  Freq2 = c(0,1,0,1,0,2),
  CLS3 = c(0,0,0,18,0,20),
  Freq3 = c(0,0,0,1,0,1)
)

字段说明:

  • CLS1:每条观测记录的最小体长值
  • Freq1:对应CLS1体长的观测频次
  • 同站点若存在更大体长的观测值,会依次记录在CLS2/Freq2、CLS3/Freq3等成对列中,值为0代表该体长组无观测

预期输出

需要生成步长为2的体长分箱频次统计表,格式如下:

cnt_table <- data.frame(
  species = c("A","A","A","B","B","B"),
  station = c(1:3,1:3),
  X11_12 = c(1,1,0,0,0,0),
  X13_14 = c(0,1,1,1,0,0),
  X15_16 = c(0,0,0,1,1,2),
  X17_18 = c(0,0,0,1,0,0),
  X19_20 = c(0,0,0,0,0,1)
)

实现代码

使用dplyr+tidyr实现,逻辑为宽转长提取有效体长观测记录、生成分箱标签、汇总频次后转回宽表,后续新增CLSx/Freqx列时无需调整代码可自动适配:

library(dplyr)
library(tidyr)

cnt_table <- df %>%
  # 宽转长,配对所有CLS和Freq列
  pivot_longer(
    cols = matches("^CLS|^Freq"),
    names_to = c(".value", "group_id"),
    names_pattern = "(CLS|Freq)(\\d+)"
  ) %>%
  # 过滤体长为0的无效记录
  filter(CLS != 0) %>%
  # 按步长2生成分箱标签,匹配示例分箱区间
  mutate(
    bin_start = floor((CLS - 11)/2)*2 + 11,
    bin_end = bin_start + 1,
    bin_col = paste0("X", bin_start, "_", bin_end)
  ) %>%
  # 按物种、站点、分箱汇总频次
  group_by(species, station, bin_col) %>%
  summarise(freq = sum(Freq), .groups = "drop") %>%
  # 长转宽生成分箱列,缺失值填充0
  pivot_wider(names_from = bin_col, values_from = freq, values_fill = 0) %>%
  # 按分箱数值从小到大排列列
  select(species, station, order(as.numeric(gsub("X(\\d+)_.*", "\\1", names(.)[3:ncol(.)])))+2)

内容的提问来源于stack exchange,提问作者TKH_9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 12:33:30