You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:按Condition分组后对Position列分箱并统计Info频次

R语言分箱分组统计解决方案

可以借助tidyverse或者data.table工具包高效完成需求,无需拆分数据框,直接通过分组统计实现:

方法一:使用tidyverse(推荐,代码可读性高)

完整代码

library(tidyverse)

# 生成测试数据
set.seed(123)
test <- data.frame(
  chr = rep("chr1", 30),
  position = sample(1:50, 30, replace = F),
  info = sample(c("X", "Y"), 30, replace = T),
  condition = sample(c("soft", "stiff"), 30, replace = T)
)

# 数据处理流程
test_processed <- test %>%
  # 计算分箱的起始和结束位置
  mutate(
    start = ((position - 1) %/% 10) * 10 + 1,
    end = start + 9
  ) %>%
  # 按分组维度统计info的出现次数
  group_by(chr, start, end, condition, info) %>%
  summarise(count = n(), .groups = "drop") %>%
  # 将info的X/Y转为列,缺失值填0
  pivot_wider(
    names_from = info,
    values_from = count,
    values_fill = 0
  ) %>%
  # 重命名列名以匹配需求
  rename(count_Y = Y, count_X = X) %>%
  # 补全所有可能的分箱+condition组合(避免缺失无数据的分组)
  complete(chr, start, end, condition, fill = list(count_Y = 0, count_X = 0)) %>%
  # 按分箱起始和condition排序
  arrange(start, condition)

# 查看结果
print(test_processed)

关键步骤说明

  1. 分箱计算:用((position - 1) %/% 10) * 10 + 1生成步长10的分箱起始值,end直接为起始值+9,确保区间是[1-10]、[11-20]等。
  2. 分组统计:按chr、分箱区间、condition和info分组,统计每个组合的数量。
  3. 格式转换:用pivot_wider将长格式转为需求的宽格式,values_fill=0保证无数据的类别显示0。
  4. 补全分组:complete函数补全所有可能的分组组合,避免某些分箱在特定condition下无数据时不显示。

方法二:使用data.table(适合大数据量,处理速度快)

完整代码

library(data.table)

# 生成测试数据
set.seed(123)
test <- data.frame(
  chr = rep("chr1", 30),
  position = sample(1:50, 30, replace = F),
  info = sample(c("X", "Y"), 30, replace = T),
  condition = sample(c("soft", "stiff"), 30, replace = T)
)

# 转为data.table格式
setDT(test)

# 计算分箱
test[, `:=`(
  start = ((position - 1) %/% 10) * 10 + 1,
  end = start + 9
)]

# 分组统计+转宽格式+补全分组
test_processed_dt <- test[, .N, by = .(chr, start, end, condition, info)] %>%
  dcast(chr + start + end + condition ~ info, value.var = "N", fill = 0) %>%
  setnames(c("Y", "X"), c("count_Y", "count_X")) %>%
  # 补全所有可能的分组组合
  .[CJ(chr = unique(chr), start = unique(start), end = unique(end), condition = unique(condition)), on = .(chr, start, end, condition)] %>%
  # 缺失值替换为0
  replace(is.na(.), 0) %>%
  # 排序
  setorder(start, condition)

# 查看结果
print(test_processed_dt)

两种方法都能输出符合需求的格式,无需手动拆分数据框,通过分组统计即可高效完成任务。

内容的提问来源于stack exchange,提问作者ZainNST

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 22:47:51