如何基于start和end列转换R语言DataFrame为条件位置标记表
问题描述
现有如下R语言DataFrame:
> df condition duration start end 1 A 2 3 4 2 B 3 8 10 3 A 2 7 8
需要将其转换为一张表格:生成从1到原数据最大end值的count列,为每个条件(A、B)标记0或1——当count处于该条件的start到end区间内时标记1,否则标记0。转换后的目标表格如下:
> df2 count A B 1 1 0 0 2 2 0 0 3 3 1 0 4 4 1 0 5 5 0 0 6 6 0 0 7 7 1 0 8 8 1 1 9 9 0 1 10 10 0 1
解决方案
方法一:使用tidyverse工具链
依赖tidyverse包,步骤清晰易读:
library(tidyverse) # 生成从1到最大end值的完整count序列 count_seq <- 1:max(df$end) # 交叉匹配count与条件,标记区间内的位置,再转宽表 df2 <- crossing(count = count_seq, condition = unique(df$condition)) %>% left_join(df, by = "condition") %>% # 判断当前count是否在对应条件的区间内 mutate(flag = ifelse(count >= start & count <= end, 1, 0)) %>% # 同一个count+condition组合取最大值(确保多区间的条件被正确标记) group_by(count, condition) %>% summarise(flag = max(flag), .groups = "drop") %>% # 转换为宽表格式 pivot_wider(names_from = condition, values_from = flag) %>% arrange(count) print(df2)
方法二:使用base R实现
无需额外加载包,适合轻量场景:
# 生成完整count序列 count_seq <- 1:max(df$end) # 获取所有唯一条件 conditions <- unique(df$condition) # 初始化全0的结果矩阵 result_mat <- matrix(0, nrow = length(count_seq), ncol = length(conditions)) colnames(result_mat) <- conditions rownames(result_mat) <- count_seq # 遍历每个条件的区间,将对应位置标记为1 for (i in seq(nrow(df))) { cond <- df$condition[i] start_pos <- df$start[i] end_pos <- df$end[i] result_mat[start_pos:end_pos, cond] <- 1 } # 转换为DataFrame并调整列顺序 df2 <- as.data.frame(result_mat) df2$count <- count_seq df2 <- df2[, c("count", conditions)] print(df2)
核心逻辑说明
两种方法的核心思路一致:
- 先生成覆盖所有位置的
count序列; - 对每个条件的所有区间,将区间内的
count位置标记为1,其余为0; - 整理成目标的宽表格式,确保同一个条件的多区间都被正确识别。
内容的提问来源于stack exchange,提问作者ghs101
相关产品推荐
相关产品推荐

