在R中按行累加坐标值生成新列的实现方法
问题描述
现有如下结构的DataFrame:
Sp1 start end A 100 1077 B 2316 4088 B 26647 28746 C 450 789 D 23 499 D 45999 60000
需要添加new_start和new_end两列,规则如下:
- 第一行的
new_start、new_end与原start、end值一致; - 后续行的
new_start为上一行new_end值加1; new_end为当前行new_start加上(原end- 原start)的差值。
示例说明
第一行结果:
Sp1 start end new_start new_end A 100 1077 100 1077
第二行(B行)结果:
Sp1 start end new_start new_end A 100 1077 100 1077 B 2316 4088 1078 2850
1078 = 1077 + 12850 = 1078 + (4088 - 2316)
最终预期结果
Sp1 start end new_start new_end A 100 1077 100 1077 B 2316 4088 1078 2850 B 26647 28746 2851 4950 C 450 789 4951 5290
DataFrame的dput格式
structure(list(Sp1 = c("A", "B", "B", "C"), start = c(100L, 2316L, 26647L, 450L), end = c(1077, 4088, 28746, 789)), class = "data.frame", row.names = c(NA, -4L))
解决方案
方法1:Base R实现
通过计算累积偏移量的方式高效实现,无需循环:
# 加载数据 df <- structure(list(Sp1 = c("A", "B", "B", "C"), start = c(100L, 2316L, 26647L, 450L), end = c(1077, 4088, 28746, 789)), class = "data.frame", row.names = c(NA, -4L)) # 计算每行的长度(end - start) df$length <- df$end - df$start # 计算new_start:第一行取原start,后续行基于上一行的长度+1累积 df$new_start <- df$start[1] + cumsum(c(0, df$length[-1] + 1)) # 计算new_end:new_start加上当前行长度 df$new_end <- df$new_start + df$length # 删除临时的length列 df <- df[, !names(df) == "length"] # 查看结果 print(df)
运行结果:
Sp1 start end new_start new_end 1 A 100 1077 100 1077 2 B 2316 4088 1078 2850 3 B 26647 28746 2851 4950 4 C 450 789 4951 5290
方法2:dplyr(tidyverse)实现
用tidyverse工具链,结合lag函数处理行间依赖:
library(dplyr) df <- structure(list(Sp1 = c("A", "B", "B", "C"), start = c(100L, 2316L, 26647L, 450L), end = c(1077, 4088, 28746, 789)), class = "data.frame", row.names = c(NA, -4L)) df <- df %>% mutate( length = end - start, new_start = case_when( row_number() == 1 ~ start, TRUE ~ lag(new_end) + 1 ), new_end = new_start + length ) %>% select(-length) # 删除临时列 # 查看结果 print(df)
该方法通过case_when处理第一行的特殊逻辑,lag(new_end)获取上一行的new_end值,同样能得到预期结果。
内容的提问来源于stack exchange,提问作者chippycentra
相关产品推荐
相关产品推荐

