You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中按行累加坐标值生成新列的实现方法

问题描述

现有如下结构的DataFrame:

Sp1 start end  
A   100   1077 
B   2316  4088
B   26647 28746
C   450    789
D   23     499
D   45999  60000

需要添加new_start和new_end两列,规则如下:

  • 第一行的new_start、new_end与原start、end值一致;
  • 后续行的new_start为上一行new_end值加1;
  • new_end为当前行new_start加上(原end - 原start)的差值。

示例说明

第一行结果:

Sp1 start end     new_start  new_end 
A   100   1077    100        1077

第二行(B行)结果:

Sp1 start end     new_start  new_end 
A   100   1077    100        1077
B   2316  4088    1078       2850
  • 1078 = 1077 + 1
  • 2850 = 1078 + (4088 - 2316)

最终预期结果

Sp1 start end     new_start  new_end 
A   100   1077    100        1077
B   2316  4088    1078       2850
B   26647 28746   2851       4950
C   450   789     4951       5290

DataFrame的dput格式

structure(list(Sp1 = c("A", "B", "B", "C"), start = c(100L, 2316L, 
26647L, 450L), end = c(1077, 4088, 28746, 789)), class = "data.frame", row.names = c(NA, 
-4L))
解决方案

方法1:Base R实现

通过计算累积偏移量的方式高效实现,无需循环:

# 加载数据
df <- structure(list(Sp1 = c("A", "B", "B", "C"), start = c(100L, 2316L, 26647L, 450L), end = c(1077, 4088, 28746, 789)), class = "data.frame", row.names = c(NA, -4L))

# 计算每行的长度(end - start)
df$length <- df$end - df$start

# 计算new_start:第一行取原start,后续行基于上一行的长度+1累积
df$new_start <- df$start[1] + cumsum(c(0, df$length[-1] + 1))

# 计算new_end:new_start加上当前行长度
df$new_end <- df$new_start + df$length

# 删除临时的length列
df <- df[, !names(df) == "length"]

# 查看结果
print(df)

运行结果:

Sp1 start   end new_start new_end
1    A   100  1077       100    1077
2    B  2316  4088      1078    2850
3    B 26647 28746      2851    4950
4    C   450   789      4951    5290

方法2:dplyr(tidyverse)实现

用tidyverse工具链,结合lag函数处理行间依赖:

library(dplyr)

df <- structure(list(Sp1 = c("A", "B", "B", "C"), start = c(100L, 2316L, 26647L, 450L), end = c(1077, 4088, 28746, 789)), class = "data.frame", row.names = c(NA, -4L))

df <- df %>%
  mutate(
    length = end - start,
    new_start = case_when(
      row_number() == 1 ~ start,
      TRUE ~ lag(new_end) + 1
    ),
    new_end = new_start + length
  ) %>%
  select(-length) # 删除临时列

# 查看结果
print(df)

该方法通过case_when处理第一行的特殊逻辑,lag(new_end)获取上一行的new_end值,同样能得到预期结果。


内容的提问来源于stack exchange,提问作者chippycentra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 02:54:24