You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中为分组数据按指定行分配整数序列?

问题描述

给定包含二元变量x的数据集:

df <- data.frame(id = c("a", "a", "a", "a", "b", "b", "b", "b"),
                 year = c("2001", "2002", "2003", "2004", "2001", "2002", "2003", "2004"),
                 x = c(0, 0, 1, 0, 0, 1, 1, 0))

数据集内容:

id year x
  a 2001 0
  a 2002 0
  a 2003 1
  a 2004 0
  b 2001 0
  b 2002 1
  b 2003 1
  b 2004 0

需要创建变量y,规则为:每个id分组中第一个x = 1的行y = 0,该行之前的行依次为递减负整数,之后的行依次为递增正整数,期望输出如下:

id year x  y
  a 2001 0 -2
  a 2002 0 -1
  a 2003 1  0
  a 2004 0  1
  b 2001 0 -1
  b 2002 1  0
  b 2003 1  1
  b 2004 0  2
解决方案

方法1:使用dplyr包(推荐,简洁易读)

借助dplyr的分组功能,先定位每组第一个x=1的行位置,再通过行索引计算y值:

library(dplyr)

df_result <- df %>%
  group_by(id) %>%
  mutate(
    # 获取每组第一个x=1的行的索引
    first_one_pos = which(x == 1)[1],
    # 当前行索引减去目标位置,得到y
    y = row_number() - first_one_pos
  ) %>%
  select(-first_one_pos) %>% # 移除临时变量
  ungroup()

# 输出结果
print(df_result)

方法2:使用Base R

通过拆分数据集、分组处理再合并的方式实现:

# 按id拆分数据集
df_groups <- split(df, df$id)

# 对每个分组计算y
df_processed <- lapply(df_groups, function(group) {
  first_one_pos <- which(group$x == 1)[1]
  group$y <- seq_len(nrow(group)) - first_one_pos
  return(group)
})

# 合并分组结果并重置行名
df_result <- do.call(rbind, df_processed)
rownames(df_result) <- NULL

# 输出结果
print(df_result)

两种方法均可得到符合要求的输出,其中dplyr方案更适合日常数据处理的流水线操作。

内容的提问来源于stack exchange,提问作者aerw4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 12:30:30