You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言tidyverse中按指定条件向前填充数据的实现求助

问题1:基于cate列规则填充result列

原始数据集

d <- data.frame(
  id = c(1,1,1,2,2,2,2,3,3,3,3,3,3,3,4,4,4,4,4),
  cate = c("yes","yes","no","no","yes","yes","no","no","yes","no","yes","yes","no","yes","yes","yes","yes","no","no"),
  result = c(1,NA,NA,NA,1,NA,NA,NA,NA,NA,1,NA,NA,NA,1,NA,NA,NA,NA)
)

需求说明

将result列中的值1向下填充,直至遇到cate列的第一个no;第一个no之后的行不再填充,保持NA。

原代码问题

原代码仅使用fill(result, .direction = 'down')会无限制向下填充,无法在遇到第一个no后停止,不符合需求。

正确实现代码

library(tidyverse)

d2 <- d %>%
  group_by(id) %>%
  mutate(
    # 标记每个result=1的位置到第一个cate=no的区间
    fill_flag = ifelse(result == 1, 
                       cumsum(cate == "no") == lag(cumsum(cate == "no"), default = 0),
                       FALSE),
    # 填充区间内的NA为1,其他保持原值
    result = ifelse(fill_flag | result == 1, 1, result)
  ) %>%
  select(-fill_flag) %>% # 移除辅助列
  ungroup()

代码逻辑解释

  1. 按id分组,确保每个id内独立处理
  2. 用cumsum(cate == "no")计算每个位置之前的no出现次数,通过比较当前次数和前一次次数,标记出从result=1到第一个no的区间
  3. 将区间内的NA替换为1,其他位置保持原始值

验证输出

运行代码后得到的结果与期望输出一致:

print(d2)
#    id cate result
# 1   1   yes      1
# 2   1   yes      1
# 3   1    no      1
# 4   2    no     NA
# 5   2   yes      1
# 6   2   yes      1
# 7   2    no      1
# 8   3    no     NA
# 9   3   yes     NA
# 10  3    no     NA
# 11  3   yes      1
# 12  3   yes      1
# 13  3    no      1
# 14  3   yes     NA
# 15  4   yes      1
# 16  4   yes      1
# 17  4   yes      1
# 18  4    no      1
# 19  4    no     NA

问题2:基于stat列规则填充ind列

原始数据集

dat <- data.frame(
  id=c(1,1,1,1,2,2,2,2,2,3,3,3,4,4,4,4), 
  stat=c('sup','unsup','unsup','sup','unsup','unsup','unsup','unsup','sup','sup','unsup','unsup','sup','sup','unsup','unsup'),
  ind=c(NA,1,NA,NA,1,1,1,NA,NA,NA,1,NA,NA,NA,1,NA)
)

需求说明

将ind列中的值1向上填充,直至遇到stat列的第一个sup;第一个sup之前的行不再填充,保持NA。

正确实现代码

dat2 <- dat %>%
  group_by(id) %>%
  mutate(
    # 从后往前标记ind=1到第一个stat=sup的区间
    fill_flag = ifelse(ind == 1,
                       rev(cumsum(rev(stat == "sup"))) == lag(rev(cumsum(rev(stat == "sup"))), default = 0),
                       FALSE),
    # 填充区间内的NA为1,其他保持原值
    ind = ifelse(fill_flag | ind == 1, 1, ind)
  ) %>%
  select(-fill_flag) %>%
  ungroup()

代码逻辑解释

  1. 按id分组,每个id独立处理
  2. 用rev(cumsum(rev(stat == "sup")))从后往前计算sup的累积次数,标记出从ind=1到上方第一个sup的区间
  3. 将区间内的NA替换为1,其他位置保持原始值

验证输出

运行代码后得到的结果与期望输出一致:

print(dat2)
#    id  stat ind
# 1   1   sup  NA
# 2   1 unsup   1
# 3   1 unsup   1
# 4   1   sup   1
# 5   2 unsup   1
# 6   2 unsup   1
# 7   2 unsup   1
# 8   2 unsup   1
# 9   2   sup   1
# 10  3   sup  NA
# 11  3 unsup   1
# 12  3 unsup   1
# 13  4   sup  NA
# 14  4   sup  NA
# 15  4 unsup   1
# 16  4 unsup   1

内容的提问来源于stack exchange,提问作者Yebelay Berehan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 12:36:19