You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅在后续值为特定值时反向填充列中的NA?

问题:基于后续特定值反向填充纵向数据的缺失值

我有包含吸烟状态列的纵向患者数据,需要仅当患者后续记录为never_smoker时,反向填充该列的缺失值(NA)。由于tidyr::fill无法按值筛选,不能直接使用。

示例数据

df <- tibble::tribble(
  ~id, ~visit, ~smoking,
  1, 1, NA,
  1, 2, NA,
  1, 3, "never_smoker",
  2, 1, NA,
  2, 2, NA,
  2, 3, "current_smoker"
)

期望结果

expected_result <- tibble::tribble(
  ~id, ~visit, ~smoking,
  1, 1, "never_smoker",
  1, 2, "never_smoker",
  1, 3, "never_smoker",
  2, 1, NA,
  2, 2, NA,
  2, 3, "current_smoker"
)

现有解法(两次反转)

df %>%
    group_by(id) %>%
    mutate(smoking = rev(accumulate(rev(smoking), ~ ifelse(is.na(.y) & .x == "never_smoker", "never_smoker", .y))))

更优解决方案

方法1:基于最终状态直接替换(最直观)

核心逻辑是先获取每个患者的最终吸烟状态,仅当该状态为never_smoker时,替换该患者所有的NA值:

library(dplyr)

df %>%
  group_by(id) %>%
  mutate(
    final_smoking = last(smoking),
    smoking = ifelse(is.na(smoking) & final_smoking == "never_smoker", final_smoking, smoking)
  ) %>%
  select(-final_smoking) %>%
  ungroup()

方法2:结合tidyr::fill的条件填充

先标记需要填充的患者组,再对这些组进行反向填充:

library(dplyr)
library(tidyr)

df %>%
  group_by(id) %>%
  mutate(needs_fill = last(smoking) == "never_smoker") %>%
  group_by(id, needs_fill) %>%
  fill(smoking, .direction = "up") %>%
  ungroup() %>%
  select(-needs_fill)

这两种方法都避免了两次反转的操作,代码可读性更强,在处理大数据集时效率也更高。

内容的提问来源于stack exchange,提问作者s_alt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 19:59:52