如何仅在后续值为特定值时反向填充列中的NA?
问题:基于后续特定值反向填充纵向数据的缺失值
我有包含吸烟状态列的纵向患者数据,需要仅当患者后续记录为never_smoker时,反向填充该列的缺失值(NA)。由于tidyr::fill无法按值筛选,不能直接使用。
示例数据
df <- tibble::tribble( ~id, ~visit, ~smoking, 1, 1, NA, 1, 2, NA, 1, 3, "never_smoker", 2, 1, NA, 2, 2, NA, 2, 3, "current_smoker" )
期望结果
expected_result <- tibble::tribble( ~id, ~visit, ~smoking, 1, 1, "never_smoker", 1, 2, "never_smoker", 1, 3, "never_smoker", 2, 1, NA, 2, 2, NA, 2, 3, "current_smoker" )
现有解法(两次反转)
df %>% group_by(id) %>% mutate(smoking = rev(accumulate(rev(smoking), ~ ifelse(is.na(.y) & .x == "never_smoker", "never_smoker", .y))))
更优解决方案
方法1:基于最终状态直接替换(最直观)
核心逻辑是先获取每个患者的最终吸烟状态,仅当该状态为never_smoker时,替换该患者所有的NA值:
library(dplyr) df %>% group_by(id) %>% mutate( final_smoking = last(smoking), smoking = ifelse(is.na(smoking) & final_smoking == "never_smoker", final_smoking, smoking) ) %>% select(-final_smoking) %>% ungroup()
方法2:结合tidyr::fill的条件填充
先标记需要填充的患者组,再对这些组进行反向填充:
library(dplyr) library(tidyr) df %>% group_by(id) %>% mutate(needs_fill = last(smoking) == "never_smoker") %>% group_by(id, needs_fill) %>% fill(smoking, .direction = "up") %>% ungroup() %>% select(-needs_fill)
这两种方法都避免了两次反转的操作,代码可读性更强,在处理大数据集时效率也更高。
内容的提问来源于stack exchange,提问作者s_alt
相关产品推荐
相关产品推荐

