开发按identifier清零Positivo后21天内have列的函数问题
问题描述
按identifier(标识)将have列(两次采集间隔天数)在Positivo(阳性)结果后21天内的值清零,无需处理Positivo行本身,仅处理之后21天的采集日期。已添加WANT列作为预期结果。
示例数据
dput(df) structure(list(result = c("Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Positivo", "Positivo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Positivo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo", "Negativo"), identifier = c("a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b"), date_collection = structure(c(1631059200, 1631664000, 1632355200, 1632787200, 1634601600, 1635292800, 1635811200, 1636416000, 1637107200, 1637712000, 1638144000, 1638748800, 1639440000, 1640044800, 1640736000, 1640044800, 1640649600, 1641254400, 1641859200, 1643241600, 1645142400, 1645660800, 1646352000, 1646784000, 1647388800, 1648080000, 1648512000, 1649203200, 1649721600, 1650326400, 1650931200 ), class = c("POSIXct", "POSIXt"), tzone = "UTC"), have = c(0, 7, 8, 5, 21, 8, 6, 7, 8, 7, 5, 7, 8, 7, 8, 0, 7, 7, 7, 16, 22, 6, 8, 5, 7, 8, 5, 8, 6, 7, 7), want = c(0, 7, 8, 5, 21, 8, 6, 0, 0, 1, 5, 7, 8, 7, 8, 0, 7, 7, 7, 16, 22, 6, 8, 5, 0, 0, 0, 8, 6, 7, 7)), class = c("tbl_df", "tbl", "data.frame"), row.names = c(NA, -31L))
尝试的错误代码
以下代码运行异常:既无法匹配正确变量名,还清零了过多观测值。
set_days_at_risk <- function(df) { for (person_id in unique(df$person)) { person_df <- df %>% filter(person == person_id) for (i in 1:(nrow(person_df) - 1)) { if (person_df$result[i] == "positive") { end_date <- person_df$date_of_collection[i] + days(21) person_df$days_at_risk[person_df$date_of_collection > person_df$date_of_collection[i] & person_df$date_of_collection <= end_date] <- 0 } } df[df$person == person_id, "days_at_risk"] <- person_df$days_at_risk } return(df) }
错误原因分析
- 变量名不匹配:代码中使用
person、date_of_collection、days_at_risk,但实际数据列名为identifier、date_collection、have;同时阳性结果的字符串是Positivo而非positive,大小写不匹配。 - 循环逻辑缺陷:嵌套循环效率低,且未考虑同一用户存在多个阳性结果的情况,容易出现覆盖错误。
正确解决方案
使用dplyr结合lubridate、purrr实现分组处理,逻辑清晰且效率更高:
library(dplyr) library(lubridate) library(purrr) # 处理数据 df_processed <- df %>% group_by(identifier) %>% mutate( # 收集当前用户所有阳性结果的采集日期 positive_dates = list(date_collection[result == "Positivo"]), # 判断当前行是否处于阳性结果后的21天风险期内,且当前行不是阳性结果 in_risk = map_lgl(date_collection, ~ any(.x > positive_dates[[1]] & .x <= positive_dates[[1]] + days(21))) & result != "Positivo", # 生成处理后的列,风险期内清零,否则保留原值 processed_have = ifelse(in_risk, 0, have) ) %>% ungroup() %>% # 移除中间辅助列 select(-positive_dates, -in_risk) # 验证结果是否与预期一致 all.equal(df_processed$processed_have, df$want)
代码逻辑说明
- 分组:按
identifier分组,确保每个用户的数据单独处理。 - 收集阳性日期:提取每个用户所有阳性结果的采集日期,存入列表列。
- 风险期判断:对每个采集日期,检查是否落在任意阳性日期的后21天范围内,同时排除阳性结果行本身。
- 值替换:根据风险期判断结果,将对应
have值清零,其余保留原值。
内容的提问来源于stack exchange,提问作者Levi M
相关产品推荐
相关产品推荐

