You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何根据时间条件将dataframe的Disease_Name转为空白/NA

处理R数据框的日期条件筛选问题

原始数据框

> dput(df)
structure(list(Person_ID = c(123L, 123L), Disease_Name = c("Heart Disease", 
"Lung Disease"), Disease_start = c("4/11/17", "4/11/17"), Procedure_start = c("4/11/18", 
"4/11/16")), class = "data.frame", row.names = c(NA, -2L))

处理规则

  • 当Disease_start早于Procedure_start时,将Disease_Name设为空白/NA
  • 当Disease_start晚于Procedure_start时,保留Disease_Name不变

期望输出

> dput(df2)
structure(list(Person_ID = c(123L, 123L), Disease_Name = c("", 
"Lung Disease"), Disease_start = c("4/11/17", "4/11/17"), Procedure_start = c("4/11/18", 
"4/11/16")), class = "data.frame", row.names = c(NA, -2L))

解决方案

注意:原始数据中的日期是字符格式,直接比较会有风险(比如不同月份的日期字符排序可能不符合实际日期顺序),所以第一步要先把日期列转为标准日期类型,完成判断后再转回原格式(如果需要和期望输出一致的话)。

方法1:基础R实现

# 转换日期列为Date类型
df$Disease_start_date <- as.Date(df$Disease_start, format = "%m/%d/%y")
df$Procedure_start_date <- as.Date(df$Procedure_start, format = "%m/%d/%y")

# 根据条件修改Disease_Name
df$Disease_Name <- ifelse(df$Disease_start_date < df$Procedure_start_date, "", df$Disease_Name)

# 移除临时日期列(可选)
df <- df[, !names(df) %in% c("Disease_start_date", "Procedure_start_date")]

# 查看结果
dput(df)

方法2:tidyverse(dplyr)实现

如果习惯用tidyverse语法,可以这样写:

library(dplyr)

df2 <- df %>%
  # 转换日期格式
  mutate(
    Disease_start_date = as.Date(Disease_start, format = "%m/%d/%y"),
    Procedure_start_date = as.Date(Procedure_start, format = "%m/%d/%y")
  ) %>%
  # 根据条件更新Disease_Name
  mutate(Disease_Name = case_when(
    Disease_start_date < Procedure_start_date ~ "",
    TRUE ~ Disease_Name
  )) %>%
  # 移除临时日期列
  select(-c(Disease_start_date, Procedure_start_date))

# 查看结果
dput(df2)

两种方法都能得到符合要求的输出,其中format = "%m/%d/%y"对应数据中的日期格式(月/日/两位年份),若你的日期格式不同,需要调整该参数。

内容的提问来源于stack exchange,提问作者Jamie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 09:23:37