如何在含多ID重复观测的数据框中根据switch变量获取个体第三次换药日期
Hey there! Let's work through this problem together. Your goal is to track cumulative medication switches per individual (only counting rows where switch = 1) and then extract the date of the third switch for each ID. Here's how to do it cleanly with dplyr:
Step 1: Prepare the Data & Calculate Cumulative Switches
First, we need to sort each individual's prescriptions by date (since the order of visits matters for tracking switches), then compute the cumulative count of switches only for rows where a switch actually happened (leaving NA for non-switch rows):
library(dplyr) # Re-create your dataset for consistency set.seed(42) ID <- sample(c("ID1", "ID2", "ID3", "ID4", "ID5", "ID6", "ID7", "ID8", "ID9", "ID10"), 40, replace = T) prescription_date <- sample(seq(as.Date('1999/01/01'), as.Date('2010/01/01'), by="month"), 40) switch <- sample(c(0, 1), 40, replace = T, prob = c(0.4, 0.6)) df <- data.frame(ID, prescription_date, switch) # Add the desired date3switch column df_processed <- df %>% group_by(ID) %>% arrange(prescription_date, .by_group = TRUE) %>% # Sort prescriptions by date per individual mutate( date3switch = case_when( switch == 1 ~ cumsum(switch), # Cumulative count only when a switch occurs TRUE ~ NA_real_ # NA for rows with no switch ) ) # Check the processed data (matches your example output) df_processed %>% print(n=24)
This code produces exactly the date3switch column you described: cumulative numbers for switch rows, and NA for non-switch rows.
Step 2: Extract the Third Switch Date
Now we can pull out the date of the third switch for each individual (or NA if they never had three switches):
third_switch_dates <- df_processed %>% summarise( third_switch_date = ifelse( any(date3switch == 3), # Check if the individual has at least 3 switches prescription_date[date3switch == 3], # Grab the date of the third switch NA_Date_ # Assign NA if no third switch exists ) ) # View the final result third_switch_dates
Why Your For Loop Didn't Work
Your initial loop had a few key issues:
- For loops in R don't return values unless explicitly specified, so assigning the loop directly to
dfwon't work. sum(data$switch)calculates total switches across the entire dataset, not per individual.- You didn't sort prescriptions by date, which is critical for tracking the correct order of switches.
The dplyr approach handles grouping, sorting, and cumulative counting in a clean, readable chain—no messy loops needed!
内容的提问来源于stack exchange,提问作者Margherita Bigossi

