如何筛选与首次日期相差365±90天的实验室检测记录?
Hey, great job using a recursive approach here—it’s such a logical fit for the sequential date selection rules you need! Let’s break down some tweaks to your current code and also share an iterative alternative that might be more robust for larger datasets.
First, Quick Notes on Your Existing Recursive Function
Your core logic is spot-on: starting from the first date, then repeatedly finding the next date that’s within 365±90 days and closest to exactly 365 days. That said, there are a few small areas to strengthen:
- Dependency on
first(): This function is from dplyr, so if someone runs the code without loading dplyr first, it’ll throw an error. Swapping it for basic subsetting ([1]) makes it more self-contained. - Unstated sorting assumption: The function works only if the dates are already sorted in ascending order. If your raw data has out-of-order dates, the results will be wrong. We should explicitly sort dates either in the function or before grouping.
- Recursion depth limits: R has a default recursion depth cap (usually 1000). While this is unlikely to hit with clinical lab data, it’s a potential edge case to keep in mind.
Improved Recursive Function
Here’s a revised version that fixes these issues:
library(dplyr) f_improved <- function(d) { # First, ensure dates are sorted ascending d_sorted <- sort(d) # Internal recursive helper recurse <- function(current_idx) { datediff <- difftime(d_sorted, d_sorted[current_idx], units = "days") # Only consider dates AFTER the current one, within 365±90 days in_range <- datediff > 0 & datediff >= 275 & datediff <= 455 if (!any(in_range)) { return(current_idx) } # Find all valid candidates, then pick the one closest to 365 days candidates <- which(in_range) closest_idx <- candidates[which.min(abs(datediff[candidates] - 365))] return(c(current_idx, recurse(closest_idx))) } # Start recursion from the first sorted date recurse(1) } # Usage with explicit sorting first df %>% arrange(PATIENT_ID, LAB_TEST_DATE) %>% group_by(PATIENT_ID) %>% slice(f_improved(LAB_TEST_DATE)) %>% ungroup()
Alternative: Iterative Implementation
If you want to avoid recursion entirely (to sidestep any depth limits), an iterative approach works just as well with the same logic:
f_iterative <- function(d) { d_sorted <- sort(d) selected_indices <- c(1) current_idx <- 1 total_dates <- length(d_sorted) while(TRUE) { datediff <- difftime(d_sorted, d_sorted[current_idx], units = "days") # Filter for valid next dates in_range <- datediff > 0 & datediff >= 275 & datediff <= 455 if (!any(in_range)) { break } candidates <- which(in_range) closest_idx <- candidates[which.min(abs(datediff[candidates] - 365))] selected_indices <- c(selected_indices, closest_idx) current_idx <- closest_idx # Safety break to prevent infinite loops (edge case) if (current_idx >= total_dates) { break } } return(selected_indices) } # Usage df %>% arrange(PATIENT_ID, LAB_TEST_DATE) %>% group_by(PATIENT_ID) %>% slice(f_iterative(LAB_TEST_DATE)) %>% ungroup()
Testing the Results
Both functions will produce the expected output for your sample data:
- Patient 1 only keeps the first test date
- Patient 2 retains 2015-05-30 (first), 2016-08-01, and 2017-07-27
- All other patients follow your rules correctly
Just remember: sorting the data by PATIENT_ID and LAB_TEST_DATE upfront is critical to ensure the functions work as intended, even if your raw data is messy.
内容的提问来源于stack exchange,提问作者HNSKD

