如何用dplyr匹配数据框时间重叠行并计算对应中位数?
Hey there! Let's walk through how to solve this problem using dplyr—since you mentioned the between() function, we'll use that alongside other intuitive dplyr verbs to get the job done easily, even as an R beginner.
Step 1: Setup Required Packages & Data
First, make sure you have dplyr installed and loaded. If you haven't installed it yet, run this first:
install.packages("dplyr") library(dplyr)
Next, let's replicate your original data frames so you can test the code directly:
# Your original data frames df1 <- tibble(time1 = c(0, 1, 2, 3, 4, 5, 6, 7, 8, 9), time2 = c(1, 2, 3, 4, 5, 6, 7, 8, 9, 10), id = c("a", "b", "c", "d", "e", "f", "g", "h", "i", "j")) df2 <- tibble(time = sort(runif(100, 0, 10)), C = rbinom(100, 1, 0.5))
Step 2: Core Calculation with rowwise() & between()
The simplest way for beginners to handle row-by-row operations in dplyr is using rowwise(). Here's the code to add your new median column:
# Add C_median column to df1 df1_with_median <- df1 %>% rowwise() %>% # Tell dplyr to process each row individually mutate( C_median = median( df2$C[between(df2$time, time1, time2)], # Filter df2$C where time is in [time1, time2] na.rm = TRUE # Handle cases where no rows fall in the interval (prevents NA errors) ) ) %>% ungroup() # Reset to normal data frame grouping after processing
How This Works
Let's break down the key parts:
rowwise(): This switches dplyr from operating on the entire data frame to handling each row one at a time.between(df2$time, time1, time2): Creates a boolean vector where each entry isTRUEifdf2$timefalls between the current row'stime1andtime2(inclusive).df2$C[...]: Extracts only theCvalues fromdf2` that match the interval condition.median(..., na.rm = TRUE): Calculates the median of those filteredCvalues. Thena.rm = TRUEensures we don't get an error if an interval has no matching rows (though with 100 rows indf2, this is unlikely here).ungroup(): Always a good idea to ungroup afterrowwise()to avoid unexpected behavior in future operations.
Alternative: Using purrr for Function-Style Processing
If you're curious about another approach, you can use purrr (part of the tidyverse) to iterate over rows. First load purrr:
install.packages("purrr") library(purrr)
Then run this code:
df1_with_median <- df1 %>% mutate( C_median = map_dbl( 1:nrow(.), # Iterate over each row index of df1 ~ median(df2$C[between(df2$time, time1[.x], time2[.x])], na.rm = TRUE) ) )
This achieves the same result but uses functional programming instead of row-wise grouping—both are valid, so pick whichever feels more intuitive to you!
Check the Result
To verify everything worked, print the first few rows of your updated data frame:
head(df1_with_median)
You should see your original df1 columns plus a new C_median column with the median value for each interval.
内容的提问来源于stack exchange,提问作者mowglis_diaper

