如何使用tidyverse实现基于时间范围从DataFrame提取对应索引?
Absolutely! There’s a straightforward tidyverse-based approach to map each time in df2 to the corresponding index from df1 where the time falls within the start-end range. Here are two reliable methods, depending on your dataset size:
Method 1: Fuzzy Join (Best for Larger Datasets)
Using fuzzyjoin (a tidyverse-compatible package) is efficient and clean, especially if you’re working with bigger datasets. It avoids creating a full cross-product of rows, which saves memory.
First, load the required packages:
library(tidyverse) library(fuzzyjoin)
Then run the code to generate df3:
# Your original data df1 <- data.frame(index = c(1,2,3,4), start = c(5,10,15,20), end = c(10,15,20,25)) df2 <- data.frame(time = c(11,17,18,5,5,22)) # Match times to their corresponding ranges df3 <- df2 %>% fuzzy_left_join( df1, by = c("time" = "start", "time" = "end"), match_fun = list(`>=`, `<=`) # Check time >= start AND time <= end ) %>% select(time, index) %>% # Keep only the columns we need arrange(match(time, df2$time)) # Preserve original order of df2's times
This will produce exactly the df3 you specified:
> df3 time index 1 11 2 2 17 3 3 18 3 4 5 1 5 5 1 6 22 4
Method 2: Rowwise Processing (Simple for Small Datasets)
If you’re working with a small dataset and prefer to stick strictly to core dplyr functions, you can use row-wise processing to check each time against df1’s ranges:
df3 <- df2 %>% rowwise() %>% mutate(index = df1$index[start <= time & end >= time]) %>% ungroup()
This works by evaluating each row in df2 individually, finding the matching index from df1, and then ungrouping to return a standard tibble/data frame.
内容的提问来源于stack exchange,提问作者Al Mac

