使用tidyverse基于df2分组时间范围提取df1的分组索引值
Solution: Extract Matching Indices Based on Time Range and Group
Got it, let's work through this problem together. The goal is to pull the index values from df1 where, for each row in df2, the time falls within the start-end range of the matching group, then package those results into df3.
First, let's confirm our input data (I've re-written it for clarity):
# Define df1: Grouped time ranges with indices df1 <- data.frame( group = c("A","A","A","A","B","B","B","B","C","C","C","C"), index = c(1,2,3,4,5,6,7,8,9,10,11,12), start = c(5,10,15,20,5,10,15,20,5,10,15,20), end = c(10,15,20,25,10,15,20,25,10,15,20,25) ) # Define df2: Grouped time points to match df2 <- data.frame( group = c("A","B","B","C","A","C"), time = c(11,17,24,5,5,22) )
Method 1: Using dplyr (Tidyverse Approach)
This is the most readable and efficient approach for most users, leveraging joins and filtering to align the data:
library(dplyr) df3 <- df2 %>% # Join df2 with df1 on the 'group' column to pair all matching group rows left_join(df1, by = "group") %>% # Keep only rows where df2's time falls within df1's start-end range filter(time >= start, time <= end) %>% # Keep only the columns we need in the final output select(group, time, index)
What this does:
left_joinpairs every row in df2 with all rows in df1 that share the samegroupfilternarrows down to only the pairs where the time from df2 is inside the start-end window from df1selectcleans up the output to just the columns we care about
Method 2: Using Base R (No External Packages)
If you prefer to stick with base R, this loop-based approach works just as well:
# Initialize an empty list to store our matches match_list <- list() # Loop through each row in df2 for (i in seq(nrow(df2))) { current_group <- df2$group[i] current_time <- df2$time[i] # Find rows in df1 that match the group AND contain the time in their range matched_indices <- df1[ df1$group == current_group & df1$start <= current_time & df1$end >= current_time, "index" ] # Add the results to our list, keeping track of the original group and time match_list[[i]] <- data.frame( group = current_group, time = current_time, index = matched_indices ) } # Combine all list elements into a single data frame df3 <- do.call(rbind, match_list)
Expected Output
Both methods will produce the same df3:
print(df3) # group time index # 1 A 11 2 # 2 B 17 7 # 3 B 24 8 # 4 C 5 9 # 5 A 5 1 # 6 C 22 12
内容的提问来源于stack exchange,提问作者Al Mac
相关产品推荐
相关产品推荐

