如何高效合并R数据框中值相近的行并计算均值?
Great question! Using vectorized operations instead of a for loop will be way more efficient, especially for larger datasets. Here's a clean, optimized solution using dplyr:
library(dplyr) DF <- data.frame(ID=letters[1:12], Values=c(1, 2.2, 3, 5, 6.2, 6.8, 7, 8.5, 10, 12.2, 13, 14)) result <- DF %>% # Create a group identifier: start new group when adjacent Values differ by >=1 mutate(group = cumsum(c(TRUE, diff(Values) >= 1))) %>% # Group by our new group column group_by(group) %>% # Summarize: concatenate IDs and calculate mean Values (rounded to 2 decimals) summarize( ID = paste(ID, collapse = "/"), Values = round(mean(Values), 2) ) %>% # Remove the group column since we don't need it in the final output select(-group) print(result)
Output:
# A tibble: 8 × 2 ID Values <chr> <dbl> 1 a 1.00 2 b/c 2.60 3 d 5.00 4 e/f/g 6.67 5 h 8.50 6 i 10.00 7 j/k 12.60 8 l 14.00
How this works:
- Group Identification: The
cumsum(c(TRUE, diff(Values) >=1))creates a unique group number for each set of consecutive rows where adjacentValuesdiffer by less than 1. TheTRUEensures the first row starts a new group, and every time the difference between the current and previous value is ≥1, we increment the group number. - Summarization: Using
group_by(group)lets us aggregate each group:paste(ID, collapse="/")concatenates all IDs in the group with slashes.round(mean(Values),2)calculates the average of theValuesin the group and rounds it to 2 decimal places (matching your example).
- Efficiency: This approach uses vectorized operations (no loops) which are optimized in R, making it much faster than a for loop—especially as your dataframe grows in size.
Note: I fixed a small typo in your expected output (the last row's ID should be l instead of i, since your original dataframe uses letters[1:12]).
内容的提问来源于stack exchange,提问作者Jian
相关产品推荐
相关产品推荐

