如何在R中依据条件数据框修改目标数据列的值?
dplyr Your loop-based solution works, but a more idiomatic and efficient way to handle this in the tidyverse is using a non-equi join to match each timestamp in data_df to the corresponding range in criteria_df, then updating the values accordingly. This avoids explicit loops and leverages vectorized operations, which are faster for larger datasets.
Here's how to implement it:
library(tidyverse) # Original data set.seed(123) data_df <- tibble(t = 1:15, value = sample(letters, 15)) # Criteria data criteria_df <- tibble(start = c(1, 3, 7), end = c(2, 5, 10), value = c('a', 'b', 'c')) # Update values using non-equi join updated_data <- data_df %>% # Join data_df with criteria_df where t falls between start and end left_join( criteria_df, join_by(between(t, start, end)), suffix = c("_original", "_new") ) %>% # Use the new value if available, else keep original mutate(value = coalesce(value_new, value_original)) %>% # Keep only the necessary columns select(t, value) # View the result updated_data
How This Works:
- Non-Equi Join: The
join_by(between(t, start, end))condition matches eachtindata_dfto the row(s) incriteria_dfwheretlies within thestart-endrange. For timestamps not in any range, thevalue_newcolumn will beNA. - Coalesce:
coalesce(value_new, value_original)picks the first non-NAvalue, which means we use the criteria-specified value when available, otherwise retain the original value. - Cleanup: We select only the
tand updatedvaluecolumns to get back the original dataframe structure.
Handling Overlapping Ranges (If Applicable):
If your criteria_df has overlapping ranges, the join will create multiple rows for timestamps that fall into multiple ranges. In that case, you can add a step to prioritize which value to use (e.g., the latest range, or the first match). For example, if you want to keep the last matching criteria:
updated_data <- data_df %>% left_join(criteria_df, join_by(between(t, start, end)), suffix = c("_original", "_new")) %>% group_by(t) %>% slice_last() %>% # Keep the last matching criteria row ungroup() %>% mutate(value = coalesce(value_new, value_original)) %>% select(t, value)
Alternative: Functional Approach with purrr
If you prefer a functional style over joins, you can use purrr::reduce to iteratively apply each criteria range to the dataframe. This is similar to your loop but more concise:
library(purrr) updated_data <- criteria_df %>% pmap(function(start, end, value) { data_df %>% mutate(value = if_else(between(t, start, end), value, value)) }) %>% reduce(full_join, by = c("t", "value")) %>% distinct(t, .keep_all = TRUE)
However, this approach is less efficient than the non-equi join for large datasets, as it involves multiple dataframe operations.
Content of the question originates from Stack Exchange, asked by mauna

