R语言按日期与price_1筛选最接近price_2的行并处理平局
Solution Using dplyr
It looks like you're aiming to group your dataframe by Date and price_1, retain rows where price_2 is closest to price_1 (by absolute difference), and average price_2 and Volat when there's a tie in the minimum difference. Your initial code has the right idea but misses critical details—like grouping by both columns and handling ties with aggregation.
Step-by-Step Code Implementation
library(dplyr) # Assume your dataframe is named 'df' result <- df %>% # Group by both Date and price_1 (your initial code only grouped by Date) group_by(Date, price_1) %>% # Calculate absolute difference between price_2 and price_1, plus the minimum difference in the group mutate(abs_diff = abs(price_2 - price_1), min_group_diff = min(abs_diff)) %>% # Filter rows where the difference equals the group's minimum filter(abs_diff == min_group_diff) %>% # Aggregate: if 1 row, mean returns the original value; if multiple rows, returns the average summarise( price_2 = mean(price_2), Volat = mean(Volat), .groups = "drop" # Ungroup after summarising ) print(result)
How This Works
- Grouping:
group_by(Date, price_1)ensures we process each unique date-price_1 pair separately—this is essential since your requirement specifies grouping by both fields. - Calculate Differences:
mutateadds two columns:abs_diff: The absolute value ofprice_2 - price_1for each row.min_group_diff: The smallestabs_diffwithin the current group.
- Filter Closest Rows:
filter(abs_diff == min_group_diff)keeps only the rows whereprice_2is closest toprice_1in the group. - Handle Ties:
summarisewithmean()automatically handles both cases:- If only one row is left (no tie),
mean()returns the original value. - If multiple rows have the same minimum difference, it averages
price_2andVolatas requested.
- If only one row is left (no tie),
Testing with Your Sample Input
Running this code on your sample data will produce exactly your expected output:
# A tibble: 3 × 4 Date price_2 price_1 Volat <chr> <dbl> <dbl> <dbl> 1 2011-07-15 215 200 5 2 2011-07-15 200 201.5 6.2 3 2011-07-16 205 205 5.75
Fixes to Your Initial Code
Your original code had two main issues:
- It only grouped by
Date, notprice_1—this would mix different price_1 groups within the same date. - The
which()syntax doesn't fit with dplyr's pipeline workflow; usingmutate()+filter()is the standard way to handle group-wise logic in dplyr.
内容的提问来源于stack exchange,提问作者Narjems
相关产品推荐
相关产品推荐

