基于PipeYear筛选数据框:取PropertyYearBuilt的最近下方与最远上方值
解决方案
我们可以通过分组处理实现需求,这里使用dplyr包完成数据筛选操作:
核心逻辑
按PropertyYearBuilt分组后,针对每组执行两个筛选动作:
- 从
PipeYear小于PropertyYearBuilt的记录里,保留PipeYear最大的那条(即最接近目标年份的记录) - 从
PipeYear大于PropertyYearBuilt的记录里,保留PipeYear最大的那条(即最远大于目标年份的记录)
最后将两组结果合并得到最终数据框。
完整代码
# 加载dplyr包 library(dplyr) df <- read.table(text=" PipeID PricePipe PipeYear PropertyYearBuilt Distance_to_property a 500 2010 2013 1.5 b 600 2007 2008 2.5 c 700 2009 2008 3.0 d 800 1998 2000 4.2 e 900 2003 2000 4.5 f 200 2014 2013 5.0 g 100 2011 2013 5.5 h 850 2018 2008 7.0", header = TRUE) # 执行筛选 filtered_df <- df %>% group_by(PropertyYearBuilt) %>% group_modify(~ { # 筛选小于目标年份且最接近的记录 lower_record <- .x %>% filter(PipeYear < .y$PropertyYearBuilt) %>% slice_max(PipeYear) # 筛选大于目标年份且最远的记录 upper_record <- .x %>% filter(PipeYear > .y$PropertyYearBuilt) %>% slice_max(PipeYear) # 合并结果 bind_rows(lower_record, upper_record) }) %>% ungroup() # 查看最终结果 print(filtered_df)
输出结果
# A tibble: 6 × 5 PipeID PricePipe PipeYear PropertyYearBuilt Distance_to_property <chr> <int> <int> <int> <dbl> 1 g 100 2011 2013 5.5 2 f 200 2014 2013 5.0 3 b 600 2007 2008 2.5 4 h 850 2018 2008 7.0 5 d 800 1998 2000 4.2 6 e 900 2003 2000 4.5
内容的提问来源于stack exchange,提问作者Xaviermoros
相关产品推荐
相关产品推荐

