R语言按州县分组创建流浪狗数量峰值日二分类变量的实现方法
解决方法
首先提前提取全局流浪狗峰值日期,再按州、县分组逐组处理逻辑即可,完整实现代码如下:
library(tidyverse) library(lubridate) # 原始数据生成代码 state <- c(rep("Alabama", 10), rep("Arizona", 10), rep("Arkansas", 10)) county <- c(rep("Baldwin", 5), rep("Barbour", 5), rep("Apache", 5), rep("Cochise", 5), rep("Arkansas", 5), rep("Ashley", 5)) date <- rep(seq(ymd('2012-04-06'),ymd('2012-04-10'),by='days'), 6) stray_dogs <- c(lag(1:3, n = 2, default = 0), floor(runif(7, min=1, max=4)), lag(1:6, n = 5, default = 0), floor(runif(4, min=1, max=18)), lag(1:2, n = 1, default = 0), floor(runif(8, min=1, max=4))) df <- data.frame(state, county, date, stray_dogs) %>% mutate(stray_dogs_max = max(stray_dogs)) %>% mutate(most_stray_dogs = case_when(stray_dogs_max == stray_dogs ~ 1, stray_dogs_max != stray_dogs ~ 0)) # 第一步:提取全局峰值日期 global_peak_date <- df %>% filter(most_stray_dogs == 1) %>% pull(date) %>% unique() # 第二步:分组计算目标二分类列county_peak df_result <- df %>% group_by(state, county) %>% mutate( # 生成过程辅助变量 county_max = max(stray_dogs), county_total = sum(stray_dogs), date_diff = abs(as.numeric(difftime(date, global_peak_date, units = "days"))), # 按规则生成目标列 county_peak = case_when( # 规则1:该县无流浪狗记录,全局峰值日标1 county_total == 0 ~ as.integer(date == global_peak_date), # 规则2:匹配县内最大值后,选离全局峰值最近的日期标1 stray_dogs == county_max ~ as.integer(row_number(date_diff) == 1), # 其余情况标0 TRUE ~ 0L ) ) %>% # 可选:删除过程辅助列 ungroup() %>% select(-county_max, -county_total, -date_diff)
逻辑说明
- 提前提取全局峰值日期避免分组内重复计算,提升运行效率
- 分组后先判断该县是否所有流浪狗记录为0,命中规则1直接赋值
- 对于存在流浪狗记录的县,先筛选出等于该县最大值的行,再按与全局峰值日的时间差升序排列,第一个出现的即为符合要求的峰值日,标记为1,其余为0
- 若存在多个日期和全局峰值日的时间差完全相同的极端情况,
row_number()会默认选择日期更早的标记为1,也可以替换为slice_min(date_diff, n=1, with_ties = FALSE)实现相同效果
内容的提问来源于stack exchange,提问作者Ludrew
相关产品推荐
相关产品推荐

