R语言创建分组唯一ID并按组内条件重编码ATM位置数据
R语言实现ATM点位数据分组与字段校正
核心处理步骤
- 分组ID生成:以
address、terminal_id两个字段作为联合分组键,为每个唯一分组分配连续整数ID - 正确值提取:每个分组内,筛选2019年1月1日之前的记录,提取非
"Financial Institution"的取值作为该组点位类型的校正基准;如果组内2019年前无其他取值,则判定原始编码正确,不做修改 - 字段生成:新增
location_corrected存储校正后的点位类型,新增location_changed标记该分组是否存在编码修正,修正过的组统一标记为"yes",未修正标记为"no"
可运行代码
library(dplyr) library(tibble) # 读入原始数据 data <- tribble( ~address, ~date, ~terminal_id, ~location_type_description, "1 GATEWAY DR OROMOCTO", "2017-01-01", "NC79", "Gas Station", "1 GATEWAY DR OROMOCTO", "2018-01-01", "NC79", "Gas Station", "1 GATEWAY DR OROMOCTO", "2019-11-01", "NC79", "Financial Institution", "1 GATEWAY DR OROMOCTO", "2020-01-01", "NC79", "Financial Institution", "1 GATEWAY DR OROMOCTO", "2020-12-01", "NC79", "Financial Institution" ) %>% mutate(across(date, as.Date)) # 执行清洗逻辑 data_clean <- data %>% group_by(address, terminal_id) %>% # 生成唯一分组ID mutate(group_identifier = cur_group_id()) %>% mutate( # 提取组内2019年前的正确类型,无符合条件值时保留原始类型 ref_type = first( location_type_description[date < as.Date("2019-01-01") & location_type_description != "Financial Institution"], default = first(location_type_description) ), location_corrected = ref_type, # 判断是否存在修正:只要2019年后有记录和参考值不一致,就标记为已修改 location_changed = ifelse( any(location_type_description[date >= as.Date("2019-01-01")] != ref_type), "yes", "no" ) ) %>% ungroup() %>% select(-ref_type)
注意事项
若单个分组内2019年之前存在多种非
"Financial Institution"的点位类型,上述代码默认取最早记录对应的类型作为校正基准,可根据实际业务规则调整first()函数的取值逻辑,比如取出现频次最高的类型。
内容的提问来源于stack exchange,提问作者dano_
相关产品推荐
相关产品推荐

