R语言实现:按工厂编号聚合产品ID并生成共享标识列
解决方案:为数据集添加共享工厂标识及关联产品列
需求说明
现有包含产品ID(id)和工厂编号(number)的数据集,需要新增两列:
shared product:二进制列,标记该产品是否存在相同工厂编号的其他产品what products:列示拥有相同工厂编号的所有产品ID(无关联产品则为空)
修正后的示例数据集
原示例代码中a,b,c等未定义,需改为字符型:
library(dplyr) library(stringr) df <- data.frame( id = c("a","b","c","d","e","f","g","a","b","d"), number = c("178","321","178","452","984","321","540","178","321","452") )
处理代码
通过dplyr和stringr包实现需求:
result_df <- df %>% # 去重:移除同一产品ID对应同一工厂的重复记录 distinct(id, number) %>% # 按工厂编号分组 group_by(number) %>% mutate( # 生成shared product列:组内不同产品数>1则为1,否则为0 `shared product` = as.integer(n_distinct(id) > 1), # 生成what products列:有共享工厂则拼接去重排序后的产品ID,否则为空 `what products` = ifelse(`shared product` == 1, str_c(sort(unique(id)), collapse = ","), "") ) %>% # 取消分组 ungroup() %>% # 按产品ID排序,匹配期望输出顺序 arrange(id) # 查看结果 print(result_df)
输出结果
# A tibble: 7 × 4 id number `shared product` `what products` <chr> <chr> <int> <chr> 1 a 178 1 a,c 2 b 321 1 b,f 3 c 178 1 a,c 4 d 452 0 "" 5 e 984 0 "" 6 f 321 1 b,f 7 g 540 0 ""
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

