如何为data.table中重复的Variables_2022列添加匹配唯一ID列?
解决方案:为重复变量分配唯一编码
直接用data.table的分组赋值功能就能实现需求,以下是完整代码:
library(data.table) dat <- fread("Variable_codes_2022 Variables_2022 Cat1_1 This_question Cat1_2 Other_question Cat2_1 One_question Cat2_2 Another_question Cat3_1 Some_question Cat3_2 Extra_question Cat3_3 This_question Cat4_1 One_question Cat4_2 Wrong_question") # 生成Common_codes_2022列 dat[, Common_codes_2022 := if (.N >= 2) paste0("Com_", .GRP) else "", by = Variables_2022]
代码说明:
by = Variables_2022:按Variables_2022的内容分组,相同内容归为一组.N:表示当前组的行数(即该变量内容出现的次数).GRP:自动生成的分组序号,每个重复组对应唯一序号- 逻辑判断
if (.N >=2):仅为出现次数≥2的重复变量分配编码,不重复的变量对应列留空
运行后得到的结果如下:
Variable_codes_2022 Variables_2022 Common_codes_2022 1: Cat1_1 This_question Com_1 2: Cat1_2 Other_question 3: Cat2_1 One_question Com_2 4: Cat2_2 Another_question 5: Cat3_1 Some_question 6: Cat3_2 Extra_question 7: Cat3_3 This_question Com_1 8: Cat4_1 One_question Com_2 9: Cat4_2 Wrong_question
内容的提问来源于stack exchange,提问作者Tom
相关产品推荐
相关产品推荐

