求助:基于personID和operationID在R DataFrame中生成计数列
解决R中按person_id分组,对operation_id按出现顺序生成统一计数的问题
你需要在person_id分组内,给每个首次出现的operation_id按顺序分配递增数值,相同的operation_id对应同一个数。用dplyr包就能轻松实现,以下是两种靠谱的方法:
方法1:用match()+unique()精准匹配出现顺序
这种方法直接基于分组内operation_id的实际出现顺序生成计数,适配所有场景:
library(dplyr) # 加载你的示例数据 person_id <- c("1", "1", "1", "2", "2", "2", "2", "3", "3") operation_id <- c("60533", "60533", "60534", "50677", "50678", "50678", "50679", "78322", "78322") row <- c("1", "2", "3", "4", "5", "6", "7", "8", "9") df <- data.frame(person_id, operation_id, row) # 生成目标列row_intend df <- df %>% group_by(person_id) %>% mutate(row_intend = match(operation_id, unique(operation_id))) %>% ungroup() # 查看结果 print(df)
方法2:用dense_rank()快速生成连续排名
如果你的operation_id出现顺序和自身数值顺序一致(比如示例数据),用dense_rank()更简洁:
df <- df %>% group_by(person_id) %>% mutate(row_intend = dense_rank(operation_id)) %>% ungroup()
结果验证
运行任意一种方法后,输出都会和你期望的完全一致:
# A tibble: 9 × 4 person_id operation_id row row_intend <chr> <chr> <chr> <int> 1 1 60533 1 1 2 1 60533 2 1 3 1 60534 3 2 4 2 50677 4 1 5 2 50678 5 2 6 2 50678 6 2 7 2 50679 7 3 8 3 78322 8 1 9 3 78322 9 1
关键说明
- 两种方法都必须先按
person_id分组,保证计数只在同一个人的范围内进行; - 方法1的核心逻辑是:
unique(operation_id)会保留分组内operation_id的出现顺序,match()返回每个值在这个唯一列表中的位置,从而得到顺序计数,不管operation_id的数值大小如何都能生效; - 方法2依赖
operation_id的数值顺序和出现顺序一致,如果你不确定数据是否满足这个条件,优先选方法1。
内容的提问来源于stack exchange,提问作者r_newbie
相关产品推荐
相关产品推荐

