如何按国家分组为人员生成字母序列标识?
解决方案
方法一:使用dplyr(tidyverse生态)
首先构造可复现的示例数据:
library(dplyr) df <- tibble( country = c(rep("A",6), rep("B",6)), person = c(rep("John",3), rep("Peter",3), "David", "Thomas", "David", "Adam", "Adam", "Thomas"), time = c(1:3, 1:3, 1,2,3,1,2,3) )
按国家分组,为每个唯一人员生成字母序列标识:
df <- df %>% group_by(country) %>% mutate( # 先匹配人员在组内唯一列表的位置,再转成对应大写字母 Letterseq = LETTERS[match(person, unique(person))] ) %>% ungroup()
输出结果:
print(df)
# A tibble: 12 × 4 country person time Letterseq <chr> <chr> <int> <chr> 1 A John 1 A 2 A John 2 A 3 A John 3 A 4 A Peter 1 B 5 A Peter 2 B 6 A Peter 3 B 7 B David 1 A 8 B Thomas 2 B 9 B David 3 A 10 B Adam 1 C 11 B Adam 2 C 12 B Thomas 3 B
方法二:使用data.table
如果习惯用data.table处理数据,可采用以下写法:
library(data.table) dt <- as.data.table(df) dt[, Letterseq := LETTERS[match(person, unique(person))], by = country]
核心逻辑说明
- 按
country分组,确保每个国家内的人员标识独立生成 unique(person)提取组内所有唯一人员,match()将每个人员映射到其在唯一列表中的位置(数字序号)- 利用R内置的
LETTERS向量,把数字序号转换为对应的大写字母
内容的提问来源于stack exchange,提问作者Victor Shin
相关产品推荐
相关产品推荐

