R语言:新增列标识同一clus_ID下animals_id是否唯一
按分组判断ID唯一性并新增标识列
原始数据集
df # A tibble: 21 × 2 animals_id clus_ID <chr> <int> 1 L085 246 2 L085 246 3 L085 246 4 L084 247 5 L084 247 6 L084 247 7 L085 249 8 L084 249 9 L084 249 10 L087 249
需求
新增一列type:同一clus_ID分组内仅含一种animals_id时标记为"A",含多种时标记为"B",预期结果如下:
animals_id clus_ID type 1 L085 246 A 2 L085 246 A 3 L085 246 A 4 L084 247 A 5 L084 247 A 6 L084 247 A 7 L085 249 B 8 L084 249 B 9 L084 249 B 10 L087 249 B
无效代码及问题分析
尝试的两段代码均未得到预期结果:
# 代码1:引用全局数据集而非分组内数据 df %>% group_by(clus_ID) %>% mutate(test = ifelse(length(unique(df[,"animals_id"]))==1, "A", "B")) # 代码2:逻辑正确但可能因数据/环境问题失效 df %>% group_by(clus_ID) %>% mutate(type = ifelse(n_distinct(animals_id) == 1, "A", "B"))
问题原因:
- 代码1错误:分组后仍调用整个数据集的
df[,"animals_id"],判断的是全局唯一值数量,而非当前分组内的数量,导致结果全部为"A"或"B"。 - 代码2异常:大概率是
animals_id存在隐形空格等格式问题,或者未正确加载dplyr包导致分组计算失效。
正确解决方案
方案1:修正分组内列引用
使用当前分组的animals_id列进行判断:
library(dplyr) df %>% group_by(clus_ID) %>% mutate(type = ifelse(length(unique(animals_id)) == 1, "A", "B")) %>% ungroup() # 可选:取消分组,避免后续操作受影响
方案2:用case_when优化逻辑并确保分组计算
如果偏好n_distinct,可结合case_when,同时先检查数据格式:
library(dplyr) # 先清洗数据,去除字符串两端隐形空格 df$animals_id <- trimws(df$animals_id) # 执行分组判断 df %>% group_by(clus_ID) %>% mutate(type = case_when(n_distinct(animals_id) == 1 ~ "A", TRUE ~ "B")) %>% ungroup()
可复现数据集
> dput(df) structure(list(animals_id = c("L085", "L085", "L085", "L084", "L084", "L084", "L085", "L084", "L084", "L087", "L084", "L084", "L084", "L084", "L084", "L084", "L084", "L084", "L084", "L084", "L084"), clus_ID = c(246L, 246L, 246L, 247L, 247L, 247L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L)), class = "data.frame", row.names = c(366428L, 366429L, 366430L, 349169L, 349170L, 349171L, 366435L, 349185L, 349186L, 378191L, 349343L, 349345L, 349346L, 349347L, 349477L, 349478L, 349479L, 349480L, 349706L, 349869L, 350121L))
内容的提问来源于stack exchange,提问作者mto23
相关产品推荐
相关产品推荐

