You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:新增列标识同一clus_ID下animals_id是否唯一

按分组判断ID唯一性并新增标识列

原始数据集

df
# A tibble: 21 × 2
   animals_id clus_ID
   <chr>        <int>
 1 L085           246
 2 L085           246
 3 L085           246
 4 L084           247
 5 L084           247
 6 L084           247
 7 L085           249
 8 L084           249
 9 L084           249
10 L087           249

需求

新增一列type:同一clus_ID分组内仅含一种animals_id时标记为"A",含多种时标记为"B",预期结果如下:

animals_id clus_ID   type
 1 L085           246   A
 2 L085           246   A
 3 L085           246   A
 4 L084           247   A
 5 L084           247   A
 6 L084           247   A
 7 L085           249   B
 8 L084           249   B
 9 L084           249   B
10 L087           249   B

无效代码及问题分析

尝试的两段代码均未得到预期结果:

# 代码1:引用全局数据集而非分组内数据
df %>% group_by(clus_ID) %>% mutate(test = ifelse(length(unique(df[,"animals_id"]))==1, "A", "B"))

# 代码2:逻辑正确但可能因数据/环境问题失效
df %>% group_by(clus_ID) %>% mutate(type = ifelse(n_distinct(animals_id) == 1, "A", "B"))

问题原因:

  • 代码1错误:分组后仍调用整个数据集的df[,"animals_id"],判断的是全局唯一值数量,而非当前分组内的数量,导致结果全部为"A"或"B"。
  • 代码2异常:大概率是animals_id存在隐形空格等格式问题,或者未正确加载dplyr包导致分组计算失效。

正确解决方案

方案1:修正分组内列引用

使用当前分组的animals_id列进行判断:

library(dplyr)

df %>% 
  group_by(clus_ID) %>% 
  mutate(type = ifelse(length(unique(animals_id)) == 1, "A", "B")) %>% 
  ungroup() # 可选:取消分组,避免后续操作受影响

方案2:用case_when优化逻辑并确保分组计算

如果偏好n_distinct,可结合case_when,同时先检查数据格式:

library(dplyr)

# 先清洗数据,去除字符串两端隐形空格
df$animals_id <- trimws(df$animals_id)

# 执行分组判断
df %>% 
  group_by(clus_ID) %>% 
  mutate(type = case_when(n_distinct(animals_id) == 1 ~ "A",
                          TRUE ~ "B")) %>% 
  ungroup()

可复现数据集

> dput(df)
structure(list(animals_id = c("L085", "L085", "L085", "L084", 
"L084", "L084", "L085", "L084", "L084", "L087", "L084", "L084", 
"L084", "L084", "L084", "L084", "L084", "L084", "L084", "L084", 
"L084"), clus_ID = c(246L, 246L, 246L, 247L, 247L, 247L, 249L, 
249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 249L, 
249L, 249L, 249L)), class = "data.frame", row.names = c(366428L, 
366429L, 366430L, 349169L, 349170L, 349171L, 366435L, 349185L, 
349186L, 378191L, 349343L, 349345L, 349346L, 349347L, 349477L, 
349478L, 349479L, 349480L, 349706L, 349869L, 350121L))

内容的提问来源于stack exchange,提问作者mto23

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 04:50:59