You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中提取无重复分类数据的技术问询

提取分类数据并去除重复项的正确方法

你需要从包含数值型和分类型数据的数据集中提取分类数据并去除重复项,以下是针对需求的正确实现方案:

示例数据集

首先修正代码中的拼写错误(aninimal改为animal),示例数据集定义如下:

df <- data.frame(
  animal = c('cat', 'cat', 'cat', 'cat', 'dog', 'dog', 'dog', 'dog', 'dog'),
  fur_col = c('tan', 'tan', 'tan', 'white', 'black', 'black', 'white', 'brown', 'brown'),
  age = c(2, 2, 3, 5, 7, 3, 1, 6, 5)
)

原代码问题分析

你使用的代码无法实现去重:group_by()仅完成分组操作,而summarize(animal, fur_col)会返回分组内的所有行,并未执行去重逻辑,因此得到重复结果:

summary <- df %>% 
  group_by(animal, fur_col) %>% 
  summarize(animal, fur_col)

正确实现方法

方法1:使用dplyr的distinct()函数

distinct()是dplyr中专门用于去除重复行的函数,可直接提取唯一的分类列组合:

# 选择分类列并去重
unique_categories <- df %>% 
  select(animal, fur_col) %>% 
  distinct()

或者更简洁的写法,直接在distinct()中指定目标列:

unique_categories <- df %>% 
  distinct(animal, fur_col, .keep_all = FALSE)

方法2:基础R的unique()函数

如果不依赖dplyr包,使用基础R的unique()函数也能实现需求:

# 提取分类列后去重
unique_categories <- unique(df[, c("animal", "fur_col")])

最终结果

两种方法都会得到你期望的无重复分类数据:

animalfur_col
cattan
catwhite
dogblack
dogwhite
dogbrown

内容的提问来源于stack exchange,提问作者Shae11

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 23:48:22