You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在data.frame中仅保留仅出现一次的ID行(优先dplyr方案)

解决方案

dplyr 方法(推荐,适配大型数据集)

通过分组后过滤组大小为1的行,精准保留仅出现一次的Product对应记录:

library(dplyr)

data <- data.frame(Product=c('A', 'B', 'B', 'C'),
                   Likeability=c(80, 80, 82, 70),
                   Score=c(31, 33, 33, 33),
                   Quality=c(16, 32, 56, 18))

result <- data %>%
  group_by(Product) %>%
  filter(n() == 1) %>%
  ungroup() # 取消分组,避免影响后续数据操作

print(result)

执行后输出:

# A tibble: 2 × 4
  Product Likeability Score Quality
  <chr>         <dbl> <dbl>   <dbl>
1 A                80    31      16
2 C                70    33      18

dplyr的分组操作针对大数据集做了性能优化,内存占用低、处理速度快,适合大规模数据场景。

Base R 替代方法

无需加载第三方包时,可通过基础R工具实现:

data <- data.frame(Product=c('A', 'B', 'B', 'C'),
                   Likeability=c(80, 80, 82, 70),
                   Score=c(31, 33, 33, 33),
                   Quality=c(16, 32, 56, 18))

# 统计每个Product的出现频次
product_freq <- table(data$Product)
# 筛选仅出现一次的Product值
target_products <- names(product_freq[product_freq == 1])
# 提取对应行数据
result <- data[data$Product %in% target_products, ]

print(result)

执行后输出:

Product Likeability Score Quality
1       A          80    31      16
4       C          70    33      18

关键说明

  • 两种方法都会完全移除所有重复出现的Product对应的行,不会保留重复组中的任意一条记录,完全匹配需求。
  • 超大型数据集场景下,优先选择dplyr方案,其底层优化能显著提升处理效率。

内容的提问来源于stack exchange,提问作者econ9595

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 14:35:20