如何在data.frame中仅保留仅出现一次的ID行(优先dplyr方案)
解决方案
dplyr 方法(推荐,适配大型数据集)
通过分组后过滤组大小为1的行,精准保留仅出现一次的Product对应记录:
library(dplyr) data <- data.frame(Product=c('A', 'B', 'B', 'C'), Likeability=c(80, 80, 82, 70), Score=c(31, 33, 33, 33), Quality=c(16, 32, 56, 18)) result <- data %>% group_by(Product) %>% filter(n() == 1) %>% ungroup() # 取消分组,避免影响后续数据操作 print(result)
执行后输出:
# A tibble: 2 × 4 Product Likeability Score Quality <chr> <dbl> <dbl> <dbl> 1 A 80 31 16 2 C 70 33 18
dplyr的分组操作针对大数据集做了性能优化,内存占用低、处理速度快,适合大规模数据场景。
Base R 替代方法
无需加载第三方包时,可通过基础R工具实现:
data <- data.frame(Product=c('A', 'B', 'B', 'C'), Likeability=c(80, 80, 82, 70), Score=c(31, 33, 33, 33), Quality=c(16, 32, 56, 18)) # 统计每个Product的出现频次 product_freq <- table(data$Product) # 筛选仅出现一次的Product值 target_products <- names(product_freq[product_freq == 1]) # 提取对应行数据 result <- data[data$Product %in% target_products, ] print(result)
执行后输出:
Product Likeability Score Quality 1 A 80 31 16 4 C 70 33 18
关键说明
- 两种方法都会完全移除所有重复出现的Product对应的行,不会保留重复组中的任意一条记录,完全匹配需求。
- 超大型数据集场景下,优先选择
dplyr方案,其底层优化能显著提升处理效率。
内容的提问来源于stack exchange,提问作者econ9595
相关产品推荐
相关产品推荐

