You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言DataFrame中保留某变量最常见的前N个观测值?

保留数据集中出现次数最多的两个城市的行

原始数据

dt <- structure(list(ID = 1:10, City = c("New York", "New York", "LA", 
"LA", "LA", "Boston", "Chicago ", "New York", "LA", "New York"
), Random_Info_To_Keep = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 
1L)), class = "data.frame", row.names = c(NA, -10L))

方法一:Base R 实现

  1. 统计各城市的出现次数
  2. 筛选出出现次数Top2的城市
  3. 保留包含这些城市的行
# 统计城市出现频次
city_freq <- table(dt$City)
# 获取Top2高频城市
top2_cities <- names(sort(city_freq, decreasing = TRUE)[1:2])
# 筛选目标行
dt2 <- dt[dt$City %in% top2_cities, ]

方法二:dplyr 实现(tidyverse风格)

借助dplyr的计数和筛选功能,代码更简洁直观:

library(dplyr)

dt2 <- dt %>%
  add_count(City) %>%  # 给每行添加对应城市的出现次数
  filter(n %in% sort(unique(n), decreasing = TRUE)[1:2]) %>%  # 筛选Top2频次的城市行
  select(-n)  # 移除临时生成的计数列

验证结果

执行dput(dt2)会得到预期输出:

structure(list(ID = c(1L, 2L, 3L, 4L, 5L, 8L, 9L, 10L), City = c("New York", 
"New York", "LA", "LA", "LA", "New York", "LA", "New York"), 
    Random_Info_To_Keep = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L)), class = "data.frame", row.names = c(NA, 
-8L))

内容的提问来源于stack exchange,提问作者Jamie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 11:20:17