You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中批量筛选数据框中频次Top5国家的全部数据?

批量筛选高频国家数据的解决方案

方法1:直接用%in%快速筛选(最简便)

不需要额外写函数,直接利用R的%in%运算符匹配多个国家值:

# 定义目标国家列表
top_countries <- c("Afghanistan", "Iraq", "Nigeria", "Yemen", "India")
# 批量提取数据
gtd_top <- subset(gtd, country_txt %in% top_countries)

方法2:封装成可复用函数

如果需要频繁针对不同国家列表筛选,可以封装一个通用函数:

# 定义筛选函数:参数为数据框、国家列名、目标国家向量
get_country_subset <- function(df, country_col_name, target_countries) {
  # 将字符串列名转为可识别的变量符号
  col_sym <- sym(country_col_name)
  # 筛选数据
  subset(df, !!col_sym %in% target_countries)
}

# 使用示例
top_countries <- c("Afghanistan", "Iraq", "Nigeria", "Yemen", "India")
gtd_top <- get_country_subset(gtd, "country_txt", top_countries)

方法3:自动化获取TopN国家并筛选

结合你之前统计高频国家的代码,直接自动化完成筛选,不用手动输入国家名:

# 自动提取出现频次最高的5个国家
top_countries <- names(sort(table(gtd$country_txt), decreasing = TRUE)[1:5])
# 批量筛选数据
gtd_top <- subset(gtd, country_txt %in% top_countries)

内容的提问来源于stack exchange,提问作者Humberto Hernandez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 10:30:04