You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言中如何让ggplot绘制的geom_bar柱状图仅显示TOP 10数据?

R语言中如何让ggplot绘制的geom_bar柱状图仅显示TOP 10数据?

嗨,我看到你尝试用top_n()但没成功,别担心,咱们一步步来搞定这个问题!其实核心问题是你之前的代码没先统计每个目的地的航班总量,直接用top_n()只会抽取10条航班记录,而不是筛选出航班量Top10的目的地~ 下面给你两种可行的解决方法:

方法一:先统计再筛选(推荐,逻辑更清晰)

先对数据做分组统计,拿到每个目的地的航班数后再取Top10,最后关联机场名称绘图:

library(nycflights13)
library(tidyverse)

# 处理数据:筛选+统计+取Top10+关联机场名
plot_2_data <- flights %>%
  filter(origin == "JFK",
         month %in% c(7, 8, 9),
         !is.na(dep_time)) %>%
  # 按目的地分组,统计航班总数
  group_by(dest) %>%
  summarise(flight_count = n(), .groups = "drop") %>%
  # 提取航班量最多的前10个目的地
  top_n(10, flight_count) %>%
  # 关联机场全称
  left_join(airports %>% select(faa, name), by = c("dest" = "faa"))

# 设置绘图尺寸并绘图
options(repr.plot.width = 9, repr.plot.height = 6)
plot_2_data %>%
  ggplot(aes(y = reorder(name, flight_count), x = flight_count)) +
  geom_bar(stat = "identity", fill = "#4e0090") +
  theme_minimal() +
  labs(title = "Top 10 Destinations people traveled to the most (summer 2013)",
       x = "Count",
       y = "Name",
       caption = "Source: Package 'nycflights13'") +
  theme(plot.title = element_text(hjust = 0.5)) + # 标题居中更美观,也可改回你原来的-0.15
  expand_limits(x = c(0, 2900))

方法二:用因子处理直接在ggplot中筛选

如果不想提前统计,也可以用fct_lump_n()把目的地名称这个因子处理成仅保留Top10,剩下的归为“Other”后过滤掉:

library(nycflights13)
library(tidyverse)

options(repr.plot.width = 9, repr.plot.height = 6)

flights %>%
  filter(origin == "JFK",
         month %in% c(7, 8, 9),
         !is.na(dep_time)) %>%
  left_join(airports %>% select(faa, name), by = c("dest" = "faa")) %>%
  # 保留航班量前10的目的地,其余设为"Other"
  mutate(name = fct_lump_n(name, n = 10, w = 1)) %>%
  filter(name != "Other") %>% # 过滤掉"Other"分组
  ggplot(aes(y = reorder(name, name, function(y) length(y)), x = ..count..)) +
  geom_bar(fill = "#4e0090") +
  theme_minimal() +
  labs(title = "Top 10 Destinations people traveled to the most (summer 2013)",
       x = "Count",
       y = "Name",
       caption = "Source: Package 'nycflights13'") +
  theme(plot.title = element_text(hjust = -0.15)) +
  expand_limits(x = c(0, 2900))

两种方法都能实现你的需求,第一种逻辑更直观,也方便后续对统计数据做其他处理,推荐优先用第一种哦~

备注:内容来源于stack exchange,提问作者Tawan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.23 11:44:30