使用dplyr筛选Top N分组后用ggplot绘图的优化方案咨询
Great question! Your current approach gets the job done, but we can streamline it to be more concise and readable. Let's start by fixing the sample data (your original tribble had a syntax issue) and then walk through a few improved methods.
First, here's the corrected sample dataset:
library(tidyverse) ap <- tribble( ~year, ~examName, 2014, "Statistics", 2015, "Statistics", 2016, "Statistics", 2016, "Statistics", 2016, "Statistics", 2016, "Statistics", 2017, "Statistics", 2017, "Statistics", 2017, "Statistics", 2017, "Statistics", 2017, "Statistics", 2013, "Macroeconomics", 2013, "Macroeconomics", 2014, "Macroeconomics", 2015, "Macroeconomics", 2016, "Macroeconomics", 2016, "Macroeconomics", 2016, "Macroeconomics", 2016, "Macroeconomics", 2016, "Macroeconomics", 2017, "Macroeconomics", 2017, "Macroeconomics", 2017, "Macroeconomics", 2017, "Macroeconomics", 2017, "Macroeconomics", 2017, "Macroeconomics", 2013, "Calculus", 2014, "Calculus", 2015, "Calculus", 2016, "Calculus", 2017, "Calculus", 2017, "Psychology", 2017, "Psychology", 2017, "Psychology", 2017, "Psychology", 2017, "Psychology", 2018, "Psychology", 2018, "Psychology" )
Method 1: Extract Top N Names First (Most Readable)
This approach separates the "get top groups" step from the filtering/plotting step, making your code easy to follow:
# Step 1: Get the top 2 exam names by occurrence top_exams <- ap %>% count(examName, sort = TRUE) %>% head(2) %>% pull(examName) # Extract just the names as a vector # Step 2: Filter the original data and plot ap %>% filter(examName %in% top_exams) %>% count(examName, year) %>% ggplot(aes(x = year, y = n, group = examName)) + geom_line() + facet_wrap(~examName)
Why this works: Using pull() gives us a simple vector of the top exam names, which we can use with %in% to filter the original dataset. This avoids creating an intermediate dataframe with inner_join, keeping the pipeline clean.
Method 2: Inline Filter with Inner Join (No Extra Variables)
If you prefer to keep everything in one pipeline, you can nest the top-N calculation directly inside inner_join:
ap %>% inner_join( ap %>% count(examName, sort = TRUE) %>% head(2), by = "examName" ) %>% count(examName, year) %>% ggplot(aes(x = year, y = n, group = examName)) + geom_line() + facet_wrap(~examName)
Why this works: This cuts out the intermediate ap_top dataframe by performing the top-N count as a subquery within inner_join. It's concise and keeps all logic in a single flow.
Method 3: Using forcats for Flexible Grouping
If you ever want to include an "Other" category for non-top groups (or just filter them out), fct_lump_n is a powerful tool:
ap %>% # Lump all non-top-2 exams into "Other" mutate(examName = fct_lump_n(examName, n = 2, ties.method = "first")) %>% # Filter out the "Other" group (remove this line if you want to keep it) filter(examName != "Other") %>% count(examName, year) %>% ggplot(aes(x = year, y = n, group = examName)) + geom_line() + facet_wrap(~examName)
Why this works: fct_lump_n automatically groups the least frequent categories into "Other". While we filter it out here, this method is great if you later want to visualize top groups alongside the rest of the data.
Recommendation
Method 1 is my go-to for most cases—it's the most readable, and separating the top-N extraction makes it easy to adjust the number of groups later (just change head(2) to head(N)).
内容的提问来源于stack exchange,提问作者talbe009

