RStudio中汇总gapminder数据GDP最值后缺失对应年份的解决方法
Hey there, great question! When you use summarize() with min() and max(), you're only aggregating the gdpPercap values, which is why the corresponding year gets lost. Here are a few clean, dplyr-native ways to get both the min/max gdpPercap AND their associated years for each country:
Method 1: Use slice_min() and slice_max() (Most Intuitive)
This approach pulls the full rows containing the min and max values, then combines them into a single result:
library(dplyr) # Load your data (adjust path if needed) gapminder <- read.csv("gapminder_data.csv") # Get rows with minimum gdpPercap per country min_gdp_data <- gapminder %>% group_by(country) %>% slice_min(gdpPercap, n = 1) %>% # n=1 keeps only one row (adjust if ties exist) select(country, year, gdpPercap) %>% rename(min_year = year, min_gdpPercap = gdpPercap) # Get rows with maximum gdpPercap per country max_gdp_data <- gapminder %>% group_by(country) %>% slice_max(gdpPercap, n = 1) %>% select(country, year, gdpPercap) %>% rename(max_year = year, max_gdpPercap = gdpPercap) # Combine results into one dataframe final_result <- min_gdp_data %>% inner_join(max_gdp_data, by = "country") # View the output head(final_result)
Note: If a country has multiple years with the same min/max gdpPercap, change n = 1 to n = Inf to keep all tied rows, then adjust the join or pivot as needed.
Method 2: Single Pipeline with filter() + pivot_wider()
This keeps everything in one chain by filtering for min/max values first, then reshaping the data:
gapminder %>% group_by(country) %>% filter(gdpPercap == min(gdpPercap) | gdpPercap == max(gdpPercap)) %>% mutate(value_type = ifelse(gdpPercap == min(gdpPercap), "min", "max")) %>% pivot_wider( names_from = value_type, values_from = c(year, gdpPercap), names_sep = "_" ) %>% ungroup()
This will give you a single row per country with columns like year_min, gdpPercap_min, year_max, gdpPercap_max.
Method 3: Direct summarize() with Indexing
If you prefer to stick close to your original code structure, you can use which.min() and which.max() to pull the corresponding year directly in summarize():
gapminder %>% group_by(country) %>% summarize( min_gdpPercap = min(gdpPercap), min_year = year[which.min(gdpPercap)], max_gdpPercap = max(gdpPercap), max_year = year[which.max(gdpPercap)] ) %>% ungroup()
Tip: If you need to handle ties (multiple years with the same min/max), replace year[which.min(gdpPercap)] with paste(year[gdpPercap == min(gdpPercap)], collapse = ", ") to list all relevant years as a single string.
内容的提问来源于stack exchange,提问作者Maerkli

