如何为重复分组的美国军事援助数据集计算年度均值(R语言实现)
Got it, let's tackle this problem step by step! You want to collapse your military aid dataset so each country has exactly one record per fiscal year, calculating the mean of the obligation amounts. Since you're already working with a tibble, we'll use tidyverse tools (specifically dplyr) to get this done cleanly.
Step 1: Fix Data Types First
Looking at your raw data, the obligation columns (Obligations..Historical.Dollars. and Obligations..Constant.Dollars.) are stored as character strings (they’re wrapped in quotes in your tribble). We need to convert these to numeric first—you can’t calculate a mean on text!
Step 2: Group and Aggregate
We’ll group the data by Fiscal Year and Country (the two columns we want unique combinations of), then compute the mean for the obligation values. For columns like Region that don’t vary within a group, we can keep the single consistent value using first() or unique().
Full Code
library(dplyr) # Clean and aggregate the data military_aggregated <- military_struct %>% # Convert obligation columns to numeric mutate( Obligations..Historical.Dollars. = as.numeric(Obligations..Historical.Dollars.), Obligations..Constant.Dollars. = as.numeric(Obligations..Constant.Dollars.) ) %>% # Group by the columns we want unique records for group_by(Fiscal.Year, Country) %>% # Calculate means and retain consistent columns summarize( Region = first(Region), Mean.Historical.Obligations = mean(Obligations..Historical.Dollars.), Mean.Constant.Obligations = mean(Obligations..Constant.Dollars.), .groups = "drop" # Remove grouping structure after aggregation ) # View the result military_aggregated
What This Does
mutate(): Converts the obligation columns from character to numeric, enabling mathematical operations.group_by(Fiscal.Year, Country): Creates groups for every unique country-fiscal year pair.summarize():Region = first(Region): Keeps the region value (since each country belongs to one region, all rows in the group share this value).mean(): Computes the average of the historical and constant obligation amounts for each group.
.groups = "drop": Removes the grouping metadata so your final tibble is a standard, ungrouped dataset.
Example Output Check
For Botswana in 2001 (which has two rows in your raw data), the historical obligation mean will be (1597000 + 663000)/2 = 1130000—the code will calculate this correctly.
内容的提问来源于stack exchange,提问作者Cristina Lopez

