R语言:基于数据框计算合并指标的实现方案
Got it, let's work through this problem together. First, I'll help you complete the incomplete sample data frame you provided, then show you two common approaches to calculate aggregated (merged) metrics grouped by the Region column.
Step 1: Complete the Sample Data Frame
Your original data frame code cuts off for the CEE, FRANCE, and UK&I regions. Here's a fully runnable version with reasonable probability distributions for each region (plus a seed to ensure reproducibility):
set.seed(123) # Fixes random sampling for consistent results df <- data.frame( Region = c(rep("NORDICS", 1100), rep("DACH", 900), rep("MED", 1800), rep("CEE", 15000), rep("FRANCE", 2000), rep("UK&I", 2500)), Score = c( sample(seq(1, 4, 1), size = 1100, replace = TRUE, prob = c(0.6, 0.2, 0.1, 0.1)), sample(seq(1, 4, 1), size = 900, replace = TRUE, prob = c(0.3, 0.3, 0.2, 0.2)), sample(seq(1, 4, 1), size = 1800, replace = TRUE, prob = c(0.8, 0.1, 0.05, 0.05)), sample(seq(1, 4, 1), size = 15000, replace = TRUE, prob = c(0.4, 0.3, 0.2, 0.1)), sample(seq(1, 4, 1), size = 2000, replace = TRUE, prob = c(0.5, 0.25, 0.15, 0.1)), sample(seq(1, 4, 1), size = 2500, replace = TRUE, prob = c(0.35, 0.3, 0.2, 0.15)) ) )
Step 2: Calculate Aggregated Metrics
We'll cover two practical methods: one using the intuitive dplyr package (part of the tidyverse) and another using only base R (no extra packages required).
Method 1: Using dplyr (Tidyverse)
First, install and load the package if you haven't already:
install.packages("dplyr") library(dplyr)
Now compute key aggregated metrics per region with readable, concise code:
aggregated_metrics <- df %>% group_by(Region) %>% summarize( Total_Samples = n(), # Total observations per region Average_Score = round(mean(Score), 2), # Mean score (rounded) Median_Score = median(Score), # Median score # Percentage of responses for each Score level Pct_Score_1 = round((sum(Score == 1)/Total_Samples)*100, 2), Pct_Score_2 = round((sum(Score == 2)/Total_Samples)*100, 2), Pct_Score_3 = round((sum(Score == 3)/Total_Samples)*100, 2), Pct_Score_4 = round((sum(Score == 4)/Total_Samples)*100, 2) ) # View the final results print(aggregated_metrics)
This groups the data by Region and calculates all the core metrics you'd likely need for analysis.
Method 2: Using Base R
If you prefer avoiding external packages, use aggregate() and table() to achieve the same result:
# Calculate core stats (count, mean, median) base_stats <- aggregate(Score ~ Region, data = df, FUN = function(x) { c( Total_Samples = length(x), Average_Score = round(mean(x), 2), Median_Score = median(x) ) }) # Clean up the output format base_stats_clean <- do.call(data.frame, base_stats) # Calculate percentage of each Score category score_proportions <- prop.table(table(df$Region, df$Score), margin = 1)*100 score_proportions_df <- round(as.data.frame.matrix(score_proportions), 2) colnames(score_proportions_df) <- paste0("Pct_Score_", colnames(score_proportions_df)) # Merge stats and proportions into one data frame final_base_metrics <- merge(base_stats_clean, score_proportions_df, by.x = "Region", by.y = "row.names") # View the results print(final_base_metrics)
Customizing Metrics
If you need additional metrics (like standard deviation, minimum/maximum score, or custom aggregations), you can easily extend either code:
- For dplyr: Add a new line in the
summarize()block (e.g.,Std_Dev = sd(Score)). - For base R: Add a new element to the function inside
aggregate().
内容的提问来源于stack exchange,提问作者Varun

