You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言:基于数据框计算合并指标的实现方案

Calculating Aggregated Metrics by Region in R

Got it, let's work through this problem together. First, I'll help you complete the incomplete sample data frame you provided, then show you two common approaches to calculate aggregated (merged) metrics grouped by the Region column.

Step 1: Complete the Sample Data Frame

Your original data frame code cuts off for the CEE, FRANCE, and UK&I regions. Here's a fully runnable version with reasonable probability distributions for each region (plus a seed to ensure reproducibility):

set.seed(123) # Fixes random sampling for consistent results
df <- data.frame(
  Region = c(rep("NORDICS", 1100),
             rep("DACH", 900),
             rep("MED", 1800),
             rep("CEE", 15000),
             rep("FRANCE", 2000),
             rep("UK&I", 2500)),
  Score = c(
    sample(seq(1, 4, 1), size = 1100, replace = TRUE, prob = c(0.6, 0.2, 0.1, 0.1)),
    sample(seq(1, 4, 1), size = 900, replace = TRUE, prob = c(0.3, 0.3, 0.2, 0.2)),
    sample(seq(1, 4, 1), size = 1800, replace = TRUE, prob = c(0.8, 0.1, 0.05, 0.05)),
    sample(seq(1, 4, 1), size = 15000, replace = TRUE, prob = c(0.4, 0.3, 0.2, 0.1)),
    sample(seq(1, 4, 1), size = 2000, replace = TRUE, prob = c(0.5, 0.25, 0.15, 0.1)),
    sample(seq(1, 4, 1), size = 2500, replace = TRUE, prob = c(0.35, 0.3, 0.2, 0.15))
  )
)

Step 2: Calculate Aggregated Metrics

We'll cover two practical methods: one using the intuitive dplyr package (part of the tidyverse) and another using only base R (no extra packages required).

Method 1: Using dplyr (Tidyverse)

First, install and load the package if you haven't already:

install.packages("dplyr")
library(dplyr)

Now compute key aggregated metrics per region with readable, concise code:

aggregated_metrics <- df %>%
  group_by(Region) %>%
  summarize(
    Total_Samples = n(), # Total observations per region
    Average_Score = round(mean(Score), 2), # Mean score (rounded)
    Median_Score = median(Score), # Median score
    # Percentage of responses for each Score level
    Pct_Score_1 = round((sum(Score == 1)/Total_Samples)*100, 2),
    Pct_Score_2 = round((sum(Score == 2)/Total_Samples)*100, 2),
    Pct_Score_3 = round((sum(Score == 3)/Total_Samples)*100, 2),
    Pct_Score_4 = round((sum(Score == 4)/Total_Samples)*100, 2)
  )

# View the final results
print(aggregated_metrics)

This groups the data by Region and calculates all the core metrics you'd likely need for analysis.

Method 2: Using Base R

If you prefer avoiding external packages, use aggregate() and table() to achieve the same result:

# Calculate core stats (count, mean, median)
base_stats <- aggregate(Score ~ Region, data = df, FUN = function(x) {
  c(
    Total_Samples = length(x),
    Average_Score = round(mean(x), 2),
    Median_Score = median(x)
  )
})

# Clean up the output format
base_stats_clean <- do.call(data.frame, base_stats)

# Calculate percentage of each Score category
score_proportions <- prop.table(table(df$Region, df$Score), margin = 1)*100
score_proportions_df <- round(as.data.frame.matrix(score_proportions), 2)
colnames(score_proportions_df) <- paste0("Pct_Score_", colnames(score_proportions_df))

# Merge stats and proportions into one data frame
final_base_metrics <- merge(base_stats_clean, score_proportions_df, by.x = "Region", by.y = "row.names")

# View the results
print(final_base_metrics)

Customizing Metrics

If you need additional metrics (like standard deviation, minimum/maximum score, or custom aggregations), you can easily extend either code:

  • For dplyr: Add a new line in the summarize() block (e.g., Std_Dev = sd(Score)).
  • For base R: Add a new element to the function inside aggregate().

内容的提问来源于stack exchange,提问作者Varun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:07:11