在R语言中按分类分组计算数值变量的均值、中位数、众数
Hey there! Let's tackle this problem step by step. You want to calculate descriptive stats (mean, median, mode) for Mass_kg grouped by Order_or_higher in R—here's how you can do it easily:
First, handle the mode (since R doesn't have a built-in mode function)
Unlike mean and median, R doesn’t come with a default function to calculate mode. Let’s make a simple one that works for your numeric Mass_kg values:
get_mode <- function(x) { # Get unique values and find which occurs most frequently unique_vals <- unique(x) unique_vals[which.max(tabulate(match(x, unique_vals)))] }
Note: This returns the first mode if there are multiple values with the same highest frequency. If you need to return all modes, use this modified version instead:
get_mode_multi <- function(x) { freq_table <- table(x) max_frequency <- max(freq_table) # Return all values that match the highest frequency names(freq_table[freq_table == max_frequency]) }
Option 1: Use dplyr (clean, readable approach)
dplyr is perfect for grouped operations. First install it if you haven’t, then load the package and compute your stats:
# Install dplyr if needed (run once) install.packages("dplyr") # Load the package library(dplyr) # Replace `your_data` with the name of your actual dataset grouped_stats <- your_data %>% group_by(Order_or_higher) %>% summarise( average_mass = mean(Mass_kg, na.rm = TRUE), # na.rm handles missing values median_mass = median(Mass_kg, na.rm = TRUE), mode_mass = get_mode(Mass_kg) # Use get_mode_multi if you need all modes ) # View the result print(grouped_stats)
Option 2: Base R (no extra packages needed)
If you prefer not to use external packages, you can use aggregate() and merge the results:
# Calculate each stat separately mean_data <- aggregate(Mass_kg ~ Order_or_higher, data = your_data, FUN = mean, na.rm = TRUE) median_data <- aggregate(Mass_kg ~ Order_or_higher, data = your_data, FUN = median, na.rm = TRUE) mode_data <- aggregate(Mass_kg ~ Order_or_higher, data = your_data, FUN = get_mode) # Merge into one dataframe grouped_stats_base <- merge(merge(mean_data, median_data, by = "Order_or_higher"), mode_data, by = "Order_or_higher") # Rename columns for clarity colnames(grouped_stats_base) <- c("Order_or_higher", "average_mass", "median_mass", "mode_mass") # View the result print(grouped_stats_base)
Just remember to replace your_data with the actual name of your dataset. The na.rm = TRUE argument is important—it ensures missing values in Mass_kg don’t break your calculations!
内容的提问来源于stack exchange,提问作者SomebodyNeedingHelp

