R语言新手:如何为汇总数据框添加极差、频率、众数统计行?
Hey there! As a fellow R learner, I totally get how frustrating it can be when pre-built packages don't check all your boxes. Let's walk through exactly how to create that custom summary table you need, step by step.
First: Fixing the Range Calculation
You're right that range() returns two values (min and max) instead of the single "range" value (max - min). The fix is super simple—you have two easy options:
- Use
diff(range(your_variable, na.rm = TRUE)): Thediff()function subtracts the first value in the range from the second, giving you the exact range you want. - Calculate it directly:
max(your_variable, na.rm = TRUE) - min(your_variable, na.rm = TRUE)
The na.rm = TRUE flag is crucial here—it ensures missing values don't break your calculations.
Next: Create a Mode Function (Since R Doesn't Have One Built-In)
Base R doesn't include a native mode function, but we can write a quick, reusable one that works for both numeric and categorical variables:
get_mode <- function(x, na.rm = FALSE) { if (na.rm) { x <- x[!is.na(x)] } unique_vals <- unique(x) unique_vals[which.max(tabulate(match(x, unique_vals)))] }
This function finds the most frequent value in your variable, and includes an option to ignore missing values if needed.
Build Your Custom Summary Table
Now let's put it all together to create a table that includes all your desired stats. I'll show you two approaches—one with base R, and one using the tidyverse (dplyr/tidyr) if you prefer that workflow.
Approach 1: Base R
Let's use the mtcars dataset as an example (replace this with your own data frame):
# Example data (use your actual data frame instead) my_data <- mtcars[, c("mpg", "disp", "hp")] # Define a function to calculate all your desired stats for a single variable calc_stats <- function(x) { if (is.numeric(x)) { c( Mean = mean(x, na.rm = TRUE), Median = median(x, na.rm = TRUE), Min = min(x, na.rm = TRUE), Max = max(x, na.rm = TRUE), Range = diff(range(x, na.rm = TRUE)), Mode = get_mode(x, na.rm = TRUE), Total_Observations = length(na.omit(x)), Mode_Frequency = max(tabulate(match(x, unique(x[!is.na(x)])))) ) } # Add an else block here if you need stats for categorical variables too! } # Apply the function to every variable in your data frame summary_table <- t(sapply(my_data, calc_stats)) # Convert to a data frame for readability summary_table <- as.data.frame(summary_table)
Here, I included both total observations and the frequency of the mode—adjust these based on exactly what you mean by "frequency"!
Approach 2: Tidyverse (dplyr + tidyr)
If you're using the tidyverse (it's great for data manipulation!), this approach is more readable:
library(dplyr) library(tidyr) # Use your own data frame instead of mtcars my_data <- mtcars[, c("mpg", "disp", "hp")] summary_table <- my_data %>% summarise(across(everything(), list( mean = ~mean(., na.rm = TRUE), median = ~median(., na.rm = TRUE), min = ~min(., na.rm = TRUE), max = ~max(., na.rm = TRUE), range = ~diff(range(., na.rm = TRUE)), mode = ~get_mode(., na.rm = TRUE), total_obs = ~length(na.omit(.)), mode_freq = ~max(tabulate(match(., unique(.[!is.na(.)])))) ))) %>% # Reshape the data into a clean table format pivot_longer(everything(), names_sep = "_", names_to = c("Variable", "Statistic")) %>% pivot_wider(names_from = "Statistic", values_from = "value")
This will give you a clean, row-per-variable table with all your custom stats.
Final Notes
- If you have categorical variables, just add an
elseclause to thecalc_statsfunction (base R) or adjust theacross()call (tidyverse) to calculate relevant stats like mode and category frequencies. - Always double-check the
na.rmflags—missing values can mess up your results if you forget them!
Content of the question comes from Stack Exchange, asked by Johnathan James Mawdsley

