如何用R语言实现双重循环绘制数值变量多统计量柱状图
Let's fix your code step by step and get those statistic bar charts working properly! The main issues with your nested loop code are: you didn't populate the gh list with actual statistical functions, your loop was overwriting values instead of tracking both variables and statistics, and the data reshaping step didn't preserve the statistic type for plotting.
First, let's recreate your dataframe for testing:
df <- data.frame( A = 1:8, B = c("asd", "fsd", "gs", "asd", "sf", "dfg", "sdfg", "fgdsgd"), C = c(29,24,46,50,43,29,32,24), D = c("sf", "gfd", "asd", "gfg", "fg", "er", "tr", "qw"), E = c(36,56,39,26,56,35,27,31), F = c(44,34,37,23,37,51,28,36) )
Solution 1: Using Base R
This approach uses only base R functions (plus ggplot2 for plotting) so you don't need extra packages:
library(ggplot2) # 1. Define the statistical functions we want to calculate stat_functions <- list(mean = mean, median = median, sd = sd) # 2. Get names of numeric columns numeric_cols <- names(Filter(is.numeric, df)) # 3. Calculate stats for each column stats_data <- lapply(numeric_cols, function(col) { col_values <- df[[col]] # Add na.rm=TRUE to handle missing values if your real data has them sapply(stat_functions, function(fun) fun(col_values, na.rm = TRUE)) }) # 4. Convert results to a dataframe and add variable names stats_df <- as.data.frame(do.call(rbind, stats_data)) stats_df$variable <- numeric_cols # 5. Reshape to long format (required for grouped ggplot bars) stats_long <- reshape( stats_df, varying = names(stat_functions), v.names = "value", timevar = "statistic", times = names(stat_functions), direction = "long" ) # Clean up the long dataframe stats_long <- stats_long[, c("variable", "statistic", "value")] # 6. Plot the grouped bar chart ggplot(stats_long, aes(x = variable, y = value, fill = statistic)) + geom_col(position = "dodge") + # Place bars side-by-side for each variable labs(y = "Statistic Value", x = "Variable", fill = "Statistic Type") + theme_minimal()
Solution 2: Using Tidyverse (More Concise)
If you're comfortable with the tidyverse (dplyr + tidyr), this code is more readable and streamlined:
library(dplyr) library(tidyr) library(ggplot2) stats_long <- df %>% # Select only numeric columns from the dataframe select(where(is.numeric)) %>% # Calculate mean, median, and sd for every numeric column summarise(across( everything(), list(mean = ~mean(., na.rm = TRUE), median = ~median(., na.rm = TRUE), sd = ~sd(., na.rm = TRUE)) )) %>% # Reshape to long format, splitting column names into variable and statistic pivot_longer( everything(), names_to = c("variable", "statistic"), names_sep = "_", values_to = "value" ) # Plot the grouped bar chart ggplot(stats_long, aes(x = variable, y = value, fill = statistic)) + geom_col(position = "dodge") + labs(y = "Statistic Value", x = "Variable", fill = "Statistic Type") + theme_bw()
What Was Wrong With Your Original Code?
- Empty
ghlist: You declaredgh <- list()but never added themean,median, orsdfunctions to it. It should have beengh <- list(mean=mean, median=median, sd=sd). - Overwriting values: Your nested loop assigned
p2[i] <- gh[[j]](df[i]), which overwrote the same variable entry with each new statistic. You need to track both the variable and statistic together. - Incomplete reshaping: Your final
stack(p2)step didn't preserve which statistic each value came from, so ggplot couldn't distinguish between mean/median/sd for the same variable.
Both solutions fix these issues by structuring the data to include both variable names and statistic types, which lets ggplot create clear, grouped side-by-side bars.
内容的提问来源于stack exchange,提问作者Rfer R

