在R中为attitude数据集按列生成含11种统计量的汇总表格
Absolutely! You can absolutely use a for loop to populate your empty measure.dataframe with the 11 summary stats you need for the attitude dataset. Let's walk through this step by step—some of these stats need special handling (looking at you, mode and correlation), so we'll cover those first.
First: Fix the Mode Function
Base R doesn't have a built-in function to calculate the value mode (the most frequent value)—the mode() function in base R returns the storage type of an object instead. So we'll start by defining a custom function to get the mode:
get_mode <- function(x) { uniq_vals <- unique(x) uniq_vals[which.max(tabulate(match(x, uniq_vals)))] }
Next: Map Columns and Define Stat Calculators
First, let's align the columns from the attitude dataset with the columns in your empty dataframe. Then we'll create a list of functions to calculate each statistic, with special handling for things like range, quantiles, and correlation:
# Align attitude columns with your dataframe's target columns attitude_cols <- c("rating", "privileges", "learning", "raises", "critical", "advance") df_target_cols <- c("ratings_measure", "priv_measure", "learn_measure", "raise_measure", "critical_measrue", "advance_measure") # Create a list of functions for each statistic stat_calculators <- list( mean = function(x) round(mean(x, na.rm = TRUE), 2), median = function(x) median(x, na.rm = TRUE), mode = get_mode, max = function(x) max(x, na.rm = TRUE), min = function(x) min(x, na.rm = TRUE), range = function(x) paste0(min(x, na.rm = TRUE), " - ", max(x, na.rm = TRUE)), quantile = function(x) paste(round(quantile(x, na.rm = TRUE), 2), collapse = ", "), IQR = function(x) IQR(x, na.rm = TRUE), var = function(x) round(var(x, na.rm = TRUE), 2), sd = function(x) round(sd(x, na.rm = TRUE), 2), cor = function(x) round(cor(attitude$rating, x, use = "complete.obs"), 2) # Correlate with 'rating' )
Note: For correlation, I chose to calculate each variable's correlation with the rating column (since correlation is a pairwise statistic). If you need something different (like a full correlation matrix), you'd need to adjust your dataframe structure, but this fits your current setup.
Finally: Run the For Loop
Now we'll use a nested for loop to iterate over each statistic and each column, filling in the values:
# Loop through each statistic first for (stat_idx in seq_along(cntrl_measures)) { current_stat <- cntrl_measures[stat_idx] calc_func <- stat_calculators[[current_stat]] # Then loop through each column to calculate the stat for (col_idx in seq_along(df_target_cols)) { target_col <- df_target_cols[col_idx] attitude_data <- attitude[[attitude_cols[col_idx]]] # Fill the cell (convert to character to handle mixed types like strings for range) measure.dataframe[stat_idx, target_col] <- as.character(calc_func(attitude_data)) } } # Check the filled dataframe print(measure.dataframe)
Optional: A More R-Friendly Alternative (No For Loops)
If you ever want to skip the for loop, you can use the purrr package to vectorize this process (more idiomatic in R):
library(purrr) library(tibble) # Calculate stats for each column summary_stats <- map_dfr(attitude, function(col) { tibble( mean = round(mean(col, na.rm = TRUE), 2), median = median(col, na.rm = TRUE), mode = get_mode(col), max = max(col, na.rm = TRUE), min = min(col, na.rm = TRUE), range = paste0(min(col, na.rm = TRUE), " - ", max(col, na.rm = TRUE)), quantile = paste(round(quantile(col, na.rm = TRUE), 2), collapse = ", "), IQR = IQR(col, na.rm = TRUE), var = round(var(col, na.rm = TRUE), 2), sd = round(sd(col, na.rm = TRUE), 2), cor = round(cor(attitude$rating, col, use = "complete.obs"), 2) ) }) # Transpose to match your original dataframe structure measure.dataframe_clean <- tibble(cntrl_measures = names(summary_stats)) %>% bind_cols(t(summary_stats)) %>% rename_with(~df_target_cols, 2:ncol(.))
内容的提问来源于stack exchange,提问作者MCP_infiltrator

