跨表计算函数开发:将逻辑变量统计结果规整为指定格式数据框
Got it, let's tackle this problem step by step. First, we'll fix the column name issue for a single logical variable, then build a reusable function to iterate through all your logical variables and combine everything into the exact data frame format you want.
Step 1: Fix Column Names for a Single Variable
Your existing code generates a column named Freq for each variable's stats. We can use dplyr's renaming tools to swap that out for the actual logical variable name, plus convert row names to a proper column for easier merging later:
library(dplyr) library(tibble) library(purrr) # Example for var_log1 single_var_output <- example %>% filter(var_log1 == TRUE) %>% pull(categ_var) %>% table(useNA = 'always') %>% as.data.frame() %>% rename(var_log1 = Freq) %>% # Replace Freq with the variable name rownames_to_column("categ_var") # Turn row names into a formal column
Step 2: Build a Function to Iterate All Logical Variables
We'll create a function that automatically detects all var_log-prefixed variables, calculates their stats, and merges everything into one wide-format data frame (with missing counts filled as 0):
get_logical_category_stats <- function(data, category_col = "categ_var", var_prefix = "var_log") { # Grab all columns matching the var_log pattern log_vars <- colnames(data)[grepl(var_prefix, colnames(data))] # Store each variable's stats in a list stats_collection <- list() for (var in log_vars) { var_stats <- data %>% filter(.data[[var]] == TRUE) %>% # Safely reference dynamic variable names pull({{category_col}}) %>% # Use quasiquotation for the category column table(useNA = 'always') %>% as.data.frame() %>% rename(!!var := Freq) %>% # Dynamically rename the count column rownames_to_column(category_col) stats_collection[[var]] <- var_stats } # Merge all stats together, fill missing values with 0, and sort combined_stats <- stats_collection %>% reduce(full_join, by = category_col) %>% mutate(across(all_of(log_vars), ~ifelse(is.na(.), 0, .))) %>% arrange(category_col) return(combined_stats) }
Step 3: Test the Function with Your Data
First, let's reconstruct your sample dataset to test with:
# Recreate your example data example <- tibble( categ_var = c("cat1", "cat2", "cat2", "cat4", "cat1", "cat3", "cat3", "cat1", "cat3", "cat5", "cat6", "cat7"), var_log1 = c(TRUE, FALSE, FALSE, TRUE, NA, TRUE, TRUE, FALSE, FALSE, TRUE, NA, TRUE), var_log2 = c(TRUE, TRUE, NA, FALSE, NA, FALSE, TRUE, TRUE, NA, FALSE, NA, FALSE), var_log3 = c(FALSE, NA, FALSE, TRUE, TRUE, FALSE, FALSE, NA, FALSE, TRUE, TRUE, FALSE), var_log4 = c(TRUE, TRUE, FALSE, FALSE, FALSE, FALSE, TRUE, TRUE, FALSE, FALSE, FALSE, FALSE) ) # Run the function final_result <- get_logical_category_stats(example) print(final_result)
This will output exactly the format you requested:
| categ_var | var_log1 | var_log2 | var_log3 | var_log4 |
|---|---|---|---|---|
| cat1 | 1 | 2 | 1 | 2 |
| cat2 | 0 | 1 | 0 | 1 |
| cat3 | 2 | 1 | 0 | 1 |
| cat4 | 1 | 0 | 1 | 0 |
| cat5 | 1 | 0 | 1 | 0 |
| cat6 | 0 | 0 | 1 | 0 |
| cat7 | 1 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 |
Key Details:
.data[[var]]safely handles dynamic variable names to avoid dplyr's non-standard evaluation quirks- Quasiquotation (
!!and{{}}) lets us dynamically rename columns and reference the category column reduce(full_join)ensures all categories (including NA) are preserved across all variablesmutate(across(...))replaces any missing counts with 0, matching your desired output
内容的提问来源于stack exchange,提问作者Piotr Konopnicki

