You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

跨表计算函数开发:将逻辑变量统计结果规整为指定格式数据框

Got it, let's tackle this problem step by step. First, we'll fix the column name issue for a single logical variable, then build a reusable function to iterate through all your logical variables and combine everything into the exact data frame format you want.


Step 1: Fix Column Names for a Single Variable

Your existing code generates a column named Freq for each variable's stats. We can use dplyr's renaming tools to swap that out for the actual logical variable name, plus convert row names to a proper column for easier merging later:

library(dplyr)
library(tibble)
library(purrr)

# Example for var_log1
single_var_output <- example %>%
  filter(var_log1 == TRUE) %>%
  pull(categ_var) %>%
  table(useNA = 'always') %>%
  as.data.frame() %>%
  rename(var_log1 = Freq) %>%  # Replace Freq with the variable name
  rownames_to_column("categ_var")  # Turn row names into a formal column

Step 2: Build a Function to Iterate All Logical Variables

We'll create a function that automatically detects all var_log-prefixed variables, calculates their stats, and merges everything into one wide-format data frame (with missing counts filled as 0):

get_logical_category_stats <- function(data, category_col = "categ_var", var_prefix = "var_log") {
  # Grab all columns matching the var_log pattern
  log_vars <- colnames(data)[grepl(var_prefix, colnames(data))]
  
  # Store each variable's stats in a list
  stats_collection <- list()
  
  for (var in log_vars) {
    var_stats <- data %>%
      filter(.data[[var]] == TRUE) %>%  # Safely reference dynamic variable names
      pull({{category_col}}) %>%  # Use quasiquotation for the category column
      table(useNA = 'always') %>%
      as.data.frame() %>%
      rename(!!var := Freq) %>%  # Dynamically rename the count column
      rownames_to_column(category_col)
    
    stats_collection[[var]] <- var_stats
  }
  
  # Merge all stats together, fill missing values with 0, and sort
  combined_stats <- stats_collection %>%
    reduce(full_join, by = category_col) %>%
    mutate(across(all_of(log_vars), ~ifelse(is.na(.), 0, .))) %>%
    arrange(category_col)
  
  return(combined_stats)
}

Step 3: Test the Function with Your Data

First, let's reconstruct your sample dataset to test with:

# Recreate your example data
example <- tibble(
  categ_var = c("cat1", "cat2", "cat2", "cat4", "cat1", "cat3", "cat3", "cat1", "cat3", "cat5", "cat6", "cat7"),
  var_log1 = c(TRUE, FALSE, FALSE, TRUE, NA, TRUE, TRUE, FALSE, FALSE, TRUE, NA, TRUE),
  var_log2 = c(TRUE, TRUE, NA, FALSE, NA, FALSE, TRUE, TRUE, NA, FALSE, NA, FALSE),
  var_log3 = c(FALSE, NA, FALSE, TRUE, TRUE, FALSE, FALSE, NA, FALSE, TRUE, TRUE, FALSE),
  var_log4 = c(TRUE, TRUE, FALSE, FALSE, FALSE, FALSE, TRUE, TRUE, FALSE, FALSE, FALSE, FALSE)
)

# Run the function
final_result <- get_logical_category_stats(example)
print(final_result)

This will output exactly the format you requested:

categ_varvar_log1var_log2var_log3var_log4
cat11212
cat20101
cat32101
cat41010
cat51010
cat60010
cat71000
0000

Key Details:

  • .data[[var]] safely handles dynamic variable names to avoid dplyr's non-standard evaluation quirks
  • Quasiquotation (!! and {{}}) lets us dynamically rename columns and reference the category column
  • reduce(full_join) ensures all categories (including NA) are preserved across all variables
  • mutate(across(...)) replaces any missing counts with 0, matching your desired output

内容的提问来源于stack exchange,提问作者Piotr Konopnicki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:32:19