如何在R中并排查看列表?多数据集变量对比需求
Great question! You can absolutely build a clean, aligned comparison using purrr, janitor, and tidyr—here's a step-by-step implementation that handles multiple data frames, aligns common variables first, and pushes unique variables to the end:
Step 1: Setup and Sample Data
First, let's adjust your sample data to ensure factors are properly defined (since your original code targets factor variables):
library(tidyverse) library(janitor) # Sample data with factor variables df1 <- data.frame( var1 = factor(c("a", "b")), var2 = factor(c("A", "B")), var3 = factor(c("checking")) ) df2 <- data.frame( var1 = factor(letters), var2 = factor(LETTERS), var4 = factor(c("testing")) ) # Store data frames in a list (works for 2+ frames) df_list <- lst(df1, df2)
Step 2: Define Ordered Variables (Common First, Unique Last)
We'll first identify which factor variables are common across all data frames and which are unique to some, then order them accordingly:
# Get factor variables for each data frame factor_vars_per_df <- df_list %>% map(~select_if(.x, is.factor) %>% names()) # All unique factor variables across all frames all_factor_vars <- factor_vars_per_df %>% flatten_chr() %>% unique() # Variables present in every data frame (common) common_factor_vars <- all_factor_vars %>% keep(~every(factor_vars_per_df, ~.x %in% .)) # Variables present in only some data frames (unique) unique_factor_vars <- setdiff(all_factor_vars, common_factor_vars) # Final ordered list: common variables first, then unique ordered_vars <- c(common_factor_vars, unique_factor_vars)
Step 3: Generate Aligned Comparison Tables
For each ordered variable, we'll create a summary table that includes counts and percentages from every data frame (filling in missing values with NA for frames that don't have the variable):
compare_summaries <- ordered_vars %>% map(function(var) { # For each data frame, get tabyl or empty placeholder if variable is missing df_list %>% map(function(df) { if (var %in% names(df)) { df %>% select(all_of(var)) %>% tabyl() %>% mutate(variable = var) } else { # Placeholder for missing variables tibble( variable = var, !!var := NA_character_, n = NA_integer_, percent = NA_real_ ) } }) %>% bind_rows(.id = "data_frame") %>% # Reshape to have each data frame's stats as separate columns pivot_wider( names_from = data_frame, values_from = c(n, percent), names_glue = "{data_frame}_{.value}" ) %>% # Reorder columns for readability select(variable, all_of(var), starts_with("df1_"), starts_with("df2_")) })
Step 4: View the Results
When you print compare_summaries, you'll get a list of tables where each table corresponds to a variable. Common variables appear first, and unique variables are at the end:
# Print all summaries compare_summaries
Example Output Snippet:
For var1 (common to both frames):
# A tibble: 26 × 5 variable var1 df1_n df1_percent df2_n df2_percent <chr> <fct> <int> <dbl> <int> <dbl> 1 var1 a 1 0.5 1 0.0385 2 var1 b 1 0.5 1 0.0385 3 var1 c NA NA 1 0.0385 ...
For var3 (unique to df1):
# A tibble: 1 × 5 variable var3 df1_n df1_percent df2_n df2_percent <chr> <fct> <int> <dbl> <int> <dbl> 1 var3 checking 1 1 NA NA
Key Features:
- Works with 2+ data frames (just add more to
df_list) - Aligns common variables at the top of the output
- Pushes variables unique to some frames to the end
- Includes counts and percentages (from
tabyl) for each frame - Fills in
NAfor variables missing from a data frame
内容的提问来源于stack exchange,提问作者user63230

