在R中统计列表内各DataFrame的Group值频次并生成新DataFrame
Hey there! Let's work through how to get that comparison DataFrame you need. You've got 80 DataFrames loaded into data_list, each with an ID and Group column, and you want to count how often each Group value shows up in each table, then combine everything into a single easy-to-compare table. Here are two approaches you can use:
Using Tidyverse (dplyr + tidyr + purrr)
This method uses the tidyverse suite of packages for a clean, readable workflow. First, make sure you have the packages installed and loaded:
# Install if needed # install.packages("tidyverse") library(tidyverse)
Step 1: Name your DataFrames
First, we'll assign clear names to each DataFrame in your list (like Table_1, Table_2, ..., Table_80). If you'd prefer to use the original filenames instead, swap out the paste0 line with basename(file_names) to use the file's base name as the column header.
# Name each DataFrame in the list named_data <- setNames(data_list, paste0("Table_", seq_along(data_list)))
Step 2: Calculate Group counts for each table
We'll use imap() (from purrr) to iterate over the named list—this lets us access both the DataFrame and its name at the same time. We'll count each Group value and rename the count column to match the table's name.
# Generate count tables for each DataFrame count_list <- imap(named_data, function(df, table_name) { df %>% count(Group) %>% rename(!!table_name := n) # Rename count column to the table's name })
Step 3: Combine counts into a single comparison table
We'll use reduce() to full-join all the count tables together (so every Group value from any table is included), then replace any missing counts (where a Group didn't appear in a table) with 0. Finally, we'll sort by Group in descending order to match your example format.
# Combine and clean up the comparison table comparison_df <- count_list %>% reduce(full_join, by = "Group") %>% mutate(across(-Group, ~replace_na(., 0))) %>% arrange(desc(Group)) # View the result head(comparison_df)
Using Base R
If you prefer not to use external packages, here's a base R approach that achieves the same result:
Step 1: Collect all unique Group values
First, we'll gather every unique Group value from all your DataFrames to ensure we don't miss any in the final table.
# Get all unique Group values across all DataFrames all_groups <- unique(unlist(lapply(data_list, function(df) df$Group)))
Step 2: Initialize and build the comparison table
We'll start with a DataFrame of all unique Group values, then add a column for each table with the count of each Group (filling in 0 for any Group that doesn't appear in the table).
# Initialize the comparison DataFrame comparison_df <- data.frame(Group = all_groups) # Add count columns for each table for (i in seq_along(data_list)) { # Create a table name (e.g., Table_1) table_name <- paste0("Table_", i) # Count Group occurrences in the current table group_counts <- table(data_list[[i]]$Group) # Match counts to all_groups, replace missing values with 0 comparison_df[[table_name]] <- as.integer(group_counts[as.character(all_groups)]) comparison_df[[table_name]][is.na(comparison_df[[table_name]])] <- 0 } # Sort by Group in descending order to match your example comparison_df <- comparison_df[order(-comparison_df$Group), ] # View the result head(comparison_df)
Both methods will give you a DataFrame exactly like the example you shared, with Group as the first column and each subsequent column showing the count of that Group in the corresponding table.
内容的提问来源于stack exchange,提问作者schande

