如何用lapply统计R语言列表中的唯一值
Got it, let's work through this problem. You mentioned your usual methods for counting unique values in data frames don't play nice when dealing with a list of data frames—and you're right, lapply() (and other apply-family functions) are exactly the tool for this job.
First, let's clean up and formalize your sample data (I filled in the missing values in species2 to make it reproducible):
# Your sample data (completed for reproducibility) species1 <- data.frame(var_1 = c("a","a","a","b", "b", "b"), var_2 = c("c","c","d", "d", "e", "e")) species2 <- data.frame(var_1 = c("f","f","f","g", "g", "g"), var_2 = c("h","h","i", "i", "j", "j")) # Bundle the data frames into a list (this is your starting point) species_list <- list(species1, species2)
Option 1: Count Unique Values Per Column (Base R)
If you want to get the number of unique values for each column in every data frame in the list, use lapply() to iterate over the list, paired with sapply() to loop through each column in the data frame:
# Calculate unique value counts per column for each data frame col_unique_counts <- lapply(species_list, function(df) { sapply(df, function(column) length(unique(column))) }) # View the result col_unique_counts # Output: # [[1]] # var_1 var_2 # 2 3 # # [[2]] # var_1 var_2 # 2 3
Option 2: Count Unique Rows (Base R)
If you need the number of unique rows in each data frame, adjust the inner function to count unique rows instead:
# Calculate unique row counts for each data frame row_unique_counts <- lapply(species_list, function(df) { nrow(unique(df)) }) # View the result row_unique_counts # Output: # [[1]] # [1] 4 # # [[2]] # [1] 4
Option 3: Tidyverse Approach (Purrr + Dplyr)
If you prefer the tidyverse ecosystem, purrr::map() (the tidy equivalent of lapply()) works seamlessly with dplyr functions for cleaner syntax:
library(purrr) library(dplyr) # Unique values per column col_unique_counts_tidy <- map(species_list, ~ summarise_all(.x, ~ length(unique(.)))) # Unique rows row_unique_counts_tidy <- map(species_list, ~ nrow(unique(.x)))
Why Apply-Family Functions?
They eliminate the need for messy manual loops, automatically handle each data frame in the list, and return results in a structured format that's easy to work with afterward.
内容的提问来源于stack exchange,提问作者Jack Dean

