将基于dplyr的R语言Venn_Counts函数转换为data.table实现
Converting Venn_Counts to Use Only data.table
Here's the fully converted function that eliminates the dplyr dependency while preserving the original functionality—and it'll be significantly faster on large datasets, which is your main goal:
library(data.table) how_much = 2000 A <- sample(how_much, replace = TRUE, x = 1:5) B <- sample(how_much, replace = TRUE, x = 1:5) C <- sample(how_much, replace = TRUE, x = 1:5) D <- sample(how_much, replace = TRUE, x = 1:5) VennData = data.table(A, B, C, D) Venn_Counts <- function(dataset, unique_number, operator) { message("Operator argument are: `==` or `<` or `<=` or `>` or `>=`") # Validate unique_number is within 1-5 range if (!data.table::inrange(unique_number, 1, 5)) { message("Unique number must be in range from 1 to 5") return(invisible(NULL)) } # Ensure input is a data.table (works with data.frames too) dt <- data.table::as.data.table(dataset) # Step 1: Apply the operator to all columns to get logical values # Step 2: Group by all logical columns and count occurrences dt[, Map(operator, .SD, unique_number)][, .(count = .N), by = names(.SD)] } # Test the function Venn_Counts(VennData, 4, `==`)
Key Changes Explained:
- Removed dplyr dependency: We replaced
mutate_all,group_by_all, andcountwith data.table-native operations optimized for speed and memory efficiency. Mapfor column-wise operations: Instead ofmutate_all, we useMapto apply your chosen operator to every column in.SD(data.table's Subset of Data). This transforms all columns to logical values without unnecessary data copying.- Grouping and counting:
by = names(.SD)groups by all the logical columns we just created, and.Ngives the number of rows in each group—this replaces the dplyr chaingroup_by_all() %>% count(). - Robust input handling:
as.data.table(dataset)ensures the function works seamlessly with both data.frames and data.tables as input. - Cleaner error handling: We use
messageinstead ofprintand returninvisible(NULL)to avoid cluttering your output with unnecessary NULL values when validation fails.
This version will handle large datasets far more efficiently because data.table avoids the overhead of dplyr's tidy evaluation framework and uses memory-efficient, in-place operations where possible.
内容的提问来源于stack exchange,提问作者Thomas Kyle
相关产品推荐
相关产品推荐

