将Python pandas分组条件列逻辑迁移至R语言的高效实现
Efficient R Implementations for Categorizing and Grouping Large Datasets (5M Rows)
Sample Data
First, generate a representative sample dataset to test the solutions:
set.seed(123) sample_data <- data.frame( Type = sample(c("A", "B", "C"), size = 1000, replace = TRUE), Value = rnorm(n = 1000, mean = 0, sd = 0.3) )
Expected Output
We aim to get a summary table counting Low, Hit, and High categories per Type, like this (values are illustrative):
| Type | Low | Hit | High |
|---|---|---|---|
| A | 101 | 702 | 197 |
| B | 98 | 695 | 207 |
| C | 105 | 689 | 206 |
1. Base R Solution
Lightweight, no external packages required:
# Step 1: Assign categories using cut() sample_data$Category <- cut( sample_data$Value, breaks = c(-Inf, -0.25, 0.25, Inf), labels = c("Low", "Hit", "High"), include.lowest = TRUE # Ensures -0.25 falls into "Hit" ) # Step 2: Group and count base_summary <- table(sample_data$Type, sample_data$Category) # Convert to tidy data frame (optional) base_summary_df <- as.data.frame(base_summary, responseName = "Count") colnames(base_summary_df)[1:2] <- c("Type", "Category")
2. data.table Solution (Most Efficient for Large Data)
Optimized for speed and memory, perfect for 5M-row datasets:
library(data.table) # Convert to data.table dt <- as.data.table(sample_data) # Create category and count in one step dt_summary <- dt[, .(Count = .N), by = .( Type, Category = cut(Value, breaks = c(-Inf, -0.25, 0.25, Inf), labels = c("Low", "Hit", "High"), include.lowest = TRUE) )] # Optional: Reshape to wide format for readability dt_summary_wide <- dcast(dt_summary, Type ~ Category, value.var = "Count", fill = 0)
3. dplyr Solution (Tidyverse Workflow)
Readable, pipe-based syntax for tidyverse users:
library(dplyr) library(tidyr) dplyr_summary <- sample_data %>% # Assign categories with case_when() (cut() works too) mutate(Category = case_when( Value < -0.25 ~ "Low", Value >= -0.25 & Value <= 0.25 ~ "Hit", Value > 0.25 ~ "High" )) %>% # Group and count group_by(Type, Category) %>% summarise(Count = n(), .groups = "drop") %>% # Optional: Reshape to wide format pivot_wider(names_from = Category, values_from = Count, values_fill = 0)
Performance Notes
- data.table: Top choice for large datasets due to optimized memory handling and fast grouping operations.
- dplyr: Efficient with modern tidyverse optimizations, offering intuitive syntax for most users.
- Base R: Lightweight but may lag in speed compared to the other two for very large datasets.
内容的提问来源于stack exchange,提问作者Marco_CH
相关产品推荐
相关产品推荐

