You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将Python pandas分组条件列逻辑迁移至R语言的高效实现

Efficient R Implementations for Categorizing and Grouping Large Datasets (5M Rows)

Sample Data

First, generate a representative sample dataset to test the solutions:

set.seed(123)
sample_data <- data.frame(
  Type = sample(c("A", "B", "C"), size = 1000, replace = TRUE),
  Value = rnorm(n = 1000, mean = 0, sd = 0.3)
)

Expected Output

We aim to get a summary table counting Low, Hit, and High categories per Type, like this (values are illustrative):

TypeLowHitHigh
A101702197
B98695207
C105689206

1. Base R Solution

Lightweight, no external packages required:

# Step 1: Assign categories using cut()
sample_data$Category <- cut(
  sample_data$Value,
  breaks = c(-Inf, -0.25, 0.25, Inf),
  labels = c("Low", "Hit", "High"),
  include.lowest = TRUE  # Ensures -0.25 falls into "Hit"
)

# Step 2: Group and count
base_summary <- table(sample_data$Type, sample_data$Category)

# Convert to tidy data frame (optional)
base_summary_df <- as.data.frame(base_summary, responseName = "Count")
colnames(base_summary_df)[1:2] <- c("Type", "Category")

2. data.table Solution (Most Efficient for Large Data)

Optimized for speed and memory, perfect for 5M-row datasets:

library(data.table)

# Convert to data.table
dt <- as.data.table(sample_data)

# Create category and count in one step
dt_summary <- dt[, .(Count = .N), by = .(
  Type,
  Category = cut(Value, breaks = c(-Inf, -0.25, 0.25, Inf), labels = c("Low", "Hit", "High"), include.lowest = TRUE)
)]

# Optional: Reshape to wide format for readability
dt_summary_wide <- dcast(dt_summary, Type ~ Category, value.var = "Count", fill = 0)

3. dplyr Solution (Tidyverse Workflow)

Readable, pipe-based syntax for tidyverse users:

library(dplyr)
library(tidyr)

dplyr_summary <- sample_data %>%
  # Assign categories with case_when() (cut() works too)
  mutate(Category = case_when(
    Value < -0.25 ~ "Low",
    Value >= -0.25 & Value <= 0.25 ~ "Hit",
    Value > 0.25 ~ "High"
  )) %>%
  # Group and count
  group_by(Type, Category) %>%
  summarise(Count = n(), .groups = "drop") %>%
  # Optional: Reshape to wide format
  pivot_wider(names_from = Category, values_from = Count, values_fill = 0)

Performance Notes

  • data.table: Top choice for large datasets due to optimized memory handling and fast grouping operations.
  • dplyr: Efficient with modern tidyverse optimizations, offering intuitive syntax for most users.
  • Base R: Lightweight but may lag in speed compared to the other two for very large datasets.

内容的提问来源于stack exchange,提问作者Marco_CH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 22:40:16