You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效处理data.frame分组统计:连续变量极值与分类变量频数占比

Hey there! Let's streamline this data summarization task so you don't have to write repetitive lines of code for each min/max value. Here's a clean, efficient approach using tidyverse tools:

First, Set Up Our Data & Tools

First, load the necessary packages (we'll use dplyr for grouping/summarizing and tidyr for reshaping data) and define your sample data:

library(dplyr)
library(tidyr)

# Sample Data
data <- data.frame(
  "group" = c(rep(0:1, 10)), 
  "value1" = c(1:10), 
  "value2" = seq(11:20), 
  "value3" = as.factor(rep(1:3, length=10))
)

1. Batch-Process Continuous Variables (Min/Max)

Instead of writing separate lines for value1_min0, value1_max1, etc., we can use group_by() + summarise(across()) to compute min and max for all continuous variables in one go:

continuous_stats <- data %>%
  group_by(group) %>%
  summarise(
    # Apply min and max to value1 and value2, auto-name columns
    across(c(value1, value2), list(min = min, max = max), .names = "{.col}_{.fn}")
  )

This outputs a data frame where each row represents a group, with columns like value1_min, value1_max, value2_min, value2_max—no repetitive code needed, even if you add more continuous variables later!

2. Calculate Factor Variable Counts & Proportions

For value3 (your factor variable), we'll count occurrences of each level per group, then compute their proportion relative to the group's total size. We'll reshape the result to wide format to match the structure of our continuous stats:

factor_stats <- data %>%
  group_by(group) %>%
  count(value3, name = "n") %>% # Count each value3 level per group
  mutate(prop = n / sum(n)) %>% # Calculate proportion within the group
  pivot_wider(
    names_from = value3,
    values_from = c(n, prop),
    names_glue = "value3_{.value}_{value3}" # Name columns clearly
  )

This gives us columns like value3_n_1, value3_prop_1, value3_n_2, etc., showing counts and proportions for each value3 level in every group.

3. Combine All Stats Into One Data Frame

Finally, merge the two summary data frames together using group as the key:

final_summary <- continuous_stats %>%
  left_join(factor_stats, by = "group")

Alternative: Variable-Wise Summary (If Your Target Table Rows Are Variables)

If your desired output has rows for each variable (instead of groups), we can reshape the data to match that structure:

variable_wise_summary <- data %>%
  # Convert to long format to handle all variables uniformly
  pivot_longer(cols = -group, names_to = "variable", values_to = "value") %>%
  group_by(group, variable) %>%
  summarise(
    min = ifelse(is.numeric(value), min(value), NA),
    max = ifelse(is.numeric(value), max(value), NA),
    n = n(),
    prop = ifelse(is.factor(value), n() / sum(n(), na.rm = TRUE), NA)
  ) %>%
  # Reshape back to wide format with group-specific stats
  pivot_wider(
    names_from = group,
    values_from = c(min, max, n, prop),
    names_glue = "{.value}_group{group}"
  )

This creates rows for value1, value2, value3, with columns like min_group0, max_group1, prop_group0—perfect if your target figure organizes stats by variable.

内容的提问来源于stack exchange,提问作者bvowe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:36:12