You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R的dplyr中遍历列与分组值生成频次表的实现问题

解决dplyr中遍历列与分组生成频次表的问题

先看我们的原始数据:

df <- data.frame(
  col1 = rep(1, 15), 
  col2 = rep(2, 15), 
  col3 = rep(3, 15), 
  group = c(rep("A", 5), rep("B", 5), rep("C", 5))
)

你的需求很明确:为每个列(col1/col2/col3)和每个分组(A/B/C)的组合生成包含group、value、计数n和占比prop的频次表,但之前的循环写法生成了9个无意义的结果——这是因为你遍历分组后对所有列都跑了频次统计,没有把列和分组的对应关系绑定好,也没正确计算占比。

下面给你两种更优雅的解决方式:

方法1:转长格式一次性生成所有组合的频次表

这种方法是最推荐的,先把宽数据转成“长数据”,让每一行对应一个「列-分组-值」的组合,然后一步到位计算所有统计量:

library(dplyr)
library(tidyr)

# 生成合并后的频次表
full_freq_table <- df %>%
  # 将col1-col3转成长格式,列名存到column,值存到value
  pivot_longer(cols = starts_with("col"), names_to = "column", values_to = "value") %>%
  # 按列和分组分组,计算每组的计数
  group_by(column, group) %>%
  summarise(n = n(), .groups = "drop") %>%
  # 计算占比:这里是按列计算每个分组的占比(比如col1总共有15条,每个分组5条,占比1/3)
  group_by(column) %>%
  mutate(prop = round(n / sum(n), 3)) %>%
  ungroup()

print(full_freq_table)

输出结果会是清晰的合并表:

# A tibble: 9 × 4
  column group     n  prop
  <chr>  <chr> <int> <dbl>
1 col1   A         5 0.333
2 col1   B         5 0.333
3 col1   C         5 0.333
4 col2   A         5 0.333
5 col2   B         5 0.333
6 col2   C         5 0.333
7 col3   A         5 0.333
8 col3   B         5 0.333
9 col3   C         5 0.333

如果你的占比是指整个数据集的占比(比如5/45≈0.111),只需要去掉group_by(column)这一行,直接用mutate(prop = round(n / sum(n), 3))即可。

方法2:为每一列生成单独的频次表

如果你需要为每个列单独输出一个频次表(比如后续要分别绘图),可以用purrr来遍历列名,配合自定义函数:

library(dplyr)
library(purrr)

# 定义生成单列表的函数
get_col_freq <- function(col_name) {
  df %>%
    select(group, all_of(col_name)) %>%
    rename(value = all_of(col_name)) %>%
    group_by(group, value) %>%
    summarise(n = n(), .groups = "drop") %>%
    # 这里的占比逻辑可以按需调整,比如改成n / nrow(df)
    mutate(prop = round(n / sum(n), 3))
}

# 遍历所有目标列,生成列表形式的结果
target_cols <- c("col1", "col2", "col3")
col_freq_list <- map(target_cols, get_col_freq)
names(col_freq_list) <- target_cols

# 查看col1的频次表
print(col_freq_list$col1)

这样col_freq_list里每个元素就是对应列的频次表,比如col_freq_list$col1会输出:

# A tibble: 3 × 4
  group value     n  prop
  <chr> <dbl> <int> <dbl>
1 A         1     5 0.333
2 B         1     5 0.333
3 C         1     5 0.333

这两种方法都避免了繁琐的for循环,用dplyr的管道语法更易读、易维护,而且能精准控制统计逻辑。

内容的提问来源于stack exchange,提问作者Larry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:43:55