You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中高效批量生成指定列的频数统计tibble?

高效生成多字段频数统计的R代码优化

问题描述

我有一份包含30人信息的数据集,字段包括Ethnicity(种族)、Gender(性别)、School.type(学校类型)、是否享受免费校餐等。目前我用以下代码逐个生成各字段的频数统计:

df <- read.csv("~file")
df %>% select(Ethnicity) %>% group_by(Ethnicity) %>% summarise(freq = n())
df %>% select(Gender) %>% group_by(Gender) %>% summarise(freq = n())
df %>% select(School.type) %>% group_by(School.type) %>% summarise(freq = n())

我希望用1-2行更高效的代码,为8个指定字段(如种族、性别、学校类型等)生成对应的频数统计tibble,同时排除邮编、姓名这类无需统计的字段。以下是种族和性别字段的统计输出示例:

种族统计示例:

# A tibble: 13 × 2
   Ethnicity                             freq
   <chr>                                <int>
 1 Asian or Asian British - Bangladeshi     1
 2 Asian or Asian British - Indian          7
 3 Asian or Asian British - Pakistani       1
 4 Black or Black British - African         5
 5 Black or Black British - Caribbean       2
 6 Chinese                                  3
 7 Mixed - White and Asian                  2
 8 Mixed - White and Black African          1
 9 Mixed - White and Black Caribbean        1
10 Not known/ prefer not to say             1
11 White British                           27
12 White Irish                              1
13 White Other                              5

性别统计示例:

# A tibble: 2 × 2
  Gender  freq
  <chr>  <int>
1 Female    36
2 Male      21

解决方案

方法1:单tibble整合所有统计结果(推荐)

通过长格式转换一次性处理所有目标字段,输出包含字段名、取值、频数的整合结果:

library(tidyverse)

# 替换为你的8个目标字段
target_cols <- c("Ethnicity", "Gender", "School.type", "free_school_meal", "field5", "field6", "field7", "field8")

freq_summary <- df %>%
  select(all_of(target_cols)) %>%
  pivot_longer(everything(), names_to = "category", values_to = "value") %>%
  group_by(category, value) %>%
  summarise(freq = n(), .groups = "drop")

如果需要拆分出单个字段的统计结果,可使用freq_summary %>% filter(category == "Ethnicity"),格式和你给出的示例一致。

方法2:生成以字段名为索引的结果列表

如果希望每个字段的统计结果单独存储,可生成命名列表,通过字段名直接调用:

library(tidyverse)

target_cols <- c("Ethnicity", "Gender", "School.type", "...") # 补充剩余字段

freq_list <- map(target_cols, ~ df %>%
                   select(.x) %>%
                   group_by(across(.x)) %>%
                   summarise(freq = n())) %>%
  set_names(target_cols)

调用示例:freq_list$Gender即可得到性别的频数统计tibble,和你原来的输出完全匹配。


内容的提问来源于stack exchange,提问作者Jess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 23:25:56