R语言多列分组计数计算占比 循环追加输出到CSV的实现方法
实现方案
直接遍历你需要统计的属性列即可,输出格式和你要求的完全匹配,代码如下:
# 加载依赖(count函数来自dplyr包) library(dplyr) # 你的示例数据 age = c(45, 21, 32, 33, 46) gender = c('female', 'female', 'male', 'male', 'female') income = c('low', 'low', 'medium', 'high', 'low') education = c('high', 'high', 'high', 'medium', 'medium') df = data.frame(age, gender ,income, education) # 配置项 stat_cols <- c("gender", "income", "education") # 要统计的属性列表 output_file <- "stat.csv" totuser <- nrow(df) # 可选:删除已存在的旧文件,避免重复追加内容 if(file.exists(output_file)) file.remove(output_file) # 循环统计写入 for (col in stat_cols) { stat_res <- df %>% count(.data[[col]]) %>% mutate(sot = n / totuser) # 追加写入,每次都保留当前属性的表头 write.table( stat_res, file = output_file, sep = ";", row.names = FALSE, col.names = TRUE, append = TRUE, quote = FALSE ) }
运行后输出的stat.csv内容如下(匹配你的示例结构):
gender;n;sot female;3;0.6 male;2;0.4 income;n;sot low;3;0.6 medium;1;0.2 high;1;0.2 education;n;sot high;3;0.6 medium;2;0.4
补充说明
- 如果你需要计数列名改为
Freq匹配你示例里的写法,只需要在count之后加一行rename(Freq = n)即可 - 如果要统计age属性,建议先做分箱转为分类变量,否则会对每个单独年龄值统计
- 代码里加了
quote=FALSE避免字符串自动被包裹引号,符合常规CSV阅读习惯
内容的提问来源于stack exchange,提问作者bountan
相关产品推荐
相关产品推荐

