在R语言中如何按多列分组计算分性别分年度的就业率
R语言实现分性别分年度就业率统计
方案1:tidyverse生态实现(推荐,语法简洁易读)
首先加载依赖包:
library(dplyr) library(tidyr)
生成长格式统计结果
result_long <- my_data %>% # 按性别、年份分组 group_by(Gender, Year) %>% summarise( # 组内总人数 total = n(), # 组内就业人数 employed_cnt = sum(Employment_status == "Employed"), # 计算就业率 employment_rate = employed_cnt / total, .groups = "drop" )
转换为宽格式(和预期输出结构一致,年份为列、性别为行)
result_wide <- result_long %>% pivot_wider( id_cols = Gender, names_from = Year, values_from = employment_rate )
方案2:Base R实现(无需安装加载第三方包)
# 统计各分组总人数 total_cnt <- aggregate(ID ~ Gender + Year, my_data, length) # 统计各分组就业人数 employed_cnt <- aggregate(Employment_status ~ Gender + Year, my_data, function(x) sum(x == "Employed")) # 合并数据计算就业率 stat_res <- merge(total_cnt, employed_cnt, by = c("Gender", "Year")) stat_res$employment_rate <- stat_res$Employment_status / stat_res$ID # 转换为宽表 stat_wide <- reshape( stat_res, idvar = "Gender", timevar = "Year", direction = "wide", drop = c("ID", "Employment_status") )
注意事项
如果你的统计口径需要排除Other类的就业状态,可在统计前先做数据过滤:
# tidyverse 过滤写法 my_data <- filter(my_data, Employment_status %in% c("Employed", "Unemployed")) # Base R 过滤写法 my_data <- my_data[my_data$Employment_status %in% c("Employed", "Unemployed"), ]
内容的提问来源于stack exchange,提问作者Merlin Nuiamäe
相关产品推荐
相关产品推荐

