基于雇佣与离职日期的月度员工数量统计R绘图开发求助
按年月统计在职员工数量并可视化(R实现)
核心思路
通过处理员工的雇佣/离职日期,生成全时间范围的年月序列,计算每个月的在职人数,最终用折线图展示趋势。重点处理无离职日期的在职员工,确保统计准确。
步骤1:准备示例数据
先构造符合需求的测试数据集:
employee_data <- data.frame( emp_id = 1:5, original_hire_date = as.Date(c("2018-03-15", "2019-01-02", "2019-07-20", "2020-02-10", "2021-05-01")), termination_date = as.Date(c("2020-12-31", NA, "2022-03-15", NA, "2023-01-01")) )
步骤2:数据清洗与预处理
加载所需包,统一日期格式,将无离职日期的员工标记为当前日期在职:
library(dplyr) library(lubridate) library(ggplot2) # 清洗日期字段,替换NA的离职日期为当前日期 employee_data_clean <- employee_data %>% mutate( original_hire_date = ymd(original_hire_date), termination_date = ifelse(is.na(termination_date), today(), ymd(termination_date)) ) %>% mutate(termination_date = ymd(termination_date))
步骤3:生成统计时间序列
确定统计的时间范围(从最早雇佣月到最晚离职月),生成所有需要统计的年月:
# 计算时间边界 min_month <- floor_date(min(employee_data_clean$original_hire_date), "month") max_month <- floor_date(max(employee_data_clean$termination_date), "month") # 生成年月序列 date_sequence <- seq(min_month, max_month, by = "month") %>% as.Date()
步骤4:计算月度在职人数
这里提供两种方法,小数据集用方法1,大数据集推荐方法2:
方法1:逐行判断法(直观易理解)
monthly_headcount <- tibble(month = date_sequence) %>% rowwise() %>% mutate( # 判断员工是否在当月在职:雇佣日期<=当月最后一天,且离职日期>=当月第一天 headcount = sum( employee_data_clean$original_hire_date <= ceiling_date(month, "month") - days(1) & employee_data_clean$termination_date >= floor_date(month, "month") ) ) %>% ungroup() %>% mutate(month_label = format(month, "%Y年%m月"))
方法2:事件流法(高效适合大数据)
将雇佣记为+1,离职记为-1,通过累加事件计算人数:
# 生成雇佣/离职事件 events <- employee_data_clean %>% transmute(date = original_hire_date, change = 1) %>% bind_rows( employee_data_clean %>% # 离职员工次月不再统计,所以用离职次月第一天作为-1事件 transmute(date = ceiling_date(termination_date, "month"), change = -1) ) %>% arrange(date) # 按月累加事件,得到月度在职人数 monthly_headcount <- events %>% mutate(month = floor_date(date, "month")) %>% group_by(month) %>% summarise(total_change = sum(change)) %>% arrange(month) %>% mutate(headcount = cumsum(total_change)) %>% mutate(month_label = format(month, "%Y年%m月"))
步骤5:可视化趋势图
生成带数据标签的折线图,清晰展示月度人数变化:
ggplot(monthly_headcount, aes(x = month, y = headcount)) + geom_line(color = "#2c3e50", linewidth = 1.2) + geom_point(color = "#e74c3c", size = 3) + geom_text(aes(label = headcount), vjust = -1, size = 3.5) + scale_x_date(date_labels = "%Y年%m月", date_breaks = "3 months") + labs( title = "公司月度在职员工数量趋势", x = "年月", y = "在职员工数" ) + theme_minimal() + theme( axis.text.x = element_text(angle = 45, hjust = 1), plot.title = element_text(hjust = 0.5, size = 14, face = "bold") )
关键注意事项
- 无离职日期的员工:用
today()替换NA,确保其被统计到当前及之后的月份 - 统计逻辑:判断员工是否在当月有至少一天在职,避免漏算入职/离职当月的员工
- 时间序列对齐:所有日期统一按「当月第一天」或「当月最后一天」判断,避免统计偏差
内容的提问来源于stack exchange,提问作者user8272537
相关产品推荐
相关产品推荐

