You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用tidyverse简化查询员工流失率最高与最低部门的代码

简化计算员工流失率最高/最低部门的Tidyverse代码

Hey there! Your original code logic is solid, but we can streamline it drastically using Tidyverse's chained operations and built-in functions—no need for multiple intermediate data frames. Let's break down the optimized version:

Full Simplified Code

library(tidyverse)

# Calculate turnover per department and grab top/bottom results
turnover_summary <- df %>%
  group_by(department) %>%
  summarise(
    total_employees = n(),
    left_count = sum(left == "yes"),
    turnover_rate = left_count / total_employees
  ) %>%
  # Combine highest and lowest turnover rows
  bind_rows(
    slice_max(., turnover_rate, n = 1),
    slice_min(., turnover_rate, n = 1)
  ) %>%
  ungroup()

print(turnover_summary)

Key Optimizations Explained

  • group_by(department): Groups the data by department so all subsequent calculations are done per department subset.
  • summarise(): Computes all necessary metrics in one step instead of splitting across data frames:
    • total_employees = n(): Counts total staff per department.
    • left_count = sum(left == "yes"): Counts employees who left (since TRUE converts to 1 and FALSE to 0, summing gives the count).
    • turnover_rate = left_count / total_employees: Calculates the turnover ratio directly.
  • slice_max() + slice_min() + bind_rows(): Fetches the row with the highest turnover and the row with the lowest turnover, then combines them into a single result data frame. The . in slice_max(., ...) refers to the data coming down the pipe.
  • ungroup(): Removes the department grouping to prevent unexpected behavior in any follow-up operations.

Even Shorter Version (Using a Quick Mean Trick)

Since left == "yes" returns logical values (1 for TRUE, 0 for FALSE), taking the mean directly gives the turnover rate. This cuts down the summarise() step even more:

# Get top turnover department
top_turnover <- df %>%
  group_by(department) %>%
  summarise(turnover_rate = mean(left == "yes")) %>%
  slice_max(turnover_rate, n = 1)

# Get bottom turnover department
bottom_turnover <- df %>%
  group_by(department) %>%
  summarise(turnover_rate = mean(left == "yes")) %>%
  slice_min(turnover_rate, n = 1)

# Combine results
final_result <- bind_rows(top_turnover, bottom_turnover)

This keeps everything concise while maintaining readability—core Tidyverse principles!

内容的提问来源于stack exchange,提问作者Mattia Valsangiacomo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 12:12:28