如何使用tidyverse简化查询员工流失率最高与最低部门的代码
简化计算员工流失率最高/最低部门的Tidyverse代码
Hey there! Your original code logic is solid, but we can streamline it drastically using Tidyverse's chained operations and built-in functions—no need for multiple intermediate data frames. Let's break down the optimized version:
Full Simplified Code
library(tidyverse) # Calculate turnover per department and grab top/bottom results turnover_summary <- df %>% group_by(department) %>% summarise( total_employees = n(), left_count = sum(left == "yes"), turnover_rate = left_count / total_employees ) %>% # Combine highest and lowest turnover rows bind_rows( slice_max(., turnover_rate, n = 1), slice_min(., turnover_rate, n = 1) ) %>% ungroup() print(turnover_summary)
Key Optimizations Explained
group_by(department): Groups the data by department so all subsequent calculations are done per department subset.summarise(): Computes all necessary metrics in one step instead of splitting across data frames:total_employees = n(): Counts total staff per department.left_count = sum(left == "yes"): Counts employees who left (sinceTRUEconverts to 1 andFALSEto 0, summing gives the count).turnover_rate = left_count / total_employees: Calculates the turnover ratio directly.
slice_max()+slice_min()+bind_rows(): Fetches the row with the highest turnover and the row with the lowest turnover, then combines them into a single result data frame. The.inslice_max(., ...)refers to the data coming down the pipe.ungroup(): Removes the department grouping to prevent unexpected behavior in any follow-up operations.
Even Shorter Version (Using a Quick Mean Trick)
Since left == "yes" returns logical values (1 for TRUE, 0 for FALSE), taking the mean directly gives the turnover rate. This cuts down the summarise() step even more:
# Get top turnover department top_turnover <- df %>% group_by(department) %>% summarise(turnover_rate = mean(left == "yes")) %>% slice_max(turnover_rate, n = 1) # Get bottom turnover department bottom_turnover <- df %>% group_by(department) %>% summarise(turnover_rate = mean(left == "yes")) %>% slice_min(turnover_rate, n = 1) # Combine results final_result <- bind_rows(top_turnover, bottom_turnover)
This keeps everything concise while maintaining readability—core Tidyverse principles!
内容的提问来源于stack exchange,提问作者Mattia Valsangiacomo
相关产品推荐
相关产品推荐

