基于员工与直属经理ID循环生成全层级经理数据的技术问询
Here's a practical, loop-based solution to generate the full manager hierarchy for each employee—this works even with random/prefixed IDs since the hierarchy is entirely data-driven:
Sample Data Setup
First, let's use your provided sample data to demonstrate:
employee_id = seq(1:10) manager_id = c(1,1,2,3,4,2,3,1,4,5) hr = data.frame(employee_id, manager_id)
Step 1: Build a Hierarchy-Traversing Function
We'll create a function that takes an employee ID and walks up the chain of managers until it reaches the root (where an employee is their own manager):
get_manager_hierarchy <- function(emp_id, hr_df) { hierarchy <- c() current_id <- emp_id # Keep moving up the chain until we hit the top-level manager while(TRUE) { # Get the direct manager of the current ID current_manager <- hr_df$manager_id[hr_df$employee_id == current_id] # Stop if we've reached the root (manager is themselves) if(current_manager == current_id) { break } # Add the manager to our hierarchy list and move up hierarchy <- c(hierarchy, current_manager) current_id <- current_manager } return(hierarchy) }
Testing this for employee 4 gives exactly the chain you expect: get_manager_hierarchy(4, hr) returns [3, 2, 1].
Step 2: Apply the Function to All Employees
Next, we'll add each employee's manager hierarchy to the data frame, then figure out how many manager levels we need columns for:
# Add a column with each employee's full manager chain hr$manager_hierarchy <- lapply(hr$employee_id, get_manager_hierarchy, hr_df = hr) # Find the deepest hierarchy to set our column count max_manager_levels <- max(sapply(hr$manager_hierarchy, length))
Step 3: Reshape to Wide Format (Your Desired Output)
Now we'll create dedicated columns for each manager level (managerL1, managerL2, etc.) and populate them:
# Create columns for each manager level for(level in 1:max_manager_levels) { hr[[paste0("managerL", level)]] <- sapply(hr$manager_hierarchy, function(chain) { # Use the manager at this level if it exists, else NA if(length(chain) >= level) chain[level] else NA }) } # Keep only the columns we care about (employee ID + manager levels) final_output <- hr[, c("employee_id", paste0("managerL", 1:max_manager_levels))]
Check the Result
For employee 4, the output matches your expected example perfectly:
final_output[final_output$employee_id == 4, ] # employee_id managerL1 managerL2 managerL3 # 4 4 3 2 1
Important Notes for Real-World Data
- Random/Prefixed IDs: This solution doesn't care about the format of your IDs (e.g., strings with prefixes like "EMP-123")—it only relies on matching
employee_idandmanager_idvalues, so it works with data-driven hierarchies. - Performance: This loop-based method is great for small to medium datasets. If you're working with 100k+ employees, consider using
data.tablefor recursive joins origraphfor graph traversal to speed things up. - Root Managers: Employees who are their own manager (like ID 1 in the sample) will have NA in all manager columns, which is correct since they have no higher-ups.
内容的提问来源于stack exchange,提问作者Cypress

